Preview · 18 tasks · N=3 seeds
Not every task needs a frontier model.
ShallowSWE selects tasks frontier agents should handle easily: everyday fixes, artifacts, workflow updates, and repo operations. That is deliberate. If frontier models struggle, the suite is measuring maximum SWE capability again; if they do not, the ranking can focus on cost per verified success.
Cost per verified success
Each row gives one model one task. When the model says it is done, hidden checks run; it keeps fixing until they pass or a cap stops it. That whole attempt is a repair loop — and every model dollar of it counts, failed tries included.
Models16/16
| Model | |||||
|---|---|---|---|---|---|
1GPT-5.6 LunalowOpenAI | $0.021 | $0.024 | 96.4% | 22.7K | 6.2 |
2GLM 5.2highZ.ai | $0.029 | $0.040 | 92.6% | 53.2K | 11.0 |
3Kimi K2.7 CodefixedMoonshot | $0.049 | $0.078 | 90.7% | 117K | 14.1 |
4Grok 4.5highxAI | $0.051 | $0.058 | 94.6% | 40.3K | 7.4 |
5Claude Sonnet 5lowAnthropic | $0.073 | $0.127 | 90.8% | 66.7K | 8.5 |
6Claude Opus 4.8lowAnthropic | $0.080 | $0.092 | 90.8% | 24.7K | 5.9 |
7Claude Sonnet 5medAnthropic | $0.081 | $0.102 | 96.3% | 47.5K | 9.2 |
8GPT-5.6 TerralowOpenAI | $0.082 | $0.095 | 94.4% | 33.1K | 7.1 |
9Claude Fable 5lowAnthropic | $0.117 | $0.136 | 96.3% | 15.3K | 4.6 |
10Claude Opus 4.8medAnthropic | $0.124 | $0.148 | 98.2% | 38.4K | 7.5 |
11GPT-5.5lowOpenAI | $0.133 | $0.153 | 100.0% | 26.0K | 7.5 |
12GPT-5.6 SollowOpenAI | $0.150 | $0.164 | 100.0% | 28.6K | 6.8 |
13GPT-5.5medOpenAI | $0.232 | $0.358 | 98.2% | 61.0K | 9.3 |
14InklinghighThinking Machines | $0.237 | $0.269 | 74.1% | 712K | 31.6 |
15Gemini 3.5 FlashmedGoogle | $0.326 | $0.380 | 90.8% | 308K | 24.9 |
16Kimi K3maxMoonshot | $0.375 | $0.453 | 87.1% | 353K | 15.3 |
The same work at two prices
Cost per success blends two things: how much work a model needs, and what its tokens cost. The chart below separates them. Each model’s measured token usage is priced twice — at its own list price, the number above, and at one shared reference rate that strips pricing out and leaves token efficiency. The gap between the two is the sticker premium: what the list price adds on top of the work itself.
▶Behind the numbersWhere the dollar goes · token price and turns vs cost · task-level breakdown
basket-weighted · 18 tasks · N=3 seeds per task/model · 95% CIs from task-resampling bootstrap, overlapping CIs are ties · mini-swe-agent · openrouter 2026-07-18 · Solved = verified successes / scored repair loops · scored failures and cap hits included in CPSC · infra exclusions tracked in raw rows and retried for coverage
The deep end against the shallow end
The same models, matched effort for effort: ranked by DeepSWE pass@1 on hard tasks, and by measured cost per success here. Watch the order flip.
effort-matched rows · deep: measured from DeepSWE v1.1 leaderboard, 2026-07-17T08:18:55.870582+00:00; intelligence axis = pass@1 · shallow: measured from the selected ShallowSWE CPSC basket
Task set
18 original tasks: three kinds of everyday work, each in small, medium, and large. Built for the Pier harness.
Artifact
Transform logs, CSVs, env files, and support data into required artifacts.
- Small · Env flags to JSON · Extract error fields
- Medium · Payout reconciliation · Access log incidents
- Large · Billing revenue rollup · Support SLA business hours
Code
Fix regressions, add coverage, split logic, and preserve user-facing behavior.
- Small · Invoice duplicate regression · Retry error fallback
- Medium · Auth token expiry regression · Split notification renderer
- Large · Settings cache invalidation · Invoice multi-source merge
Workflow
Apply status updates, reconcile tickets, merge config branches, and preserve state.
- Small · Post build status · Ticket from bug report
- Medium · Ignored config flag · Ticket update without duplicate
- Large · Merge config branches · Config key rollover
How a task earns its spot: admitted by a frozen frontier ceiling, sized by a measured floor panel, and written from scratch — never adapted from public repos or benchmarks. The full admission and quality gates are in the methodology.
How the score is measured
A scored row is one model working one task from a fixed starting seed. Each time the agent declares done, the benchmark runs hidden programmatic tests. If they fail, the same agent keeps working until success or a cap. CPSC measures model API spend, not developer time or total cost of ownership. Read the full methodology →
Weighted views use the same ratio form: weighted mean spend divided by weighted solve rate. The site does not average per-task CPSC values directly. This preview charges failures their actual spend (realized CPSC); the v1 headline will be reference-budget CPSC, which needs calibrated task budgets that don't exist yet.
- Scored by hidden tests: no LLM judges or human grading; failures receive only coarse, non-oracle feedback, never the failing check
- One model per row: no fallbacks, ensembles, or handoffs to a stronger model mid-task
- Controlled scaffold: one agent harness and prompt template, with only provider-required configuration differences
- Giving up still costs: hitting a spend, verifier-submission, or agent-step cap scores as a failure, and its spend stays in the bill
- Prices are pinned: dollars are token counts times the dated openrouter price sheet, downloadable below
- Outages don't count against models: provider and harness errors are retried, not scored
- Read close ranks as ties: per-task samples are small; exact attempt counts ship in the data files
Measured. Every ShallowSWE number on this page comes from the bounded repair-loop rows, with dollar values derived from the dated openrouter 2026-07-18 price sheet. All exports, price sheets, and licenses live on the data page.
repair-loop rowsEverything behind the numbers
The five-minute version: the repair loop, the metric family, and the calibration gates.
→Web version and PDF of the white paper, with version history and citation.
→The Kaggle development shakedown is complete; report-grade execution remains pending.
→Row-level exports, price sheets, licenses, and provenance for every number shown.
→