Preview · 18 tasks · N=3 seeds

Not every task needs a frontier model.

ShallowSWE selects tasks frontier agents should handle easily: everyday fixes, artifacts, workflow updates, and repo operations. That is deliberate. If frontier models struggle, the suite is measuring maximum SWE capability again; if they do not, the ranking can focus on cost per verified success.

Depth gaugemeasured
$0.020$0.050$0.100$0.200SURFACE · CHEAPEST $ / SUCCESS16× SPREADGPT-5.6 Luna low$0.021GLM 5.2 high$0.029Kimi K2.7 Code$0.049Grok 4.5 high$0.051Sonnet 5 low$0.073Opus 4.8 low$0.080GPT-5.5 low$0.133GPT-5.6 Sol low$0.150GPT-5.5 med$0.232Gemini 3.5 Flash med$0.326
Cost per successful completion · log scale · Representative models · 10 of 16 shown
Supported by
Kaggle

Sponsorship funds compute for evaluation runs. Sponsors have no role in task selection, scoring, or rankings — every displayed number traces to the public data manifest.

The leaderboard

Cost per verified success

Each row gives one model one task. When the model says it is done, hidden checks run; it keeps fixing until they pass or a cap stops it. That whole attempt is a repair loop — and every model dollar of it counts, failed tries included.

853 / 864 scored repair loops passed in the full basket; failures stay in the cost numerator.how it’s measured ↓
Models16/16
Model filter
applies to every chart
Measured leaderboard for the selected basket weights, sortable
Model
1GPT-5.6 LunalowOpenAI
$0.021$0.02496.4%22.7K6.2
2GLM 5.2highZ.ai
$0.029$0.04092.6%53.2K11.0
3Kimi K2.7 CodefixedMoonshot
$0.049$0.07890.7%117K14.1
4Grok 4.5highxAI
$0.051$0.05894.6%40.3K7.4
5Claude Sonnet 5lowAnthropic
$0.073$0.12790.8%66.7K8.5
6Claude Opus 4.8lowAnthropic
$0.080$0.09290.8%24.7K5.9
7Claude Sonnet 5medAnthropic
$0.081$0.10296.3%47.5K9.2
8GPT-5.6 TerralowOpenAI
$0.082$0.09594.4%33.1K7.1
9Claude Fable 5lowAnthropic
$0.117$0.13696.3%15.3K4.6
10Claude Opus 4.8medAnthropic
$0.124$0.14898.2%38.4K7.5
11GPT-5.5lowOpenAI
$0.133$0.153100.0%26.0K7.5
12GPT-5.6 SollowOpenAI
$0.150$0.164100.0%28.6K6.8
13GPT-5.5medOpenAI
$0.232$0.35898.2%61.0K9.3
14InklinghighThinking Machines
$0.237$0.26974.1%712K31.6
15Gemini 3.5 FlashmedGoogle
$0.326$0.38090.8%308K24.9
16Kimi K3maxMoonshot
$0.375$0.45387.1%353K15.3

The same work at two prices

Cost per success blends two things: how much work a model needs, and what its tokens cost. The chart below separates them. Each model’s measured token usage is priced twice — at its own list price, the number above, and at one shared reference rate that strips pricing out and leaves token efficiency. The gap between the two is the sticker premium: what the list price adds on top of the work itself.

The sticker premium
Cost per success priced twice: at each model’s list price, and at one shared reference rate
at reference ratesat list price
$0.0050$0.010$0.020$0.050$0.100$0.200$0.500$1.00Fable 5 low12×GPT-5.6 Sol low7.6×GPT-5.5 low6.7×GPT-5.5 med6.4×Opus 4.8 med6.0×Opus 4.8 low5.9×GPT-5.6 Terra low3.8×Sonnet 5 med3.4×Kimi K3 max3.1×Sonnet 5 low3.1×Grok 4.5 high2.2×Gemini 3.5 Flash med1.9×GPT-5.6 Luna low1.5×Inkling high1.1×GLM 5.2 high1.0×Kimi K2.7 Code0.9×
every model’s measured token usage repriced at z-ai/glm-5.2 list rates (openrouter 2026-07-18) · sorted by premium · log scale · the default anchor is the cheapest open-weight sheet on the panel · a premium over the reference token rates, not a serving margin — models may differ in what they cost to run
Behind the numbersWhere the dollar goes · token price and turns vs cost · task-level breakdown
Where the dollar goes
Mean cost of one repair loop, split by token type
Cache readsCache writesFresh inputOutput
$0$0.100$0.200$0.300$0.400GPT-5.6 Luna low$0.021GLM 5.2 high$0.028Kimi K2.7 Code$0.049Grok 4.5 high$0.051Sonnet 5 low$0.073Opus 4.8 low$0.080GPT-5.6 Terra low$0.081Sonnet 5 med$0.081Fable 5 low$0.117Opus 4.8 med$0.124GPT-5.5 low$0.133GPT-5.6 Sol low$0.150Inkling high$0.203GPT-5.5 med$0.232Gemini 3.5 Flash med$0.326Kimi K3 max$0.375
one bar per model-effort row, cheapest first · basket-weighted mean $ per scored repair loop · fresh input excludes cache reads and writes · reconstructed from token means × openrouter 2026-07-18 price sheet
Token price against cost
Cost per success against output-token list price
OUTPUT PRICE / 1M TOKENS - LOGCOST PER SUCCESS - LOG▲ cheaper per task · ◀ cheaper output tokens$5.00$10.00$20.00$50.00$0.020$0.050$0.100$0.200$0.500Fable 5 lowSonnet 5 lowSonnet 5 medOpus 4.8 lowOpus 4.8 medGPT-5.5 lowGPT-5.5 medGPT-5.6 Luna lowGPT-5.6 Terra lowGPT-5.6 Sol lowGemini 3.5 Flash medGLM 5.2 highKimi K2.7 CodeKimi K3 maxGrok 4.5 highInkling high
one dot per model-effort row · x: list output-token price · y: cost per success · bubble size: tokens per verified success · dashed cuts are panel medians; the shaded quadrant beats both
Task-level breakdown
Cost and tokens per success, task by task
ArtifactProduce checked files
$0.010$0.020$0.050$0.100$0.200$0.500$1.00$2.00▲ cheaperS1S2M1M2L1L2GPT-5.6Gemini
CodeRepair code paths
$0.010$0.020$0.050$0.100$0.200$0.500$1.00$2.00▲ cheaperS1S2M1M2L1L2GPT-5.6Inkling
WorkflowOperate repo state
$0.010$0.020$0.050$0.100$0.200$0.500$1.00$2.00▲ cheaperS1S2M1M2L1L2GPT-5.5GPT-5.6

basket-weighted · 18 tasks · N=3 seeds per task/model · 95% CIs from task-resampling bootstrap, overlapping CIs are ties · mini-swe-agent · openrouter 2026-07-18 · Solved = verified successes / scored repair loops · scored failures and cap hits included in CPSC · infra exclusions tracked in raw rows and retried for coverage

DeepSWE comparison

The deep end against the shallow end

The same models, matched effort for effort: ranked by DeepSWE pass@1 on hard tasks, and by measured cost per success here. Watch the order flip.

Rank translation
DeepSWE ability rank against ShallowSWE price rank
DEEP ENDpass@1 · DeepSWESHALLOW END$ / success · ShallowSWE#1#2#3#4#5#6#7#8#9#10#11#12#13#14#15Kimi K3 max68.5%Kimi K3 max$0.375Fable 5 low59.6%Fable 5 low$0.117GPT-5.5 med54.0%GPT-5.5 med$0.232Grok 4.5 high53.8%Grok 4.5 high$0.051Opus 4.8 med48.7%Opus 4.8 med$0.124GPT-5.6 Sol low45.4%GPT-5.6 Sol low$0.150Opus 4.8 low40.8%Opus 4.8 low$0.080Sonnet 5 med39.8%Sonnet 5 med$0.081Gemini 3.5 Flash med37.4%Gemini 3.5 Flash med$0.326GLM 5.2 high36.3%GLM 5.2 high$0.029Kimi K2.7 Code30.5%Kimi K2.7 Code$0.049Sonnet 5 low30.5%Sonnet 5 low$0.073GPT-5.5 low27.0%GPT-5.5 low$0.133GPT-5.6 Terra low24.1%GPT-5.6 Terra low$0.082GPT-5.6 Luna low1.5%GPT-5.6 Luna low$0.021
rank 1 = highest DeepSWE pass@1 on the left, lowest ShallowSWE CPSC on the right · dashed line = a family’s low-effort variant

effort-matched rows · deep: measured from DeepSWE v1.1 leaderboard, 2026-07-17T08:18:55.870582+00:00; intelligence axis = pass@1 · shallow: measured from the selected ShallowSWE CPSC basket

The tasks

Task set

18 original tasks: three kinds of everyday work, each in small, medium, and large. Built for the Pier harness.

Artifact

Produce checked files

Transform logs, CSVs, env files, and support data into required artifacts.

  • Small · Env flags to JSON · Extract error fields
  • Medium · Payout reconciliation · Access log incidents
  • Large · Billing revenue rollup · Support SLA business hours

Code

Repair code paths

Fix regressions, add coverage, split logic, and preserve user-facing behavior.

  • Small · Invoice duplicate regression · Retry error fallback
  • Medium · Auth token expiry regression · Split notification renderer
  • Large · Settings cache invalidation · Invoice multi-source merge

Workflow

Operate repo state

Apply status updates, reconcile tickets, merge config branches, and preserve state.

  • Small · Post build status · Ticket from bug report
  • Medium · Ignored config flag · Ticket update without duplicate
  • Large · Merge config branches · Config key rollover

How a task earns its spot: admitted by a frozen frontier ceiling, sized by a measured floor panel, and written from scratch — never adapted from public repos or benchmarks. The full admission and quality gates are in the methodology.

Method

How the score is measured

A scored row is one model working one task from a fixed starting seed. Each time the agent declares done, the benchmark runs hidden programmatic tests. If they fail, the same agent keeps working until success or a cap. CPSC measures model API spend, not developer time or total cost of ownership. Read the full methodology →

the metric · realized CPSC
CPSC=total scored repair-loop spendverified successes

Weighted views use the same ratio form: weighted mean spend divided by weighted solve rate. The site does not average per-task CPSC values directly. This preview charges failures their actual spend (realized CPSC); the v1 headline will be reference-budget CPSC, which needs calibrated task budgets that don't exist yet.

  • Scored by hidden tests: no LLM judges or human grading; failures receive only coarse, non-oracle feedback, never the failing check
  • One model per row: no fallbacks, ensembles, or handoffs to a stronger model mid-task
  • Controlled scaffold: one agent harness and prompt template, with only provider-required configuration differences
  • Giving up still costs: hitting a spend, verifier-submission, or agent-step cap scores as a failure, and its spend stays in the bill
  • Prices are pinned: dollars are token counts times the dated openrouter price sheet, downloadable below
  • Outages don't count against models: provider and harness errors are retried, not scored
  • Read close ranks as ties: per-task samples are small; exact attempt counts ship in the data files

Measured. Every ShallowSWE number on this page comes from the bounded repair-loop rows, with dollar values derived from the dated openrouter 2026-07-18 price sheet. All exports, price sheets, and licenses live on the data page.

repair-loop rows