What 341,054 public AI coding-agent runs across 29 dataset/model/scaffold groups say about where the money goes — and which of the patterns actually generalise.
We took every public dataset of AI coding-agent runs we could find — 341,054 runs from 11 datasets, spanning 2024 Llama SWE-agent runs, GPT-4o, Claude 3.5 and 3.7 Sonnet, Qwen3-Coder-480B and gpt-5-nano, on four scaffold families — and asked one question: where does the money go, and does it buy anything?
of estimated agent spend goes to the longest fifth of runs (36% under cached pricing). Range across the 29 groups: 21–72%, median 50%.
Those runs succeed at about a third the rate of the shortest fifth at the median: 19% vs 59% for Claude 3.7 Sonnet, 30% vs 57% for Open-SWE v1.0, 3% vs 22% for Llama 70B. Resolve falls from shortest to longest in 14 of the 15 groups with resolve labels, but the size of the fall ranges from 0.16× to 0.91×.
fewer resolved tasks per dollar in the longest fifth than the shortest. It is the worst bucket per task solved in 13 of those 15 groups.
Three things follow, in plain language:
You don't need to understand agents to understand this: a growing line item where two-fifths of the money buys the fewest results is a procurement problem, not a technology problem.
Three questions for your AI or engineering lead:
Claude 3.7 Sonnet, 14,374 runs, tool-call format. Runs split into five equal groups by length. Bars: share of estimated spend with cached input pricing. Line: share of runs that resolved the task. Resolve falls from the shortest to the longest fifth in 14 of the 15 groups that carry resolve labels; the exception is GPT-4o under a 30-step cap, whose shortest runs resolve nothing at all.
Loops: identical action three or more times with no state-changing action between. Blind retries: identical action immediately after an error observation. Big output: any single observation over 20,000 characters. Ended without result: context exhausted, iteration or cost limit, or cut off mid-call. Mechanical waste: cost of loop and retry steps plus the cost of dragging oversized output through remaining steps. Resolve: task passed hidden tests where the dataset records it. Cost basis: characters/4, full context re-sent each step, Sonnet-class rates; quote the percentages. Groups under 1,000 runs omitted; everything is in index.json. Ingest is capped at the first 20,000 rows of each dataset/config/split in the publisher's own streaming order, so the larger sources are a head sample, not the whole dataset. Edit thrash (the same file edited six or more times) is reported in the repo tables but is not counted as mechanical waste and not counted as a finding: it fires on over 30% of runs in 20 of the 29 groups and peaks at 78%, which makes it a description of how a scaffold edits files rather than a detection of waste. The agentic-coding-trajectories rows are that dataset's own provenance labels and resample populations already in the set — read them as replication, not as extra independent evidence, and do not pool them with their sources.
| model | scaffold | dataset | runs | med steps | loops | blind retries | big output | ended w/o result | mech waste | failed-run spend | resolve |
|---|---|---|---|---|---|---|---|---|---|---|---|
| claude-3-5-sonnet-20241022 | swe-agent/ticks | SWE-smith-trajectories | 4,033 | 17 | 0% | 1% | 10% | 0% | 3% | 0% | 41% |
| claude-3-5-sonnet-20241022 | swe-agent/tool | SWE-smith-trajectories | 5,098 | 14 | 0% | 4% | 16% | 0% | 4% | 0% | 41% |
| claude-3-5-sonnet-20241022 | swe-agent/xml | SWE-smith-trajectories | 5,198 | 15 | 0% | 2% | 16% | 0% | 4% | 0% | 42% |
| claude-3-7-sonnet-20250219 | swe-agent/ticks | SWE-smith-trajectories | 15,914 | 29 | 0% | 1% | 22% | 0% | 9% | 0% | 45% |
| claude-3-7-sonnet-20250219 | swe-agent/tool | SWE-smith-trajectories | 14,374 | 30 | 0% | 1% | 20% | 0% | 6% | 0% | 38% |
| claude-3-7-sonnet-20250219 | swe-agent/xml | SWE-smith-trajectories | 14,485 | 29 | 0% | 1% | 21% | 0% | 8% | 0% | 43% |
| kwai-klear-swe-smith-mini | mini-swe-agent | agentic-coding-trajectories | 5,000 | 24 | 0% | 0% | 0% | 11% | 0% | 8% | — |
| nebius-swe-rebench-openhands | openhands | agentic-coding-trajectories | 5,000 | 61 | 1% | 1% | 72% | 9% | 5% | 17% | 49% |
| swe-smith-claude-3-7-sonnet | swe-agent | agentic-coding-trajectories | 5,000 | 27 | 1% | 1% | 18% | 0% | 7% | 0% | 44% |
| gpt-4o-2024-08-06 | openhands | OpenHands-Sampled-Trajectories | 5,826 | 15 | 2% | 9% | 18% | 53% | 5% | 55% | 7% |
| Open-SWE v1.0 (see card) | openhands | Open-SWE-Traces | 20,000 | 54 | 1% | 1% | 35% | 0% | 1% | 0% | 44% |
| Open-SWE v1.0 (see card) | sweagent | Open-SWE-Traces | 20,000 | 67 | 4% | 1% | 54% | 0% | 10% | 0% | 48% |
| Open-SWE v1.1 (see card) | minisweagent | Open-SWE-Traces | 20,000 | 51 | 0% | 0% | 0% | 0% | 0% | 0% | — |
| Open-SWE v1.1 (see card) | openhands | Open-SWE-Traces | 20,000 | 77 | 0% | 0% | 50% | 0% | 2% | 0% | — |
| Open-SWE v1.1 (see card) | sweagent | Open-SWE-Traces | 20,000 | 76 | 0% | 2% | 55% | 0% | 2% | 0% | — |
| Open-SWE v1.2 (see card) | minisweagent | Open-SWE-Traces | 20,000 | 47 | 0% | 0% | 0% | 0% | 0% | 0% | — |
| Qwen3-Coder-480B | openhands | SWE-rebench-openhands-trajectories | 20,000 | 61 | 1% | 1% | 72% | 9% | 5% | 17% | 48% |
| Qwen3-Coder-480B | openhands | SWE-Hero-openhands-trajectories | 20,000 | 60 | 0% | 0% | 58% | 0% | 4% | 0% | — |
| Qwen3-Coder-480B | openhands | SWE-Zero-openhands-trajectories | 20,000 | 31 | 0% | 0% | 7% | 0% | 1% | 0% | — |
| swe-agent-llama-70b | swe-agent | SWE-agent-trajectories | 18,582 | 17 | 14% | 22% | 1% | 32% | 23% | 77% | 16% |
| swe-agent-llama-8b | swe-agent | SWE-agent-trajectories | 1,092 | 25 | 18% | 31% | 1% | 50% | 35% | 91% | 21% |
| qwen3-coder-30b | mini-swe-agent | mini-coder-trajs-400k | 20,000 | 30 | 1% | 1% | 0% | 2% | 0% | 2% | 19% |
| gpt-5-nano-2025-08-07 | terminus-2 | AgentTrove | 20,000 | 5 | 6% | 5% | 0% | 33% | 3% | 65% | — |
| mini-coder-1.7B | mini-swe-agent-1 | SWE-ZERO-12M-trajectories | 20,000 | 15 | 1% | 1% | 0% | 97% | 0% | 98% | — |
nvidia/Open-SWE-Traces. The dataset card does not attribute a generating model per config and the traces carry no model field, so rows are labelled by trace version and compare scaffolds within a version, never models. Every run in this dataset is a completed submission (SFT curation), so failed-run cost is zero by construction. v1.0 was filtered after release to remove runs with "git hacking" behavior.
| version | generating model | scaffold | runs | med steps | big output | mech waste | longest fifth's spend (cached) | resolve |
|---|---|---|---|---|---|---|---|---|
| v1.0 | not attributed on the card | openhands | 20,000 | 54 | 35% | 1% | 39% | 44% |
| v1.0 | not attributed on the card | sweagent | 20,000 | 67 | 54% | 10% | 40% | 48% |
| v1.1 | not attributed on the card | minisweagent | 20,000 | 51 | 0% | 0% | 37% | — |
| v1.1 | not attributed on the card | openhands | 20,000 | 77 | 50% | 2% | 28% | — |
| v1.1 | not attributed on the card | sweagent | 20,000 | 76 | 55% | 2% | 30% | — |
| v1.2 | not attributed on the card | minisweagent | 20,000 | 47 | 0% | 0% | 53% | — |
Bash-only runs are a third shorter with zero oversized output; the scaffold eliminates a waste class by construction. Its long tail is wider, so its spend is more concentrated in the longest fifth, not less. Which scaffold is cheapest per resolved task needs resolve labels v1.1 doesn't carry.
Within each group, runs are split into five equal-count buckets by step count. Two cost bases: full context re-sent every step, and cached input (re-read tokens at 10%). Resolved per $ is relative within a group. Groups with resolve labels and at least 500 runs per bucket.
| model, split | quintile | steps | spend share, no cache | spend share, cached | resolve | resolved / $ |
|---|---|---|---|---|---|---|
| gpt-4o-2024-08-06 train.raw | 1 | 3–4 | 0% | 0% | 0% | 0.00 |
| 2 | 4–10 | 1% | 1% | 5% | 2.40 | |
| 3 | 10–19 | 8% | 12% | 18% | 0.66 | |
| 4 | 19–31 | 24% | 26% | 10% | 0.13 | |
| 5 | 31–99 | 67% | 60% | 4% | 0.02 | |
| claude-3-5-sonnet-20241022 ticks | 1 | 5–11 | 2% | 4% | 40% | 3.12 |
| 2 | 11–14 | 4% | 7% | 45% | 1.81 | |
| 3 | 14–20 | 8% | 10% | 41% | 0.96 | |
| 4 | 20–30 | 16% | 17% | 44% | 0.51 | |
| 5 | 30–138 | 70% | 61% | 37% | 0.09 | |
| claude-3-7-sonnet-20250219 ticks | 1 | 7–20 | 5% | 6% | 66% | 1.65 |
| 2 | 20–26 | 8% | 10% | 55% | 0.74 | |
| 3 | 26–34 | 14% | 15% | 46% | 0.39 | |
| 4 | 34–48 | 23% | 24% | 36% | 0.17 | |
| 5 | 48–151 | 50% | 44% | 23% | 0.05 | |
| claude-3-5-sonnet-20241022 tool | 1 | 4–10 | 2% | 5% | 50% | 4.36 |
| 2 | 10–12 | 4% | 7% | 47% | 2.41 | |
| 3 | 12–17 | 7% | 10% | 40% | 1.27 | |
| 4 | 17–26 | 14% | 17% | 37% | 0.55 | |
| 5 | 26–127 | 72% | 62% | 31% | 0.09 | |
| claude-3-7-sonnet-20250219 tool | 1 | 8–20 | 4% | 6% | 59% | 1.55 |
| 2 | 20–26 | 8% | 10% | 47% | 0.65 | |
| 3 | 26–34 | 13% | 15% | 37% | 0.31 | |
| 4 | 34–48 | 24% | 24% | 28% | 0.13 | |
| 5 | 48–151 | 51% | 45% | 19% | 0.04 | |
| claude-3-5-sonnet-20241022 xml | 1 | 4–10 | 2% | 5% | 48% | 4.12 |
| 2 | 10–13 | 4% | 7% | 50% | 2.41 | |
| 3 | 13–18 | 7% | 10% | 42% | 1.23 | |
| 4 | 18–27 | 15% | 17% | 40% | 0.55 | |
| 5 | 27–127 | 72% | 61% | 31% | 0.09 | |
| claude-3-7-sonnet-20250219 xml | 1 | 6–19 | 4% | 6% | 65% | 1.70 |
| 2 | 19–26 | 8% | 10% | 53% | 0.71 | |
| 3 | 26–34 | 14% | 15% | 44% | 0.36 | |
| 4 | 34–48 | 24% | 24% | 33% | 0.16 | |
| 5 | 48–151 | 50% | 44% | 21% | 0.05 | |
| swe-agent-llama-70b train | 1 | 1–10 | 1% | 3% | 22% | 5.20 |
| 2 | 10–14 | 3% | 6% | 27% | 2.80 | |
| 3 | 14–21 | 7% | 11% | 19% | 0.87 | |
| 4 | 21–37 | 23% | 26% | 8% | 0.12 | |
| 5 | 37–398 | 64% | 55% | 3% | 0.02 | |
| Qwen3-Coder-480B SWE-rebench | 1 | 21–48 | 10% | 11% | 65% | 0.26 |
| 2 | 48–56 | 14% | 15% | 57% | 0.16 | |
| 3 | 56–66 | 18% | 18% | 49% | 0.11 | |
| 4 | 66–81 | 24% | 23% | 41% | 0.07 | |
| 5 | 81–114 | 35% | 32% | 27% | 0.03 | |
| Open-SWE v1.0 openhands | 1 | 11–39 | 7% | 9% | 57% | 0.31 |
| 2 | 39–49 | 12% | 13% | 49% | 0.15 | |
| 3 | 49–60 | 16% | 17% | 43% | 0.10 | |
| 4 | 60–75 | 23% | 23% | 37% | 0.05 | |
| 5 | 75–200 | 42% | 39% | 30% | 0.02 | |
| Open-SWE v1.0 sweagent | 1 | 8–47 | 6% | 7% | 64% | 0.27 |
| 2 | 47–60 | 11% | 12% | 53% | 0.12 | |
| 3 | 60–75 | 16% | 17% | 46% | 0.06 | |
| 4 | 75–95 | 24% | 23% | 41% | 0.04 | |
| 5 | 95–200 | 43% | 40% | 33% | 0.02 | |
| qwen3-coder-30b train | 1 | 3–19 | 4% | 6% | 30% | 2.06 |
| 2 | 19–26 | 8% | 11% | 25% | 0.86 | |
| 3 | 26–34 | 14% | 16% | 18% | 0.37 | |
| 4 | 34–48 | 23% | 24% | 14% | 0.16 | |
| 5 | 48–256 | 51% | 44% | 9% | 0.05 |
Four claims from earlier versions of this work did not survive the full run. They're recorded so nobody re-derives them from an old version of the tools, or from an old version of this page.
| dataset | models | scaffold | runs used | outcome label |
|---|---|---|---|---|
| SWE-bench/SWE-smith-trajectories (tool, xml, ticks) | Claude 3.5 / 3.7 Sonnet, GPT-4o | SWE-agent | 60,000 | resolved |
| nvidia/Open-SWE-Traces v1.0, v1.1, v1.2 | not attributed per config on the dataset card; labelled by trace version | OpenHands, SWE-agent, mini-SWE-agent | 120,000 | v1.0 only |
| nebius/SWE-agent-trajectories | Llama 3.1 8B / 70B / 405B | SWE-agent | 20,000 | target |
| nebius/SWE-rebench-openhands-trajectories | Qwen3-Coder-480B | OpenHands | 20,000 | resolved |
| nvidia/SWE-Hero, nvidia/SWE-Zero | Qwen3-Coder-480B | OpenHands | 40,000 | none |
| SWE-Gym/OpenHands-Sampled-Trajectories | GPT-4o, Claude 3.5 Sonnet | OpenHands, maxiter 30 | 6,055 | resolved |
| thoughtworks/agentic-coding-trajectories | Claude 3.7 (SWE-smith), Qwen (rebench), Klear mini | mixed | 15,000 | resolved |
| open-thoughts/AgentTrove | gpt-5-nano | terminus-2 | 20,000 | none |
| ricdomolm/mini-coder-trajs-400k | qwen3-coder-30b | mini-SWE-agent | 20,000 | verified |
| AlienKevin/SWE-ZERO-12M-trajectories | 1.7B model | mini-SWE-agent | 20,000 | none |
nebius-swe-rebench-openhands, swe-smith-claude-3-7-sonnet and kwai-klear-swe-smith-mini are the source labels thoughtworks/agentic-coding-trajectories carries for its own rows; the first matches nebius/SWE-rebench-openhands-trajectories to within sampling noise on every statistic, so treat it as a resample of a population already here rather than as independent evidence.index.json.git clone https://github.com/metermaidai/audit && cd audit python -m pip install datasets pyarrow duckdb python pipeline.py sweep --limit 20000 # streams every registered dataset into data/runs/*.parquet python pipeline.py report # DuckDB over Parquet -> report/index.md, index.json
Everything above came from public runs. The same detectors, plus the spend-side audit (over-tier models, missing prompt caching, batch-eligible jobs, unowned keys), run locally against your own traces and your Anthropic or OpenAI admin API. Nothing leaves your machine.
python metermaid_audit.py --anon # spend audit, ids hashed python trajectory_audit.py ./your-traces # loops, retries, big output, ended without result
Send the anonymized output and get your resolve-by-length curve and your position against the Index. The next edition adds contributed production data, which is the only way current frontier models get into it.
The Agent Waste Index, edition 1. metermaid, September 2026. Data, code, and every retraction are in the repo.