/methodology
The index ranks currently purchasable configurations by task dollars per score point on one named job. Two measurements sit under that table. F1 asks what a fixed capability costs over time. F2 asks how often the cheapest way to buy that capability changes hands. Everything below is derived from the archive named on /status; nothing on this page is estimated, and where a number does not exist it says so.
F1 — the cost of a fixed capability
Hold a capability fixed by picking a score threshold on one benchmark. Walk the release-date axis and keep the running minimum cost per task among configurations that clear the threshold. Every strictly cheaper qualifying release is a new step. That staircase is the curve, and its half-life is the number.
Thresholds are not a grid of round percentages. They are exactly the distinct scores the benchmark's costed points achieved, so the scan is the whole population of curves rather than a sample — and any fixed step width would be an unmeasured constant smuggled into the headline. Adjacent thresholds usually collapse onto the same shape; the label is the most demanding of them, because every lower threshold buys at least that capability.
the gate, printed with its numbers
a staircase carries a half-life only with at least 4 steps AND at least 90 days between its first and its last; anything shorter is a curve whose whole life is briefer than the number it would claim.
A threshold that fails prints the condition it failed and the number it failed against, and its staircase is still drawn.
the headline is a band
the headline is a BAND, never a single number: this pool's curves run from 25–394 days depending on the threshold, so a lone median would be a choice presented as a measurement.
Two half-life definitions are published for every curve: the endpoint definition (first and last step, the one the band is made of) and a log-linear least-squares fit over every step. The frontier panel also prints the same curve with its first step dropped and with its last step dropped, because one revised endpoint can move the number by more than half its value.
what the corpus is, and what it excludes
| snapshot | epoch_zip@506703a3c54b4021 |
|---|---|
| fetched | 2026-08-17T06:11:03Z |
| result rows | 781 |
| points that entered the scan | 764 |
| excluded, contested | 17 |
| excluded, undated | 0 |
| release-date disputes | 2 |
| cost basis: list price × tokens | 66 |
| cost basis: unknown | 698 |
| excluded for a published cost of $0 | 1 |
| models with an availability window | 171 |
a published cost of $0 is not a measured spend of zero, it is a price nobody can buy — the ':free' hazard class — and one such point would pin a threshold's frontier at $0 forever. They are excluded and counted, never plotted.
'imputed_local' is a price nobody can buy; WeirdML's local and quantized runs are invisible through the mirror and may sit on the frontier edge.
every measurer is its own channel: two measurers never share a frontier or a dominance timeline.
the x axis is the MODEL'S RELEASE DATE, not the day anyone measured it or could have bought it (§5-F1 trap 6).
purchasability
OFF by decision (S29): the filter does not enter F1, F2 or the event fold, and the surface marks the stale frontier instead.
1 of 155 valid curves (gate_all: every curve past the gate, ARC mirrors included) is held by a retired config at the corpus's now; every holder has availability data
weirdml, score ≥ 0.2318: Gemini 1.5 Flash 002 (google/gemini-1.5-flash-002), at 2026-07-31.
price_fact.valid_to IS NULL and availability.available_to IS NULL mean OPEN AT archive_as_of, not open today. A model with no availability row is UNKNOWN, not available.
cost per task of a fixed capability
capability threshold — every configuration on the staircase scores at least this
Choosing a benchmark and moving the threshold needs JavaScript. Without it this page shows one
staircase — the one closest to ale-bench's own median half-life — and the table
below lists every point on it. Every staircase, with its threshold, its steps and the pooled
bands, is machine-readable in site.json. The raw inputs they are
derived from are in dump.sqlite — the measurements in
benchmark_result and the release dates in model.released_on. The
dump carries no curve, no threshold and no half-life: those are derived here.
- frontier step
- no cheaper configuration in the corpus since — not a measurement past that date
- last step held by a configuration that has left the market
selected ale-bench score ≥ 1162.45 → 98 d · 4 steps · 358 d span · 12.5× cheaper
3 thresholds produce this same staircase; the label is the most demanding of them, because every lower threshold buys at least this capability.
the same staircase, four ways to take its half-life · ale-bench, score ≥ 1162.45
- endpoint
- 98 d — first and last step only, the definition the headline band is made of
- least squares
- 96 d — log-linear fit over every step
- drop the first step
- 81 d
- drop the last step
- 92 d
One revised endpoint can move the number by more than half its value. That is why the headline is a band and this panel exists.
Last step held by DeepSeek-V4 Flash (0731) (max) (deepseek/deepseek-v4-flash-0731, on sale at the archive's as-of).
the plotted points, as a table
| released | US$ / task | fall | configuration | model_uid |
|---|---|---|---|---|
| 2025-08-07 | 0.244948 | — | GPT-5 (high) | openai/gpt-5-20250807 |
| 2025-12-11 | 0.142947 | 1.71× | GPT-5.2 (medium) | openai/gpt-5.2-20251211 |
| 2025-12-17 | 0.091067 | 1.57× | Gemini 3 Flash Preview | google/gemini-3-flash-preview |
| 2026-07-31 | 0.019526 | 4.66× | DeepSeek-V4 Flash (0731) (max) | deepseek/deepseek-v4-flash-0731 |
ale-bench (coding) · 66 costed points · 66 thresholds → 33 distinct staircases · 15 past the gate, 18 refused · newest release 2026-07-31
Across its 15 valid curves: median 98 d · IQR 77–162 · full band 55–394 d.
In pools: baseline_2026_08_15, gate_no_arc_mirrors, gate_all.
This number depends on the threshold. Move the slider and it changes: on this pool the same method yields anything from 25–394 days depending on which capability you hold fixed. That is why F1 is published as a band.
the x axis is the MODEL'S RELEASE DATE, not the day anyone measured it or could have bought it (§5-F1 trap 6).
Measured history, not extrapolation. Every step is a model that was released and priced; the curve says nothing about what comes next.
a published cost of $0 is not a measured spend of zero, it is a price nobody can buy — the ':free' hazard class — and one such point would pin a threshold's frontier at $0 forever. They are excluded and counted, never plotted. On this corpus that is 1 point.
'imputed_local' is a price nobody can buy; WeirdML's local and quantized runs are invisible through the mirror and may sit on the frontier edge.
F2 — how often the frontier changes hands
a KICK is the moment a configuration became dominated: something at least as strong arrived at a strictly lower cost, or something strictly stronger arrived at no higher cost. Overlapping score intervals suppress the event, and a victim is kicked once.
the window ends on the NEWEST KICK, the corpus's own now — a wall-clock anchor would make the same database print a different frequency every quiet week. Frequency is counted in EVENT-DAYS, not kicks: twelve kicks in one afternoon are one day the channel spoke.
| window | days | kicks | event-days | per month | silent |
|---|---|---|---|---|---|
| 12 mo, ARC excluded | 365 | 98 | 19 | 1.58 | 94.8% |
| 12 mo, ARC included | 365 | 164 | 23 | 1.92 | 93.7% |
| 6 mo, ARC excluded | 182 | 74 | 12 | 2.01 | 93.4% |
| benchmark | configs | kicks | event-days | born dominated | suppressed by overlapping intervals |
|---|---|---|---|---|---|
| aider-polyglot | 46 | 12 | 7 | 26 | 0 |
| ale-bench | 66 | 21 | 9 | 33 | 0 |
| arc-agi | 138 | 36 | 13 | 82 | 0 |
| arc-agi-2 | 133 | 45 | 14 | 74 | 0 |
| critpt | 10 | 0 | 0 | 4 | 0 |
| cursorbench | 27 | 9 | 4 | 7 | 0 |
| deepswe | 45 | 12 | 3 | 10 | 31 |
| exploitbench | 14 | 4 | 2 | 3 | 0 |
| osworld-2 | 4 | 0 | 0 | 1 | 0 |
| proofbench | 33 | 13 | 4 | 1 | 36 |
| surface-evolver-bench | 14 | 4 | 2 | 4 | 0 |
| the-agent-company | 6 | 4 | 4 | 0 | 0 |
| weirdml | 118 | 51 | 24 | 55 | 0 |
54 of 66 configs are dominated (82%) — ale-bench, measured by Epoch AI.
dominated is not the same as bad: a dominated configuration can still be the right purchase for a reason this archive cannot see.
2 points excluded for a published cost of $0, counted over the cost F2 compares, which is per-task where the benchmark publishes one and the suite total otherwise. That is a WIDER set than F1's y axis, so this figure can exceed f1.corpus.zero_cost_excluded over the same corpus; the two are not a contradiction, they are two scopes.
saturation: the evidence, not a label
S28 (19 Aug 2026): NO saturation label is published in Phase 1a and no curve is suppressed. The evidence is printed instead, because the threshold that would separate 'saturating' from 'saturated' has not been measured, and a guessed threshold whose documented consequence is DO-NOT-PLOT would delete benchmarks from charts on the strength of a number nobody chose.
revisited at T+90 (≈2026-11-13); until then a threshold is set only if it derives from a measurement.
| benchmark | best score | ceiling | headroom | highest valid threshold | state |
|---|---|---|---|---|---|
| aider-polyglot | 88 | 100 | 12.0% | — | no label published |
| ale-bench | 2176.88 | ∞ | — | 1244.92 | no label published |
| arc-agi | 0.98 | 1 | 2.0% | 0.905 | no label published |
| arc-agi-2 | 0.925 | 1 | 7.5% | 0.7208 | no label published |
| critpt | 0.323 | 1 | 67.7% | — | no label published |
| cursorbench | 0.729 | 1 | 27.1% | 0.648 | no label published |
| deepswe | 0.7364864864864865 | 1 | 26.4% | 0.5812917594654788 | no label published |
| exploitbench | 0.738 | 1 | 26.2% | — | no label published |
| osworld-2 | 0.206 | 1 | 79.4% | — | no label published |
| proofbench | 0.78 | 1 | 22.0% | 0.54 | no label published |
| surface-evolver-bench | 0.95 | 1 | 5.0% | — | no label published |
| the-agent-company | 0.518 | 1 | 48.2% | 0.08 | no label published |
| weirdml | 0.9194 | 1 | 8.1% | 0.721 | no label published |
-
aider-polyglot— not measured: no threshold produced a staircase past the gate -
ale-bench— unbounded above (a rating, not a fraction); not defined: headroom needs a ceiling, and this benchmark has none -
critpt— not measured: no threshold produced a staircase past the gate -
exploitbench— not measured: no threshold produced a staircase past the gate -
osworld-2— not measured: no threshold produced a staircase past the gate -
surface-evolver-bench— not measured: no threshold produced a staircase past the gate
data quality: rows this archive does not stand behind
S31 (19 Aug 2026): price_fact is NOT corrected, filtered or deleted — the archive records what the catalogues actually published. Doubt is this separate, derived claim, and it is presented per INCIDENT (seller × upstream commit), never by blaming the seller whose arrival made an old broken row visible again.
31 of 6702 price rows (0.46%) carry a suspect mark. They are published rather than deleted, and every consumer of the dumps is expected to filter them before computing anything.
Identity, checked on every export so it can actually fail: union == M + R + C − MR − MC − RC + MRC, with each rule set built in its own pass over the ledger so the equation can actually fail — 31 == 15 + 9 + 10 − 3 − 0 − 0 + 0, holds.
| rule | definition | rows |
|---|---|---|
m magnitude two-tail | 1 ≤ price_micro ≤ 999 (a nonzero price under $0.001/Mtok) or price_micro ≥ 1e9 (at or above $1000/Mtok) measured: 15 of 15 independently confirmed wrong, and the nearest legitimate row sits 2.8× above the top of the bad band. | 15 |
r residue ratio | the parser recorded what rounding threw away: residue ≠ 0 and (price_micro = 0 or 1e6·|residue| > price_micro), in exact rational arithmetic a WITNESS, not a detector: its recall over the classified upstream errors is 3 of 24, because an over-scale by a power of ten on a round decimal lands on a whole µUSD and leaves no residue at all. That blindness is why the third rule is a hand-curated file. | 9 |
c curated incident | an entry in suspects.toml, each carrying a written rationale and its evidence, each required to select exactly one price_fact row a stale entry marks nothing while the count still claims it did, so a curated entry that selects zero or several rows fails the export outright. | 10 |
Overlaps: 3 rows fire on both magnitude and residue, 0 on magnitude and curation, 0 on residue and curation, 0 on all three.
100% precision is a MEASURED PROPERTY of this corpus, not a guarantee. A future error with a mid-sized residue would land in the empty band between the benign artefacts and the real ones, and would want a re-measurement rather than a silent pass.
the curated incidents
An incident is a seller and one upstream commit instant — never the seller whose arrival made an old broken row visible again. A months-old wrong row prints a fresh headline every time a new, blameless seller lists the same model, so attributing the fault to the entering seller names the wrong party.
helicone-2025-10-23-mistral-off-band
helicone · upstream commit 2025-10-23T20:43:27Z · still open at the archive's as-of, 297.5 days and counting · 4 rows
| model_uid | meter | published | why it is marked |
|---|---|---|---|
mistral/mistral-nemo | input_text | $20/Mtok | $20/Mtok against a five-seller consensus of $0.15 — 133x, not a clean power of ten, in a catalogue whose other 94 comparable rows sit at median ratio 1.00. No correction observed through 2026-08-17, the archive's as-of date. docs/DECISIONS.md 2026-08-19, S31 measurement round; drives NETSPLIT 25cb43927a99. |
mistral/mistral-nemo | output_text | $40/Mtok | $40/Mtok against a consensus of $0.15 — 267x, log10 closeness 0.426, so not a decimal shift. No correction observed through 2026-08-17, the archive's as-of date, and the row was still driving NETSPLIT events in the archive's final fortnight. docs/DECISIONS.md 2026-08-19, S31 measurement round; drives NETSPLIT fc06d4725f67, 3cfc8d37f133, 330f15318621. |
mistral/mistral-small | input_text | $75/Mtok | $75/Mtok against a consensus of $1.00 — 75x, log10 closeness 0.125. No correction observed through 2026-08-17, the archive's as-of date. Note the consensus side of this model carries its own identity smear (11 floating aliases across two namespaces), which is why the entry names the row and not a corrected price. docs/DECISIONS.md 2026-08-19, S31 measurement round; drives NETSPLIT 5a80a91c2525, afca45144ce9. |
mistral/mistral-small | output_text | $200/Mtok | $200/Mtok against a consensus of $3.00 — 67x, log10 closeness 0.176. No correction observed through 2026-08-17, the archive's as-of date. docs/DECISIONS.md 2026-08-19, S31 measurement round; drives NETSPLIT b41cfc097a65, 2594c7fbf79f. |
oci-2025-08-07-decimal-shift
oracle-oci · upstream commit 2025-08-07T16:45:17Z · still open at the archive's as-of, 374.6 days and counting · 2 rows
| model_uid | meter | published | why it is marked |
|---|---|---|---|
xai/grok-3 | output_text | $0.15/Mtok | $0.15/Mtok against a five-seller consensus of $15.00 — exactly 10^-2, while the same listing's input_text meter is correct at ratio 1.000. A two-cell decimal error in litellm provider 'oci'; no correction observed through 2026-08-17, the archive's as-of date. docs/DECISIONS.md 2026-08-19, S31 measurement round; oci book-wide 32 comparable era-rows, median ratio 1.00, only these two off-band. |
xai/grok-4-0709 | output_text | $0.15/Mtok | $0.15/Mtok against a five-seller consensus of $15.00 — exactly 10^-2, same commit and same two-cell error as xai/grok-3; the listing's input_text meter is correct. docs/DECISIONS.md 2026-08-19, S31 measurement round; still open at the 2026-08-17 archive as-of (374+ days), driver of NETSPLIT events 1b99d94cf7a2, 96c32ec44388, 4daa48899afb, cc4bad402b98. |
watsonx-2025-10-06-bad-import
ibm-watsonx · upstream commit 2025-10-06T19:33:05Z · corrected upstream after 11 days, on 2025-10-17 · 4 rows
| model_uid | meter | published | why it is marked |
|---|---|---|---|
mistral/mistral-small-2503 | input_text | $200/Mtok | $200/Mtok, 200x a three-seller consensus; corrected upstream 11 days later to $0.10/Mtok (÷2000). Part of one litellm commit that wrote 8 wrong rows over 4 models. docs/DECISIONS.md 2026-08-19, S31 measurement round; window closes 2025-10-17T18:51:22Z, the incident's own correction commit. |
mistral/mistral-small-2503 | output_text | $600/Mtok | $600/Mtok, 200x consensus; corrected upstream 11 days later to $0.30/Mtok (÷2000). Same commit as the input_text row. docs/DECISIONS.md 2026-08-19, S31 measurement round; window closes 2025-10-17T18:51:22Z. |
mistral/pixtral-12b-2409 | input_text | $150/Mtok | $150/Mtok, exactly 1000x a four-seller consensus; corrected upstream 11 days later to $0.35/Mtok. The full Class-1 signature — single seller, clean power of ten, later corrected. docs/DECISIONS.md 2026-08-19, S31 measurement round; drives NETSPLIT 8ed462c8dfd6. |
mistral/pixtral-12b-2409 | output_text | $150/Mtok | $150/Mtok, exactly 1000x consensus; corrected upstream 11 days later to $0.35/Mtok. Same commit as the input_text row. docs/DECISIONS.md 2026-08-19, S31 measurement round; drives NETSPLIT e1f9455d63b8. |
coverage of the check itself
- source refs
- 7349
- cited observations
- 2896
- rows citing nothing
- 0
- dangling source refs
- 0
- matched pairs
- 7325
- micro fallback pairs
- 0
- ambiguous micro pairs
- 0
- rows without comparable entry
- 16
- rows with residue
- 71
- rows residue cleared
- 62
- rows with several firing readings
- 0
- rows with unreadable residue
- 0
A row without a comparable reading is reported as unchecked, not as clean. That is the safe direction and it is counted here rather than hidden.
licence and attribution
The derived dataset is published under
CC-BY-4.0. Attribute
as: SCROLLBACK (mirc.tech), CC-BY-4.0.
CC-BY-4.0 covers this derived dataset. Upstream attribution is a condition: the measurements in benchmark_result derive entirely from Epoch AI's CC-BY-4.0 benchmark corpus (the descriptive benchmark and benchmark_version rows are this archive's own curated registry, not Epoch data); the price layer derives from the LiteLLM (MIT) and models.dev (MIT) catalogues, and NO price_fact row cites a vendor page; official vendor DEPRECATION pages back 44 availability rows and the 44 event rows derived from them. See DATA-LICENSE.
- Epoch AI, Benchmarking Hub (CC-BY-4.0) — the sole source of every measurement in
benchmark_result, all 2343 rows of it. The descriptivebenchmarkandbenchmark_versiontables are this repository's own curated registry, not Epoch data; attributing them upstream would be over-attribution. - LiteLLM model catalogue (MIT), excluding the repository's
enterprise/directory, which this project does not read. - models.dev catalogue (MIT).
- Official vendor deprecation pages (developers.openai.com, platform.claude.com) back
44
availabilityrows and the 44eventrows derived from them — not the layers, which are 1554 and 3588 rows. Only dates and retirement kinds are reproduced; no page is republished, and no price row cites a vendor page. - The WeirdML CSV carries no licence statement upstream. It is archived locally and is never redistributed; it produces no derived row. The quality axis comes from Epoch's CC-BY-4.0 corpus, including the WeirdML measurements Epoch mirrors under that licence.
- The raw archive is not part of the dataset. Archived source bodies are stored for measurement and reproducibility and are not redistributed; the evidence handles in the dumps name archive rows, not resolvable joins.
- The dumps republish upstream's own citation strings.
benchmark_result.evidence_urlis Epoch's citation, verbatim, on 1245 of 2343 rows. Of those, 336 point atarcprize.organd 318 point athtihle.github.io;event.source_urlcarries 32 ofarcprize.organd 48 ofhtihle.github.io. None of those URLs was fetched to build this dataset: the strings travel exactly as Epoch published them, and they are citations rather than data collected from those pages.
The full text travels with the dumps as DATA-LICENSE,
served beside them — CC-BY-4.0 §3(a) asks that the notice reach whoever takes the data, and a
file that is only in the repository does not. Source code is separate and is MIT.