mirc.tech — /methodology archive as of 2026-08-17T07:36:29Z · schema 2

/methodology

The index ranks currently purchasable configurations by task dollars per score point on one named job. Two measurements sit under that table. F1 asks what a fixed capability costs over time. F2 asks how often the cheapest way to buy that capability changes hands. Everything below is derived from the archive named on /status; nothing on this page is estimated, and where a number does not exist it says so.

F1 — the cost of a fixed capability

Hold a capability fixed by picking a score threshold on one benchmark. Walk the release-date axis and keep the running minimum cost per task among configurations that clear the threshold. Every strictly cheaper qualifying release is a new step. That staircase is the curve, and its half-life is the number.

Thresholds are not a grid of round percentages. They are exactly the distinct scores the benchmark's costed points achieved, so the scan is the whole population of curves rather than a sample — and any fixed step width would be an unmeasured constant smuggled into the headline. Adjacent thresholds usually collapse onto the same shape; the label is the most demanding of them, because every lower threshold buys at least that capability.

the gate, printed with its numbers

a staircase carries a half-life only with at least 4 steps AND at least 90 days between its first and its last; anything shorter is a curve whose whole life is briefer than the number it would claim.

A threshold that fails prints the condition it failed and the number it failed against, and its staircase is still drawn.

the headline is a band

the headline is a BAND, never a single number: this pool's curves run from 25–394 days depending on the threshold, so a lone median would be a choice presented as a measurement.

Two half-life definitions are published for every curve: the endpoint definition (first and last step, the one the band is made of) and a log-linear least-squares fit over every step. The frontier panel also prints the same curve with its first step dropped and with its last step dropped, because one revised endpoint can move the number by more than half its value.

what the corpus is, and what it excludes

The quality axis behind every curve on this site.
snapshotepoch_zip@506703a3c54b4021
fetched2026-08-17T06:11:03Z
result rows781
points that entered the scan764
excluded, contested17
excluded, undated0
release-date disputes2
cost basis: list price × tokens 66
cost basis: unknown698
excluded for a published cost of $0 1
models with an availability window 171

a published cost of $0 is not a measured spend of zero, it is a price nobody can buy — the ':free' hazard class — and one such point would pin a threshold's frontier at $0 forever. They are excluded and counted, never plotted.

'imputed_local' is a price nobody can buy; WeirdML's local and quantized runs are invisible through the mirror and may sit on the frontier edge.

every measurer is its own channel: two measurers never share a frontier or a dominance timeline.

the x axis is the MODEL'S RELEASE DATE, not the day anyone measured it or could have bought it (§5-F1 trap 6).

purchasability

OFF by decision (S29): the filter does not enter F1, F2 or the event fold, and the surface marks the stale frontier instead.

1 of 155 valid curves (gate_all: every curve past the gate, ARC mirrors included) is held by a retired config at the corpus's now; every holder has availability data

weirdml, score ≥ 0.2318: Gemini 1.5 Flash 002 (google/gemini-1.5-flash-002), at 2026-07-31.

price_fact.valid_to IS NULL and availability.available_to IS NULL mean OPEN AT archive_as_of, not open today. A model with no availability row is UNKNOWN, not available.

cost per task of a fixed capability

Choosing a benchmark and moving the threshold needs JavaScript. Without it this page shows one staircase — the one closest to ale-bench's own median half-life — and the table below lists every point on it. Every staircase, with its threshold, its steps and the pooled bands, is machine-readable in site.json. The raw inputs they are derived from are in dump.sqlite — the measurements in benchmark_result and the release dates in model.released_on. The dump carries no curve, no threshold and no half-life: those are derived here.

Cost per task of a fixed capability on ale-bench, score ≥ 1162.45ale-bench, score ≥ 1162.45: 4 steps on the staircase, 2025-08-07 to 2026-07-31, 0.244948 down to 0.019526 US dollars per task. Endpoint half-life 98 days over 358 days, a 12.5-fold fall. The x axis is the model's release date; the y axis is US dollars per task on a logarithmic scale. The same points are listed in the table below the chart.$0.001$0.003$0.01$0.03$0.1$0.3$1$3May 25Sep 25Jan 26May 26US$ / task (log)model release dateGPT-5 (high) — released 2025-08-07, $0.244948 per taskGPT-5.2 (medium) — released 2025-12-11, $0.142947 per taskGemini 3 Flash Preview — released 2025-12-17, $0.091067 per taskDeepSeek-V4 Flash (0731) (max) — released 2026-07-31, $0.019526 per task
  • frontier step
  • no cheaper configuration in the corpus since — not a measurement past that date
  • last step held by a configuration that has left the market

selected ale-bench score ≥ 1162.45 98 d · 4 steps · 358 d span · 12.5× cheaper

3 thresholds produce this same staircase; the label is the most demanding of them, because every lower threshold buys at least this capability.

the same staircase, four ways to take its half-life · ale-bench, score ≥ 1162.45
endpoint
98 d — first and last step only, the definition the headline band is made of
least squares
96 d — log-linear fit over every step
drop the first step
81 d
drop the last step
92 d

One revised endpoint can move the number by more than half its value. That is why the headline is a band and this panel exists.

Last step held by DeepSeek-V4 Flash (0731) (max) (deepseek/deepseek-v4-flash-0731, on sale at the archive's as-of).

the plotted points, as a table
Every plotted point: ale-bench, score ≥ 1162.45. Cost is US dollars per task, exactly as the archive holds it; the fall column is against the step above.
releasedUS$ / taskfallconfigurationmodel_uid
2025-08-070.244948GPT-5 (high)openai/gpt-5-20250807
2025-12-110.1429471.71×GPT-5.2 (medium)openai/gpt-5.2-20251211
2025-12-170.0910671.57×Gemini 3 Flash Previewgoogle/gemini-3-flash-preview
2026-07-310.0195264.66×DeepSeek-V4 Flash (0731) (max)deepseek/deepseek-v4-flash-0731

ale-bench (coding) · 66 costed points · 66 thresholds → 33 distinct staircases · 15 past the gate, 18 refused · newest release 2026-07-31

Across its 15 valid curves: median 98 d · IQR 77–162 · full band 55–394 d.

In pools: baseline_2026_08_15, gate_no_arc_mirrors, gate_all.

This number depends on the threshold. Move the slider and it changes: on this pool the same method yields anything from 25–394 days depending on which capability you hold fixed. That is why F1 is published as a band.

the x axis is the MODEL'S RELEASE DATE, not the day anyone measured it or could have bought it (§5-F1 trap 6).

Measured history, not extrapolation. Every step is a model that was released and priced; the curve says nothing about what comes next.

a published cost of $0 is not a measured spend of zero, it is a price nobody can buy — the ':free' hazard class — and one such point would pin a threshold's frontier at $0 forever. They are excluded and counted, never plotted. On this corpus that is 1 point.

'imputed_local' is a price nobody can buy; WeirdML's local and quantized runs are invisible through the mirror and may sit on the frontier edge.

F2 — how often the frontier changes hands

a KICK is the moment a configuration became dominated: something at least as strong arrived at a strictly lower cost, or something strictly stronger arrived at no higher cost. Overlapping score intervals suppress the event, and a victim is kicked once.

the window ends on the NEWEST KICK, the corpus's own now — a wall-clock anchor would make the same database print a different frequency every quiet week. Frequency is counted in EVENT-DAYS, not kicks: twelve kicks in one afternoon are one day the channel spoke.

Kick frequency over three windows. The newest kick in the corpus is 2026-07-31.
window days kicks event-days per month silent
12 mo, ARC excluded 365 98 19 1.58 94.8%
12 mo, ARC included 365 164 23 1.92 93.7%
6 mo, ARC excluded 182 74 12 2.01 93.4%
Per benchmark. A configuration that arrives already dominated is not a kick, and a victim is kicked once.
benchmark configs kicks event-days born dominated suppressed by overlapping intervals
aider-polyglot 46 12 7 26 0
ale-bench 66 21 9 33 0
arc-agi 138 36 13 82 0
arc-agi-2 133 45 14 74 0
critpt 10 0 0 4 0
cursorbench 27 9 4 7 0
deepswe 45 12 3 10 31
exploitbench 14 4 2 3 0
osworld-2 4 0 0 1 0
proofbench 33 13 4 1 36
surface-evolver-bench 14 4 2 4 0
the-agent-company 6 4 4 0 0
weirdml 118 51 24 55 0

54 of 66 configs are dominated (82%) — ale-bench, measured by Epoch AI.

dominated is not the same as bad: a dominated configuration can still be the right purchase for a reason this archive cannot see.

2 points excluded for a published cost of $0, counted over the cost F2 compares, which is per-task where the benchmark publishes one and the suite total otherwise. That is a WIDER set than F1's y axis, so this figure can exceed f1.corpus.zero_cost_excluded over the same corpus; the two are not a contradiction, they are two scopes.

saturation: the evidence, not a label

S28 (19 Aug 2026): NO saturation label is published in Phase 1a and no curve is suppressed. The evidence is printed instead, because the threshold that would separate 'saturating' from 'saturated' has not been measured, and a guessed threshold whose documented consequence is DO-NOT-PLOT would delete benchmarks from charts on the strength of a number nobody chose.

revisited at T+90 (≈2026-11-13); until then a threshold is set only if it derives from a measurement.

Headroom is what is left between the best score in the corpus and the benchmark's ceiling. The last column is the most demanding threshold that still produced a staircase past the gate — read them together: a benchmark can sit at 2% headroom and still carry valid curves.
benchmark best score ceiling headroom highest valid threshold state
aider-polyglot 88 100 12.0% no label published
ale-bench 2176.88 1244.92 no label published
arc-agi 0.98 1 2.0% 0.905 no label published
arc-agi-2 0.925 1 7.5% 0.7208 no label published
critpt 0.323 1 67.7% no label published
cursorbench 0.729 1 27.1% 0.648 no label published
deepswe 0.7364864864864865 1 26.4% 0.5812917594654788 no label published
exploitbench 0.738 1 26.2% no label published
osworld-2 0.206 1 79.4% no label published
proofbench 0.78 1 22.0% 0.54 no label published
surface-evolver-bench 0.95 1 5.0% no label published
the-agent-company 0.518 1 48.2% 0.08 no label published
weirdml 0.9194 1 8.1% 0.721 no label published
  • aider-polyglot — not measured: no threshold produced a staircase past the gate
  • ale-bench — unbounded above (a rating, not a fraction); not defined: headroom needs a ceiling, and this benchmark has none
  • critpt — not measured: no threshold produced a staircase past the gate
  • exploitbench — not measured: no threshold produced a staircase past the gate
  • osworld-2 — not measured: no threshold produced a staircase past the gate
  • surface-evolver-bench — not measured: no threshold produced a staircase past the gate

data quality: rows this archive does not stand behind

S31 (19 Aug 2026): price_fact is NOT corrected, filtered or deleted — the archive records what the catalogues actually published. Doubt is this separate, derived claim, and it is presented per INCIDENT (seller × upstream commit), never by blaming the seller whose arrival made an old broken row visible again.

31 of 6702 price rows (0.46%) carry a suspect mark. They are published rather than deleted, and every consumer of the dumps is expected to filter them before computing anything.

Identity, checked on every export so it can actually fail: union == M + R + C − MR − MC − RC + MRC, with each rule set built in its own pass over the ledger so the equation can actually fail — 31 == 15 + 9 + 10 − 3 − 0 − 0 + 0, holds.

The three rules, each built in its own pass over the ledger.
rule definition rows
m magnitude two-tail 1 ≤ price_micro ≤ 999 (a nonzero price under $0.001/Mtok) or price_micro ≥ 1e9 (at or above $1000/Mtok)
measured: 15 of 15 independently confirmed wrong, and the nearest legitimate row sits 2.8× above the top of the bad band.
15
r residue ratio the parser recorded what rounding threw away: residue ≠ 0 and (price_micro = 0 or 1e6·|residue| > price_micro), in exact rational arithmetic
a WITNESS, not a detector: its recall over the classified upstream errors is 3 of 24, because an over-scale by a power of ten on a round decimal lands on a whole µUSD and leaves no residue at all. That blindness is why the third rule is a hand-curated file.
9
c curated incident an entry in suspects.toml, each carrying a written rationale and its evidence, each required to select exactly one price_fact row
a stale entry marks nothing while the count still claims it did, so a curated entry that selects zero or several rows fails the export outright.
10

Overlaps: 3 rows fire on both magnitude and residue, 0 on magnitude and curation, 0 on residue and curation, 0 on all three.

100% precision is a MEASURED PROPERTY of this corpus, not a guarantee. A future error with a mid-sized residue would land in the empty band between the benign artefacts and the real ones, and would want a re-measurement rather than a silent pass.

the curated incidents

An incident is a seller and one upstream commit instant — never the seller whose arrival made an old broken row visible again. A months-old wrong row prints a fresh headline every time a new, blameless seller lists the same model, so attributing the fault to the entering seller names the wrong party.

helicone-2025-10-23-mistral-off-band

helicone · upstream commit 2025-10-23T20:43:27Z · still open at the archive's as-of, 297.5 days and counting · 4 rows

model_uid meter published why it is marked
mistral/mistral-nemo input_text $20/Mtok $20/Mtok against a five-seller consensus of $0.15 — 133x, not a clean power of ten, in a catalogue whose other 94 comparable rows sit at median ratio 1.00. No correction observed through 2026-08-17, the archive's as-of date.
docs/DECISIONS.md 2026-08-19, S31 measurement round; drives NETSPLIT 25cb43927a99.
mistral/mistral-nemo output_text $40/Mtok $40/Mtok against a consensus of $0.15 — 267x, log10 closeness 0.426, so not a decimal shift. No correction observed through 2026-08-17, the archive's as-of date, and the row was still driving NETSPLIT events in the archive's final fortnight.
docs/DECISIONS.md 2026-08-19, S31 measurement round; drives NETSPLIT fc06d4725f67, 3cfc8d37f133, 330f15318621.
mistral/mistral-small input_text $75/Mtok $75/Mtok against a consensus of $1.00 — 75x, log10 closeness 0.125. No correction observed through 2026-08-17, the archive's as-of date. Note the consensus side of this model carries its own identity smear (11 floating aliases across two namespaces), which is why the entry names the row and not a corrected price.
docs/DECISIONS.md 2026-08-19, S31 measurement round; drives NETSPLIT 5a80a91c2525, afca45144ce9.
mistral/mistral-small output_text $200/Mtok $200/Mtok against a consensus of $3.00 — 67x, log10 closeness 0.176. No correction observed through 2026-08-17, the archive's as-of date.
docs/DECISIONS.md 2026-08-19, S31 measurement round; drives NETSPLIT b41cfc097a65, 2594c7fbf79f.

oci-2025-08-07-decimal-shift

oracle-oci · upstream commit 2025-08-07T16:45:17Z · still open at the archive's as-of, 374.6 days and counting · 2 rows

model_uid meter published why it is marked
xai/grok-3 output_text $0.15/Mtok $0.15/Mtok against a five-seller consensus of $15.00 — exactly 10^-2, while the same listing's input_text meter is correct at ratio 1.000. A two-cell decimal error in litellm provider 'oci'; no correction observed through 2026-08-17, the archive's as-of date.
docs/DECISIONS.md 2026-08-19, S31 measurement round; oci book-wide 32 comparable era-rows, median ratio 1.00, only these two off-band.
xai/grok-4-0709 output_text $0.15/Mtok $0.15/Mtok against a five-seller consensus of $15.00 — exactly 10^-2, same commit and same two-cell error as xai/grok-3; the listing's input_text meter is correct.
docs/DECISIONS.md 2026-08-19, S31 measurement round; still open at the 2026-08-17 archive as-of (374+ days), driver of NETSPLIT events 1b99d94cf7a2, 96c32ec44388, 4daa48899afb, cc4bad402b98.

watsonx-2025-10-06-bad-import

ibm-watsonx · upstream commit 2025-10-06T19:33:05Z · corrected upstream after 11 days, on 2025-10-17 · 4 rows

model_uid meter published why it is marked
mistral/mistral-small-2503 input_text $200/Mtok $200/Mtok, 200x a three-seller consensus; corrected upstream 11 days later to $0.10/Mtok (÷2000). Part of one litellm commit that wrote 8 wrong rows over 4 models.
docs/DECISIONS.md 2026-08-19, S31 measurement round; window closes 2025-10-17T18:51:22Z, the incident's own correction commit.
mistral/mistral-small-2503 output_text $600/Mtok $600/Mtok, 200x consensus; corrected upstream 11 days later to $0.30/Mtok (÷2000). Same commit as the input_text row.
docs/DECISIONS.md 2026-08-19, S31 measurement round; window closes 2025-10-17T18:51:22Z.
mistral/pixtral-12b-2409 input_text $150/Mtok $150/Mtok, exactly 1000x a four-seller consensus; corrected upstream 11 days later to $0.35/Mtok. The full Class-1 signature — single seller, clean power of ten, later corrected.
docs/DECISIONS.md 2026-08-19, S31 measurement round; drives NETSPLIT 8ed462c8dfd6.
mistral/pixtral-12b-2409 output_text $150/Mtok $150/Mtok, exactly 1000x consensus; corrected upstream 11 days later to $0.35/Mtok. Same commit as the input_text row.
docs/DECISIONS.md 2026-08-19, S31 measurement round; drives NETSPLIT e1f9455d63b8.
coverage of the check itself
source refs
7349
cited observations
2896
rows citing nothing
0
dangling source refs
0
matched pairs
7325
micro fallback pairs
0
ambiguous micro pairs
0
rows without comparable entry
16
rows with residue
71
rows residue cleared
62
rows with several firing readings
0
rows with unreadable residue
0

A row without a comparable reading is reported as unchecked, not as clean. That is the safe direction and it is counted here rather than hidden.

licence and attribution

The derived dataset is published under CC-BY-4.0. Attribute as: SCROLLBACK (mirc.tech), CC-BY-4.0.

CC-BY-4.0 covers this derived dataset. Upstream attribution is a condition: the measurements in benchmark_result derive entirely from Epoch AI's CC-BY-4.0 benchmark corpus (the descriptive benchmark and benchmark_version rows are this archive's own curated registry, not Epoch data); the price layer derives from the LiteLLM (MIT) and models.dev (MIT) catalogues, and NO price_fact row cites a vendor page; official vendor DEPRECATION pages back 44 availability rows and the 44 event rows derived from them. See DATA-LICENSE.

  • Epoch AI, Benchmarking Hub (CC-BY-4.0) — the sole source of every measurement in benchmark_result, all 2343 rows of it. The descriptive benchmark and benchmark_version tables are this repository's own curated registry, not Epoch data; attributing them upstream would be over-attribution.
  • LiteLLM model catalogue (MIT), excluding the repository's enterprise/ directory, which this project does not read.
  • models.dev catalogue (MIT).
  • Official vendor deprecation pages (developers.openai.com, platform.claude.com) back 44 availability rows and the 44 event rows derived from them — not the layers, which are 1554 and 3588 rows. Only dates and retirement kinds are reproduced; no page is republished, and no price row cites a vendor page.
  • The WeirdML CSV carries no licence statement upstream. It is archived locally and is never redistributed; it produces no derived row. The quality axis comes from Epoch's CC-BY-4.0 corpus, including the WeirdML measurements Epoch mirrors under that licence.
  • The raw archive is not part of the dataset. Archived source bodies are stored for measurement and reproducibility and are not redistributed; the evidence handles in the dumps name archive rows, not resolvable joins.
  • The dumps republish upstream's own citation strings. benchmark_result.evidence_url is Epoch's citation, verbatim, on 1245 of 2343 rows. Of those, 336 point at arcprize.org and 318 point at htihle.github.io; event.source_url carries 32 of arcprize.org and 48 of htihle.github.io. None of those URLs was fetched to build this dataset: the strings travel exactly as Epoch published them, and they are citations rather than data collected from those pages.

The full text travels with the dumps as DATA-LICENSE, served beside them — CC-BY-4.0 §3(a) asks that the notice reach whoever takes the data, and a file that is only in the repository does not. Source code is separate and is MIT.