What a task costs — and what a unit of score costs
the index ranks currently purchasable configurations on one named benchmark by task dollars per score point, lowest first. two efforts of the same model are two rows. a composite score is not computed.
Code · ALE-Bench
66 configurations · rating · measured by Epoch AI · ranked by $ / rating point
task cost divided by ALE-Bench rating — dollars per rating point. Task costs: Epoch AI, basis
list_price_x_tokens, community-submitted. Displayed values
are rounded; exact values in site.json and the dumps.
66 ranked, 0 retired excluded, 0 dropped for missing cost, 0 for a published cost of $0, 0 for a non-positive score
| # | ||||
|---|---|---|---|---|
| 1 | gpt-oss-20b | 566.05 | $0.000717 | $0.000001267 |
| 2 | gpt-oss-120b | 575.62 | $0.001303 | $0.000002264 |
| 3 | DeepSeek-V4 Flash (none) | 324.98 | $0.000882 | $0.000002714 |
| 4 | DeepSeek-V4 Flash (high) | 678.2 | $0.005179 | $0.000007636 |
| 5 | GPT-5 nano (high) | 718.67 | $0.007194 | $0.00001001 |
| 6 | DeepSeek-V4 Flash (0731) (max) | 1306.08 | $0.019526 | $0.00001495 |
| 7 | Grok Code Fast 1 | 587.73 | $0.008947 | $0.00001522 |
| 8 | Mistral Large (2512) | 264.7 | $0.005104 | $0.00001928 |
| 9 | Codestral (2508) | 137.78 | $0.00278 | $0.00002018 |
| 10 | Grok 4.1 Fast (reasoning) | 394.93 | $0.009221 | $0.00002335 |
| 11 | DeepSeek-V3.1-Terminus | 745.17 | $0.01883 | $0.00002527 |
| 12 | GPT-5 mini (high) | 799.77 | $0.020554 | $0.0000257 |
| 13 | GPT-5.4 nano (high) | 1004.52 | $0.027084 | $0.00002696 |
| 14 | Mistral Small (2603) | 497.62 | $0.015587 | $0.00003132 |
| 15 | Gemini 3.1 Flash-Lite | 797.73 | $0.025178 | $0.00003156 |
| 16 | DeepSeek-V4 Pro (none) | 521.67 | $0.0166 | $0.00003182 |
| 17 | Mistral Medium 3.1 (2508) | 210.18 | $0.007372 | $0.00003507 |
| 18 | GPT-5.4 (none) | 1086.03 | $0.052314 | $0.00004817 |
| 19 | GPT-4.1 | 558.1 | $0.026917 | $0.00004823 |
| 20 | o4-mini (high) | 826.17 | $0.044719 | $0.00005413 |
| 21 | GPT-5 (minimal) | 807.65 | $0.043744 | $0.00005416 |
| 22 | DeepSeek-R1 (May 2025) | 804.12 | $0.043794 | $0.00005446 |
| 23 | Grok 4.3 | 944.17 | $0.055799 | $0.0000591 |
| 24 | Gemini 3.5 Flash-Lite (high) | 765.27 | $0.045726 | $0.00005975 |
| 25 | Claude Opus 4.6 | 996.5 | $0.061688 | $0.0000619 |
| 26 | GPT-5.5 (none) | 1127.58 | $0.073484 | $0.00006517 |
| 27 | Gemini 3 Flash Preview | 1367.2 | $0.091067 | $0.00006661 |
| 28 | Gemini 2.5 Flash | 661.88 | $0.068155 | $0.000103 |
| 29 | Claude Haiku 4.5 | 653.48 | $0.071831 | $0.0001099 |
| 30 | o3 (high) | 933.55 | $0.103474 | $0.0001108 |
| 31 | GPT-5.2 (medium) | 1249.83 | $0.142947 | $0.0001144 |
| 32 | Grok 4.20 (reasoning) | 1150.28 | $0.13838 | $0.0001203 |
| 33 | Grok 4.5 (high) | 1309.28 | $0.176767 | $0.000135 |
| 34 | DeepSeek-V4 Pro (high) | 1006.08 | $0.148062 | $0.0001472 |
| 35 | Claude Opus 4.7 | 1323.05 | $0.224549 | $0.0001697 |
| 36 | GPT-5.4 (medium) | 1520.72 | $0.26258 | $0.0001727 |
| 37 | Claude Sonnet 4.5 | 796.15 | $0.151558 | $0.0001904 |
| 38 | GPT-5.5 (medium) | 1589.38 | $0.328965 | $0.000207 |
| 39 | GPT-5.1-Codex-Max | 1208.83 | $0.253186 | $0.0002094 |
| 40 | GPT-5 (high) | 1162.45 | $0.244948 | $0.0002107 |
| 41 | GPT-5.6 Luna (max) | 1667.4 | $0.359171 | $0.0002154 |
| 42 | Gemini 3 Pro Preview | 1176.75 | $0.259385 | $0.0002204 |
| 43 | GPT-5.4 mini (high) | 1188.58 | $0.271408 | $0.0002283 |
| 44 | GPT-5.2 (high) | 1293.55 | $0.296865 | $0.0002295 |
| 45 | GPT-5.1 (high) | 1192.15 | $0.315236 | $0.0002644 |
| 46 | Claude Sonnet 4 | 655.35 | $0.185869 | $0.0002836 |
| 47 | Claude Opus 4.5 | 1025.38 | $0.292547 | $0.0002853 |
| 48 | Claude Opus 4.8 (none) | 1411.83 | $0.421746 | $0.0002987 |
| 49 | Gemini 3.1 Pro Preview | 1160.6 | $0.388403 | $0.0003347 |
| 50 | GPT-5.3-Codex (xhigh) | 1655.22 | $0.559265 | $0.0003379 |
| 51 | Gemini 3.6 Flash (high) | 715.52 | $0.244468 | $0.0003417 |
| 52 | Gemini 2.5 Pro | 785.52 | $0.26988 | $0.0003436 |
| 53 | Mistral Medium (2604) | 763.98 | $0.272986 | $0.0003573 |
| 54 | GPT-5.4 (high) | 1607 | $0.612174 | $0.0003809 |
| 55 | GPT-5.1-Codex | 1244.92 | $0.512093 | $0.0004113 |
| 56 | GPT-5.2-Codex | 1299.9 | $0.587753 | $0.0004522 |
| 57 | Gemini 3.5 Flash (high) | 911.02 | $0.441921 | $0.0004851 |
| 58 | Claude Opus 4.8 (high) | 1563.83 | $0.963945 | $0.0006164 |
| 59 | GPT-5.6 Terra (max) | 1951.42 | $1.25405 | $0.0006426 |
| 60 | Claude Opus 4.1 | 674.77 | $0.434347 | $0.0006437 |
| 61 | Claude Sonnet 5 (high) | 1463.12 | $0.96046 | $0.0006564 |
| 62 | GPT-5.6 Sol (max) | 2176.88 | $1.53408 | $0.0007047 |
| 63 | GPT-5.5 (xhigh) | 1942.97 | $1.51713 | $0.0007808 |
| 64 | Claude Opus 5 (high) | 2164.57 | $1.89051 | $0.0008734 |
| 65 | Claude Fable 5 (high) | 2041.3 | $2.10882 | $0.001033 |
| 66 | Claude Sonnet 4.6 (medium) | 1327.3 | $1.77075 | $0.001334 |
Reason · ProofBench
33 configurations · accuracy · measured by Epoch AI · ranked by $ / correct answer
task cost divided by accuracy — expected spend per correct answer at this accuracy. Task costs: Epoch AI, basis
unknown, leaderboard-maintainer-verified. Displayed values
are rounded; exact values in site.json and the dumps.
33 ranked, 0 retired excluded, 0 dropped for missing cost, 1 for a published cost of $0, 0 for a non-positive score
| # | ||||
|---|---|---|---|---|
| 1 | GPT-5.4 nano (high) | 0.05 | $0.015894 | $0.3179 |
| 2 | GPT-5 nano (high) | 0.12 | $0.049218 | $0.4102 |
| 3 | GPT-5.6 Luna (max) | 0.54 | $0.449325 | $0.8321 |
| 4 | Gemini 3.5 Flash-Lite (high) | 0.13 | $0.132788 | $1.021 |
| 5 | GPT-5.4 mini (xhigh) | 0.21 | $0.249441 | $1.188 |
| 6 | Gemini 3.6 Flash (high) | 0.36 | $0.522059 | $1.45 |
| 7 | DeepSeek-V3.2 (none) | 0.08 | $0.11773 | $1.472 |
| 8 | Grok 4.5 (high) | 0.3 | $0.484548 | $1.615 |
| 9 | Gemini 3.1 Pro Preview (high) | 0.26 | $0.525953 | $2.023 |
| 10 | Claude Fable 5 (max) | 0.77 | $1.63546 | $2.124 |
| 11 | Gemini 3 Flash Preview (high) | 0.15 | $0.327286 | $2.182 |
| 12 | Grok 4.3 (high) | 0.11 | $0.244122 | $2.219 |
| 13 | GPT-5.6 Sol (max) | 0.77 | $1.7366 | $2.255 |
| 14 | Grok 4.20 (reasoning) | 0.14 | $0.31815 | $2.272 |
| 15 | Grok 4.1 Fast (reasoning) | 0.04 | $0.096578 | $2.414 |
| 16 | GPT-5 mini (high) | 0.09 | $0.22021 | $2.447 |
| 17 | Claude Opus 5 (max) | 0.78 | $2.10619 | $2.7 |
| 18 | Claude Opus 4.8 (max) | 0.69 | $2.29313 | $3.323 |
| 19 | Claude Opus 4.7 (max) | 0.54 | $1.80779 | $3.348 |
| 20 | GPT-5.5 (xhigh) | 0.5 | $1.68179 | $3.364 |
| 21 | Claude Sonnet 5 (max) | 0.66 | $2.24876 | $3.407 |
| 22 | Claude Sonnet 4.5 (unknown) | 0.19 | $0.65341 | $3.439 |
| 23 | Claude Opus 4.5 (high) | 0.36 | $1.26704 | $3.52 |
| 24 | GPT-5.6 Terra (xhigh) | 0.71 | $2.91861 | $4.111 |
| 25 | GPT-5 (high) | 0.18 | $0.852361 | $4.735 |
| 26 | Claude Sonnet 4.6 (max) | 0.45 | $2.16211 | $4.805 |
| 27 | Gemini 3.5 Flash (high) | 0.29 | $1.39997 | $4.828 |
| 28 | Claude Opus 4.6 (max) | 0.5 | $2.70358 | $5.407 |
| 29 | GPT-5.4 (xhigh) | 0.56 | $3.20201 | $5.718 |
| 30 | GPT-5.1-Codex-Max (high) | 0.09 | $0.527356 | $5.86 |
| 31 | Gemini 3 Pro Preview (high) | 0.2 | $1.22908 | $6.145 |
| 32 | GPT-5.2 (xhigh) | 0.15 | $1.05969 | $7.065 |
| 33 | Mistral Medium (2604) (high) | 0.1 | $1.50143 | $15.01 |
Agent · DeepSWE
45 configurations · pass@1 · measured by Epoch AI · ranked by $ / solved task
task cost divided by pass@1 — expected spend per solved task at this pass rate. Task costs: Epoch AI, basis
unknown, leaderboard-maintainer-verified. Displayed values
are rounded; exact values in site.json and the dumps.
45 ranked, 0 retired excluded, 0 dropped for missing cost, 0 for a published cost of $0, 0 for a non-positive score
| # | ||||
|---|---|---|---|---|
| 1 | GPT-5.6 Terra (medium) | 0.351 | $0.583249 | $1.661 |
| 2 | GPT-5.6 Luna (high) | 0.442 | $0.777901 | $1.758 |
| 3 | GPT-5.6 Terra (low) | 0.241 | $0.427745 | $1.778 |
| 4 | GPT-5.6 Luna (medium) | 0.113 | $0.21631 | $1.917 |
| 5 | GPT-5.6 Terra (high) | 0.538 | $1.13438 | $2.11 |
| 6 | GPT-5.6 Sol (low) | 0.454 | $1.07427 | $2.369 |
| 7 | GPT-5.6 Luna (xhigh) | 0.569 | $1.53562 | $2.701 |
| 8 | Claude Opus 5 (low) | 0.581 | $1.66262 | $2.86 |
| 9 | GPT-5.6 Sol (medium) | 0.611 | $1.86203 | $3.049 |
| 10 | GPT-5.6 Terra (xhigh) | 0.602 | $2.12719 | $3.535 |
| 11 | GPT-5.5 (low) | 0.27 | $1.2002 | $4.447 |
| 12 | Grok 4.5 (high) | 0.538 | $2.41575 | $4.493 |
| 13 | GPT-5.6 Luna (max) | 0.672 | $3.02812 | $4.507 |
| 14 | GPT-5.6 Luna (low) | 0.0155 | $0.0724062 | $4.675 |
| 15 | Claude Opus 5 (medium) | 0.689 | $3.28979 | $4.774 |
| 16 | GPT-5.6 Sol (high) | 0.694 | $3.46983 | $5 |
| 17 | GPT-5.5 (medium) | 0.54 | $2.74934 | $5.093 |
| 18 | Claude Opus 4.8 (low) | 0.408 | $2.29337 | $5.621 |
| 19 | Claude Fable 5 (low) | 0.596 | $3.75787 | $6.307 |
| 20 | GPT-5.6 Sol (xhigh) | 0.707 | $4.70366 | $6.65 |
| 21 | Claude Opus 4.8 (medium) | 0.487 | $3.44386 | $7.076 |
| 22 | GPT-5.6 Terra (max) | 0.696 | $4.94585 | $7.104 |
| 23 | Claude Sonnet 5 (low) | 0.305 | $2.18656 | $7.166 |
| 24 | Gemini 3.6 Flash (high) | 0.486 | $3.53236 | $7.274 |
| 25 | GPT-5.5 (high) | 0.644 | $5.10044 | $7.922 |
| 26 | Claude Opus 4.8 (high) | 0.518 | $4.28194 | $8.271 |
| 27 | Claude Opus 5 (high) | 0.728 | $6.07608 | $8.343 |
| 28 | Claude Fable 5 (medium) | 0.654 | $6.08819 | $9.314 |
| 29 | Claude Sonnet 5 (medium) | 0.398 | $4.07904 | $10.25 |
| 30 | GPT-5.5 (xhigh) | 0.67 | $7.22624 | $10.78 |
| 31 | GPT-5.4 (xhigh) | 0.518 | $5.65246 | $10.92 |
| 32 | GPT-5.6 Sol (max) | 0.727 | $8.38644 | $11.54 |
| 33 | Claude Opus 5 (xhigh) | 0.732 | $9.0722 | $12.4 |
| 34 | Claude Fable 5 (high) | 0.686 | $9.17764 | $13.38 |
| 35 | Claude Opus 4.8 (xhigh) | 0.544 | $8.00636 | $14.73 |
| 36 | Claude Sonnet 5 (high) | 0.482 | $7.42556 | $15.4 |
| 37 | Claude Opus 5 (max) | 0.736 | $11.8376 | $16.07 |
| 38 | Claude Sonnet 4.6 (high) | 0.299 | $5.52235 | $18.45 |
| 39 | Claude Fable 5 (xhigh) | 0.699 | $13.4145 | $19.19 |
| 40 | Gemini 3.5 Flash (medium) | 0.374 | $7.34182 | $19.64 |
| 41 | Claude Opus 4.8 (max) | 0.59 | $13.2226 | $22.42 |
| 42 | Claude Sonnet 5 (xhigh) | 0.497 | $11.8906 | $23.94 |
| 43 | Claude Fable 5 (max) | 0.697 | $21.6347 | $31.03 |
| 44 | Claude Sonnet 5 (max) | 0.538 | $26.3999 | $49.03 |
| 45 | Gemini 3.1 Pro Preview (high) | 0.118 | $9.48152 | $80.68 |
How the ratio is defined: usd_per_point = cost_per_task / score; rank 1 is the lowest ratio among configurations that are not retired at the archive as-of. models with no availability window are unknown, not retired. The half-life band and the threshold scan live on /methodology.