GPT-5.6 Luna (max)
openai/gpt-5.6-luna
| job | benchmark | # | score | task $ | ratio $ | per | released |
|---|---|---|---|---|---|---|---|
| Code | ALE-Bench | 41 | 1667.4 | $0.359171 | $0.0002154 | rating point | 2026-07-09 |
| Reason | ProofBench | 3 | 0.54 | $0.449325 | $0.8321 | correct answer | 2026-07-09 |
| Agent | DeepSWE | 2 | 0.442 | $0.777901 | $1.758 | solved task | 2026-07-09 |
| Agent | DeepSWE | 4 | 0.113 | $0.21631 | $1.917 | solved task | 2026-07-09 |
| Agent | DeepSWE | 7 | 0.569 | $1.53562 | $2.701 | solved task | 2026-07-09 |
| Agent | DeepSWE | 13 | 0.672 | $3.02812 | $4.507 | solved task | 2026-07-09 |
| Agent | DeepSWE | 14 | 0.0155 | $0.0724062 | $4.675 | solved task | 2026-07-09 |
These are the numbers on #frontier. How a fixed capability's cost falls over time is on /methodology.