Leaderboard / Model
GPT-5.6 Luna
OpenAI · Closed weights · Released Jul 9, 2026 · $0.20 in · $1.20 out per 1M tokens
41Overall · rank 25
Score by job
Coding
52 · rank 21 · 3 of 5 benchmarks
| Benchmark | Result | Variant | z |
|---|---|---|---|
| DeepSWE | 67.2% | gpt-5.6-luna_max | 0.64 |
| FrontierCode | 39.8% | gpt-5.6-luna_unknown | 0.03 |
| LMArena Coding | 1463 rating | gpt-5.6-luna-xhigh | -1.32 |
Agents
Not enough data · 0 of 4 benchmarks
Reasoning
29 · rank 30 · 3 of 4 benchmarks
| Benchmark | Result | Variant | z |
|---|---|---|---|
| GPQA Diamond | 91.6% | gpt-5.6-luna_max | 0.02 |
| ARC-AGI-2 | 59.5% | gpt-5.6-luna_max | -0.20 |
| SimpleBench | 46.8% | gpt-5.6-luna_unknown | -1.86 |
Math
58 · rank 15 · 5 of 5 benchmarks
| Benchmark | Result | Variant | z |
|---|---|---|---|
| FrontierMath Tiers 1–3 | 82.1% | gpt-5.6-luna_max | 0.97 |
| FrontierMath Tier 4 | 61.0% | gpt-5.6-luna_max | 0.84 |
| OTIS Mock AIME | 98.3% | gpt-5.6-luna_max | 0.68 |
| ProofBench | 60.0% | gpt-5.6-luna_max | 0.48 |
| LMArena Math | 1457 rating | gpt-5.6-luna-xhigh | -0.86 |
Writing
35 · rank 39 · 2 of 2 benchmarks
| Benchmark | Result | Variant | z |
|---|---|---|---|
| LMArena Creative Writing | 1396 rating | gpt-5.6-luna-xhigh | -1.65 |
| LMArena Instruction Following | 1435 rating | gpt-5.6-luna-xhigh | -1.09 |
Design
33 · rank 28 · 1 of 1 benchmarks
| Benchmark | Result | Variant | z |
|---|---|---|---|
| LMArena WebDev | 1519 rating | gpt-5.6-luna-xhigh (codex-harness) | -0.38 |