Leaderboard / Model
GPT-5.5
OpenAI · Closed weights · Released Apr 23, 2026 · $5 in · $30 out per 1M tokens
60Overall · rank 10
Score by job
Coding
79 · rank 6 · 5 of 5 benchmarks
| Benchmark | Result | Variant | z |
|---|---|---|---|
| DeepSWE | 67.0% | gpt-5.5_xhigh | 0.63 |
| Terminal-Bench | 84.7% | gpt-5.5_unknown | 1.29 |
| FrontierCode | 43.0% | gpt-5.5_unknown | 0.35 |
| SWE-bench Verified | 80.6% | gpt-5.5-pre-release_xhigh | 1.27 |
| LMArena Coding | 1495 rating | gpt-5.5-high | 0.22 |
Agents
38 · rank 11 · 3 of 4 benchmarks
| Benchmark | Result | Variant | z |
|---|---|---|---|
| APEX-Agents | 55.1% | gpt-5.5_unknown | 0.26 |
| OSWorld 2.0 | 13.0% | gpt-5.5_xhigh | -0.54 |
| Remote Labor Index | 6.3% | gpt-5.5_unknown | -0.35 |
Reasoning
72 · rank 10 · 3 of 4 benchmarks
| Benchmark | Result | Variant | z |
|---|---|---|---|
| GPQA Diamond | 94.0% | gpt-5.5-pre-release_xhigh | 0.98 |
| ARC-AGI-2 | 85.0% | gpt-5.5_xhigh | 0.97 |
| SimpleBench | 69.0% | gpt-5.5_unknown | 0.36 |
Math
69 · rank 7 · 5 of 5 benchmarks
| Benchmark | Result | Variant | z |
|---|---|---|---|
| FrontierMath Tiers 1–3 | 85.3% | gpt-5.5_xhigh | 1.19 |
| FrontierMath Tier 4 | 72.5% | gpt-5.5_xhigh | 1.32 |
| OTIS Mock AIME | 100.0% | gpt-5.5-pre-release_xhigh | 1.09 |
| ProofBench | 50.0% | gpt-5.5_xhigh | 0.09 |
| LMArena Math | 1486 rating | gpt-5.5 | 0.23 |
Writing
68 · rank 12 · 2 of 2 benchmarks
| Benchmark | Result | Variant | z |
|---|---|---|---|
| LMArena Creative Writing | 1454 rating | gpt-5.5-high | 0.08 |
| LMArena Instruction Following | 1478 rating | gpt-5.5-high | 0.51 |
Design
31 · rank 32 · 1 of 1 benchmarks
| Benchmark | Result | Variant | z |
|---|---|---|---|
| LMArena WebDev | 1510 rating | gpt-5.5-xhigh (codex-harness) | -0.47 |