Leaderboard / Model
GPT-5.2
OpenAI · Closed weights · Released Dec 11, 2025 · $1.75 in · $14 out per 1M tokens
24Overall · rank 32
Score by job
Coding
27 · rank 28 · 3 of 5 benchmarks
| Benchmark | Result | Variant | z |
|---|---|---|---|
| Terminal-Bench | 64.9% | gpt-5.2-2025-12-11_unknown | -0.40 |
| SWE-bench Verified | 73.8% | gpt-5.2-2025-12-11_high | -0.78 |
| LMArena Coding | 1446 rating | gpt-5.2-chat-latest-20260210 | -2.12 |
Agents
28 · rank 12 · 2 of 4 benchmarks
| Benchmark | Result | Variant | z |
|---|---|---|---|
| Remote Labor Index | 2.5% | gpt-5.2-2025-12-11_medium | -0.92 |
| METR Time Horizons | 5.9 h | gpt-5.2-2025-12-11_high | -0.06 |
Reasoning
22 · rank 33 · 4 of 4 benchmarks
| Benchmark | Result | Variant | z |
|---|---|---|---|
| GPQA Diamond | 91.4% | gpt-5.2-2025-12-11_xhigh | -0.07 |
| ARC-AGI-2 | 52.9% | gpt-5.2-2025-12-11_xhigh | -0.51 |
| SimpleBench | 45.8% | gpt-5.2-2025-12-11_high | -1.96 |
| Humanity’s Last Exam | 27.8% | gpt-5.2-2025-12-11_unknown | -1.16 |
Math
26 · rank 31 · 5 of 5 benchmarks
| Benchmark | Result | Variant | z |
|---|---|---|---|
| FrontierMath Tiers 1–3 | 67.4% | gpt-5.2-2025-12-11_xhigh | -0.04 |
| FrontierMath Tier 4 | 31.7% | gpt-5.2-2025-12-11_xhigh | -0.39 |
| OTIS Mock AIME | 96.1% | gpt-5.2-2025-12-11_high | 0.13 |
| ProofBench | 15.0% | gpt-5.2-2025-12-11_xhigh | -1.29 |
| LMArena Math | 1441 rating | gpt-5.2-high | -1.46 |
Writing
30 · rank 40 · 2 of 2 benchmarks
| Benchmark | Result | Variant | z |
|---|---|---|---|
| LMArena Creative Writing | 1402 rating | gpt-5.2-chat-latest-20260210 | -1.48 |
| LMArena Instruction Following | 1417 rating | gpt-5.2-chat-latest-20260210 | -1.77 |
Design
10 · rank 42 · 1 of 1 benchmarks
| Benchmark | Result | Variant | z |
|---|---|---|---|
| LMArena WebDev | 1417 rating | gpt-5.2 | -1.40 |