Leaderboard / Model
GPT-5.1
OpenAI · Closed weights · Released Nov 13, 2025 · $1.25 in · $10 out per 1M tokens
13Overall · rank 34
Score by job
Coding
0 · rank 29 · 3 of 5 benchmarks
| Benchmark | Result | Variant | z |
|---|---|---|---|
| Terminal-Bench | 47.6% | gpt-5.1-2025-11-13_medium | -1.89 |
| SWE-bench Verified | 68.0% | gpt-5.1-2025-11-13_high | -2.52 |
| LMArena Coding | 1453 rating | gpt-5.1-high | -1.80 |
Agents
Not enough data · 0 of 4 benchmarks
Reasoning
0 · rank 40 · 4 of 4 benchmarks
| Benchmark | Result | Variant | z |
|---|---|---|---|
| GPQA Diamond | 87.6% | gpt-5.1-2025-11-13_high | -1.59 |
| ARC-AGI-2 | 17.6% | gpt-5.1-2025-11-13_high | -2.13 |
| SimpleBench | 53.2% | gpt-5.1-2025-11-13_high | -1.22 |
| Humanity’s Last Exam | 23.7% | gpt-5.1-2025-11-13_unknown | -1.64 |
Math
Not enough data · 2 of 5 benchmarks
| Benchmark | Result | Variant | z |
|---|---|---|---|
| OTIS Mock AIME | 88.6% | gpt-5.1-2025-11-13_high | -1.71 |
| LMArena Math | 1444 rating | gpt-5.1-high | -1.35 |
Writing
47 · rank 33 · 2 of 2 benchmarks
| Benchmark | Result | Variant | z |
|---|---|---|---|
| LMArena Creative Writing | 1427 rating | gpt-5.1-high | -0.72 |
| LMArena Instruction Following | 1441 rating | gpt-5.1-high | -0.87 |
Design
4 · rank 45 · 1 of 1 benchmarks
| Benchmark | Result | Variant | z |
|---|---|---|---|
| LMArena WebDev | 1392 rating | gpt-5.1-medium | -1.65 |