Benchfolio

Leaderboard / Model

GPT-5.1

OpenAI · Closed weights · Released Nov 13, 2025 · $1.25 in · $10 out per 1M tokens

13Overall · rank 34

Score by job

Coding

0 · rank 29 · 3 of 5 benchmarks

BenchmarkResultVariantz
Terminal-Bench47.6%gpt-5.1-2025-11-13_medium-1.89
SWE-bench Verified68.0%gpt-5.1-2025-11-13_high-2.52
LMArena Coding1453 ratinggpt-5.1-high-1.80

Agents

Not enough data · 0 of 4 benchmarks

Reasoning

0 · rank 40 · 4 of 4 benchmarks

BenchmarkResultVariantz
GPQA Diamond87.6%gpt-5.1-2025-11-13_high-1.59
ARC-AGI-217.6%gpt-5.1-2025-11-13_high-2.13
SimpleBench53.2%gpt-5.1-2025-11-13_high-1.22
Humanity’s Last Exam23.7%gpt-5.1-2025-11-13_unknown-1.64

Math

Not enough data · 2 of 5 benchmarks

BenchmarkResultVariantz
OTIS Mock AIME88.6%gpt-5.1-2025-11-13_high-1.71
LMArena Math1444 ratinggpt-5.1-high-1.35

Writing

47 · rank 33 · 2 of 2 benchmarks

BenchmarkResultVariantz
LMArena Creative Writing1427 ratinggpt-5.1-high-0.72
LMArena Instruction Following1441 ratinggpt-5.1-high-0.87

Design

4 · rank 45 · 1 of 1 benchmarks

BenchmarkResultVariantz
LMArena WebDev1392 ratinggpt-5.1-medium-1.65