Benchfolio

Leaderboard / Model

GPT-5.4

OpenAI · Closed weights · Released Mar 5, 2026 · $2.50 in · $15 out per 1M tokens

51Overall · rank 18

Score by job

Coding

66 · rank 12 · 4 of 5 benchmarks

BenchmarkResultVariantz
DeepSWE51.8%gpt-5.4-2026-03-05_xhigh-0.34
Terminal-Bench81.8%gpt-5.4-2026-03-05_unknown1.04
SWE-bench Verified76.9%gpt-5.4-2026-03-05_high0.15
LMArena Coding1496 ratinggpt-5.4-high0.24

Agents

43 · rank 9 · 2 of 4 benchmarks

BenchmarkResultVariantz
APEX-Agents52.4%gpt-5.4-2026-03-05_unknown0.01
METR Time Horizons5.7 hgpt-5.4-2026-03-05_unknown-0.15

Reasoning

59 · rank 14 · 3 of 4 benchmarks

BenchmarkResultVariantz
GPQA Diamond93.3%gpt-5.4-2026-03-05_xhigh0.70
ARC-AGI-274.0%gpt-5.4-2026-03-05_xhigh0.46
Humanity’s Last Exam36.2%gpt-5.4-2026-03-05_xhigh-0.19

Math

59 · rank 13 · 5 of 5 benchmarks

BenchmarkResultVariantz
FrontierMath Tiers 1–378.6%gpt-5.4-2026-03-05_xhigh0.73
FrontierMath Tier 449.0%gpt-5.4-2026-03-05_xhigh0.34
OTIS Mock AIME97.8%gpt-5.4-2026-03-05_high0.54
ProofBench56.0%gpt-5.4-2026-03-05_xhigh0.33
LMArena Math1490 ratinggpt-5.4-high0.37

Writing

61 · rank 24 · 2 of 2 benchmarks

BenchmarkResultVariantz
LMArena Creative Writing1440 ratinggpt-5.4-high-0.32
LMArena Instruction Following1470 ratinggpt-5.4-high0.21

Design

20 · rank 38 · 1 of 1 benchmarks

BenchmarkResultVariantz
LMArena WebDev1463 ratinggpt-5.4-high (codex-harness)-0.94