Benchfolio

Leaderboard / Model

GPT-5.6 Luna

OpenAI · Closed weights · Released Jul 9, 2026 · $0.20 in · $1.20 out per 1M tokens

41Overall · rank 25

Score by job

Coding

52 · rank 21 · 3 of 5 benchmarks

BenchmarkResultVariantz
DeepSWE67.2%gpt-5.6-luna_max0.64
FrontierCode39.8%gpt-5.6-luna_unknown0.03
LMArena Coding1463 ratinggpt-5.6-luna-xhigh-1.32

Agents

Not enough data · 0 of 4 benchmarks

Reasoning

29 · rank 30 · 3 of 4 benchmarks

BenchmarkResultVariantz
GPQA Diamond91.6%gpt-5.6-luna_max0.02
ARC-AGI-259.5%gpt-5.6-luna_max-0.20
SimpleBench46.8%gpt-5.6-luna_unknown-1.86

Math

58 · rank 15 · 5 of 5 benchmarks

BenchmarkResultVariantz
FrontierMath Tiers 1–382.1%gpt-5.6-luna_max0.97
FrontierMath Tier 461.0%gpt-5.6-luna_max0.84
OTIS Mock AIME98.3%gpt-5.6-luna_max0.68
ProofBench60.0%gpt-5.6-luna_max0.48
LMArena Math1457 ratinggpt-5.6-luna-xhigh-0.86

Writing

35 · rank 39 · 2 of 2 benchmarks

BenchmarkResultVariantz
LMArena Creative Writing1396 ratinggpt-5.6-luna-xhigh-1.65
LMArena Instruction Following1435 ratinggpt-5.6-luna-xhigh-1.09

Design

33 · rank 28 · 1 of 1 benchmarks

BenchmarkResultVariantz
LMArena WebDev1519 ratinggpt-5.6-luna-xhigh (codex-harness)-0.38