Leaderboard / Model
DeepSeek V4 Pro 0813
DeepSeek · Open weights · Released Aug 13, 2026 · Price not listed
45Overall · rank 22
Score by job
Coding
Not enough data · 2 of 5 benchmarks
| Benchmark | Result | Variant | z |
|---|---|---|---|
| FrontierCode | 28.5% | deepseek-v4-pro-0813_high | -1.14 |
| LMArena Coding | 1467 rating | deepseek-v4-pro-high-20260813 | -1.13 |
Agents
Not enough data · 1 of 4 benchmarks
| Benchmark | Result | Variant | z |
|---|---|---|---|
| APEX-Agents | 47.3% | deepseek-v4-pro-0813_unknown | -0.48 |
Reasoning
48 · rank 22 · 2 of 4 benchmarks
| Benchmark | Result | Variant | z |
|---|---|---|---|
| GPQA Diamond | 91.7% | deepseek-v4-pro-0813_max | 0.04 |
| ARC-AGI-2 | 61.3% | deepseek-v4-pro-0813_max | -0.13 |
Math
37 · rank 23 · 5 of 5 benchmarks
| Benchmark | Result | Variant | z |
|---|---|---|---|
| FrontierMath Tiers 1–3 | 64.6% | deepseek-v4-pro-0813_max | -0.24 |
| FrontierMath Tier 4 | 26.8% | deepseek-v4-pro-0813_max | -0.59 |
| OTIS Mock AIME | 98.6% | deepseek-v4-pro-0813_max | 0.75 |
| ProofBench | 50.0% | deepseek-v4-pro-0813_unknown | 0.09 |
| LMArena Math | 1446 rating | deepseek-v4-pro-high-20260813 | -1.28 |
Writing
49 · rank 32 · 2 of 2 benchmarks
| Benchmark | Result | Variant | z |
|---|---|---|---|
| LMArena Creative Writing | 1435 rating | deepseek-v4-pro-high-20260813 | -0.49 |
| LMArena Instruction Following | 1443 rating | deepseek-v4-pro-high-20260813 | -0.82 |
Design
47 · rank 18 · 1 of 1 benchmarks
| Benchmark | Result | Variant | z |
|---|---|---|---|
| LMArena WebDev | 1581 rating | deepseek-v4-pro-high-20260813 | 0.22 |