Table of contents
Case Study Llama
Analysis of Individual writings
Appendix

DeepSeek V4 Pro reaches 59.7 on the Tuesday Work Index

The DeepSeek V4 family has made a substantial jump since its preview releases:

DeepSeek V4 Flash preview: 47.5
DeepSeek V4 Pro preview: 49.1
DeepSeek V4 Pro: 59.7

The released V4 Pro gains 10.6 points over the Pro preview and 12.2 points over the Flash preview.

At 59.7, DeepSeek V4 Pro now sits among the stronger frontier operating points on the Tuesday Work Index.

It scores above Qwen 3.8 Max at 58.7, Gemini 3.7 Flash High at 58.2, and Kimi K3 Max at 56.5. The leaders remain meaningfully ahead, with Fable 5 at 66.8 and GPT 5.6 Sol Max at 66.7.

The notable story is not that DeepSeek has reached the top of the leaderboard. It is how much of the gap it has closed since the V4 previews.

DeepSeek V4 Pro across the Surge benchmarks

DeepSeek V4 Pro performs competitively across a broad range of professional and reasoning benchmarks:

Tuesday Work Index: 59.7
ComplexConstraints: 42.1%
HANDBOOK.md: 26.5%
Riemann-bench: 38.4%
EnterpriseBench: 53.3%
Hemingway-bench: 1047 Elo
Antidote: 1027 Elo

The results put DeepSeek close to frontier models across several very different capabilities, from enterprise agents and instruction following to research mathematics and writing.

But the most interesting part of the scorecard is what happens when cost is included.

DeepSeek gets 83% of the top ComplexConstraints score for under 10% of the cost

DeepSeek V4 Pro scores 42.1% on ComplexConstraints, our benchmark for professional instruction following where requirements interact, trigger conditionally, and must often be inferred from context.

The leading operating point scores 50.5%.

DeepSeek therefore reaches roughly 83% of the top score.

A full DeepSeek V4 Pro benchmark run costs approximately $34. The leading operating point costs hundreds of dollars.

That means DeepSeek reaches 83% of the leading score for under 10% of the benchmark run cost.

It sits just below the cost-performance Pareto frontier, rather than directly on it, as GPT 5.5 offers a slightly better combination of score and cost.

The broader DeepSeek family also illustrates how capability and cost have scaled together. DeepSeek V3.2 anchors the extreme low-cost end of the curve, while V4 Pro has moved dramatically upward in performance without moving into the cost range of the most expensive frontier models.

Same Riemann-bench score, 11x the cost

DeepSeek V4 Pro's strongest cost-performance result comes on Riemann-bench, our benchmark for frontier research mathematics.

Two models score exactly 38.4%:

DeepSeek V4 Pro: 38.4% for $11.04
Grok 4.5: 38.4% for $122.38

The measured capability is identical. The benchmark run cost differs by roughly 11x.

That puts DeepSeek V4 Pro directly on the Riemann-bench cost-performance Pareto frontier. Among the operating points we evaluated, there is no model that both scores higher and costs less.

The surrounding results make the operating point even more notable.

Gemini 3.7 Flash High scores 39.2%, just 0.8 percentage points higher, at roughly $18 per benchmark run.

GPT 5.6 Sol Max reaches 74.4%, demonstrating that the absolute capability frontier remains much higher. But its benchmark run costs roughly $172.

DeepSeek therefore occupies a very different point on the curve. It delivers substantial research-mathematics capability at a fraction of the cost of many frontier operating points.

A strong open-weight contender

DeepSeek V4 Pro does not lead every benchmark, and the highest-performing frontier models remain meaningfully ahead on several measures.

But absolute capability is only one axis.

DeepSeek combines a 59.7 Tuesday Work Index score with unusually strong cost-performance on demanding reasoning tasks.

On ComplexConstraints, it reaches 83% of the leading score for under 10% of the cost.

On Riemann-bench, it matches Grok 4.5's score at roughly one-eleventh the benchmark run cost and lands directly on the Pareto frontier.

Overall, DeepSeek V4 Pro is one of the strongest open-weight contenders we've tested, particularly for workloads where complex professional reasoning and cost both matter.

Follow us
/surge-ai
@hellosurgeai

Read what frontier labs read.

We publish 1-2 deep posts every month on Al evaluation, post-training, and pushing the frontier.

Subscription confirmed

You'll get updates when we post
Oops! Something went wrong while submitting the form.

More Posts

Appendix