Qwen 3.8 Max scores 58.7 on the Tuesday Work Index, an 8.6-point improvement over Qwen 3.7 Max and a 22.4-point improvement over Qwen 3.5 Plus.
Its strongest result comes on ComplexConstraints, where Qwen 3.8 Max scores 45.5% and lands on the cost-performance Pareto frontier. It reaches roughly 90% of the leading score for 32% of the benchmark run cost.

Qwen gains 22.4 points across three generations
The clearest measure of Qwen’s progress is the Tuesday Work Index, our composite measure of frontier AI performance at work.
Qwen 3.5 Plus — 36.3
Qwen 3.7 Max — 50.1
Qwen 3.8 Max — 58.7

Qwen 3.7 Max improved by 13.8 points over Qwen 3.5 Plus. Qwen 3.8 Max adds another 8.6 points.
At 58.7, Qwen 3.8 Max now sits in the same cluster as several other frontier operating points: DeepSeek V4 Pro High scores 59.7, Gemini 3.7 Flash High scores 58.2, and Kimi K3 Max scores 56.5.
The leaders remain meaningfully ahead, with Fable 5 Adaptive Max at 66.8 and GPT 5.6 Sol Max at 66.7.
Qwen 3.8 Max lands on the ComplexConstraints cost-performance Pareto frontier
Qwen 3.8 Max’s strongest result is on ComplexConstraints, our benchmark for professional instruction following where constraints depend on one another, trigger conditionally, and often need to be inferred from context.
It scores 45.5%, up substantially across recent Qwen generations:
Qwen 3.5 Plus — 17.8%
Qwen 3.7 Max — 33.5%
Qwen 3.8 Max — 45.5%

Running Qwen 3.8 Max across ComplexConstraints costs $119. GPT 5.6 Sol Max, the highest-scoring operating point, reaches 50.5% at a run cost of $367.
Qwen therefore achieves approximately 90% of the leading score for 32% of the cost.
Graphical reasoning improves sharply
Qwen 3.8 Max also makes a large jump on Chartography, our benchmark for understanding the specialized charts and visualizations professionals use to make real decisions.
The available Qwen progression is:
Qwen 3.5 Plus — 15.9%
Qwen 3.7 Plus — 15.8%
Qwen 3.8 Max — 29.1%

Performance was effectively flat across the earlier two evaluated generations before jumping 13.3 percentage points with Qwen 3.8 Max.
Qwen does not quite reach the Chartography cost-performance Pareto frontier: Gemini 3.6 Flash Medium scores higher at a slightly lower evaluation cost. But the generational improvement is one of the largest in the model’s scorecard.
The gains are concentrated
Qwen 3.8 Max is not better everywhere.
On Riemann-bench, our benchmark for frontier research mathematics, it scores 15.2%, unchanged from Qwen 3.7 Max.
On Antidote, our expert-graded benchmark of real-world answer quality, Qwen 3.8 Max scores 990 Elo, down from 1052 for Qwen 3.7 Max.
The model’s strongest gains instead cluster around structured constraints, professional instruction following, and graphical reasoning.
Bottom line
Qwen 3.8 Max is the strongest Qwen generation we’ve measured on the Tuesday Work Index.
Across three generations, Qwen has moved from 36.3 to 58.7, narrowing a large share of the gap to the leading frontier models.
And on ComplexConstraints, Qwen 3.8 Max establishes a genuinely frontier-level cost-performance operating point.
The frontier is still higher in absolute capability. But Qwen is getting much closer, and on some professional workloads, it is doing so at a fraction of the cost.




