New Frontier data and RL environments, off the shelf

Benchmarks

Greatness isn't accidental. How we measure it shouldn't be either. If we want AGI that builds billion-dollar enterprises and globe-spanning infrastructure, we need benchmarks that test for intelligence and sophistication.

This is our ranking of models, measured by their capacity for rigorous reasoning and real-world mastery.
View by :
Professional Spreadsheet Reasoning

GDP.xlsx

A spreadsheet might hold a company’s financial model, a laboratory’s results, or an engineer’s calculations. GDP.xlsx puts models to work inside these files, testing whether they can follow the evidence across a workbook and produce answers a professional could rely on.

Rank
Model
Score
Gemini 4 Argon (High)
38.3
%
Gemini 4 Argon (High)
Claude Opus 5.5 (Adaptive/Max)
30.3
%
Claude Opus 5.5 (Adaptive/Max)
Claude Sonnet 5.5
29.1
%
Claude Sonnet 5.5
Claude Fable 5.1 (Adaptive/Max)
23.1
%
Claude Fable 5.1 (Adaptive/Max)
GPT 6 Astra (Max)
22.9
%
GPT 6 Astra (Max)
Muse Spark 1.3 (Max)
22.6
%
Muse Spark 1.3 (Max)
Grok 4.7 (xHigh)
22.3
%
Grok 4.7 (xHigh)
GPT 6.1 Sol (Max)
22
%
GPT 6.1 Sol (Max)
GLM 5.3 (Max)
21.4
%
GLM 5.3 (Max)
Qwen 3.8 Max
20.3
%
Qwen 3.8 Max
Hy4 Preview
18.3
%
Hy4 Preview
Gemini 3.8 Flash (High)
18.3
%
Gemini 3.8 Flash (High)
GPT 6 Sol (Max)
17.1
%
GPT 6 Sol (Max)
Claude Sonnet 5 (Adaptive/Max)
16.9
%
Claude Sonnet 5 (Adaptive/Max)
GPT 6 Luna (Max)
14.9
%
GPT 6 Luna (Max)
Kimi K3 (Max)
13.4
%
Kimi K3 (Max)
Gemini 3.1 Pro
10.3
%
Gemini 3.1 Pro
DeepSeek V4 Pro (Max)
6.9
%
DeepSeek V4 Pro (Max)
Nemotron 3 Ultra
4.9
%
Nemotron 3 Ultra
Mistral Large 3
0
%
Mistral Large 3
Agentic Healthcare

DAYJOB: Healthcare

DAYJOB: Healthcare evaluates long-horizon healthcare agents across clinical, operational, payer, pharmacy, and compliance workflows. Rather than telling models exactly what to do, it asks whether agents can figure out what needs to be done, and reason through it until the end.

Rank
Model
Score
Claude Opus 5.5 (Adaptive/Max)
24.7
%
Claude Opus 5.5 (Adaptive/Max)
GPT 6 Astra (Max)
11.6
%
GPT 6 Astra (Max)
Claude Fable 5.1 (Adaptive/Max)
9.6
%
Claude Fable 5.1 (Adaptive/Max)
Gemini 4 Argon (High)
9.6
%
Gemini 4 Argon (High)
Claude Opus 5 (Adaptive/Max)
8.4
%
Claude Opus 5 (Adaptive/Max)
Grok 4.7 (xHigh)
8.4
%
Grok 4.7 (xHigh)
Muse Spark 1.3 (Max)
7.8
%
Muse Spark 1.3 (Max)
Claude Fable 5 (Adaptive/Max)
5.6
%
Claude Fable 5 (Adaptive/Max)
GPT 6 Sol (Max)
3.6
%
GPT 6 Sol (Max)
Grok 4.6 (xHigh)
2.8
%
Grok 4.6 (xHigh)
Muse Spark 1.2 (xHigh)
2
%
Muse Spark 1.2 (xHigh)
Qwen 3.8 Max
2
%
Qwen 3.8 Max
GPT 5.6 Sol (Max)
1.6
%
GPT 5.6 Sol (Max)
GPT 5.6 Terra (xHigh)
1.6
%
GPT 5.6 Terra (xHigh)
Kimi K3 (Max)
1.2
%
Kimi K3 (Max)
Gemini 3.8 Flash (High)
0.8
%
Gemini 3.8 Flash (High)
Hy4 Preview
0.8
%
Hy4 Preview
GLM 5.3 (Max)
0.4
%
GLM 5.3 (Max)
GLM 5.3 Flash (Max)
0.4
%
GLM 5.3 Flash (Max)
GPT 5.6 Luna (Max)
0.4
%
GPT 5.6 Luna (Max)
DeepSeek V4 Pro (Max)
0.4
%
DeepSeek V4 Pro (Max)
Hy3 (High)
0.4
%
Hy3 (High)
Claude Sonnet 5 (Adaptive/Max)
0
%
Claude Sonnet 5 (Adaptive/Max)
Gemini 3.7 Flash (High)
0
%
Gemini 3.7 Flash (High)
Inkling (Max)
0
%
Inkling (Max)
DeepSeek V4 Flash (Max)
0
%
DeepSeek V4 Flash (Max)
Gemini 3.1 Pro
0
%
Gemini 3.1 Pro
Kimi K2.7 Code (Max)
0
%
Kimi K2.7 Code (Max)
Muse Glimmer 30B (xHigh)
0
%
Muse Glimmer 30B (xHigh)
Nemotron 3 Ultra
0
%
Nemotron 3 Ultra
Mistral Large 3
0
%
Mistral Large 3
GPT 6 Luna (Max)
0
%
GPT 6 Luna (Max)
Professional Graphical Reasoning

Chartography

Chartography is our benchmark for professional chart understanding. It tests whether frontier models can read the Kaplan-Meier curves, candlestick charts, contour maps, Sankey diagrams, Bode plots, and other specialized graphics that professionals use to make real decisions every day. It evaluates visual perception, domain-aware interpretation, and multi-step graphical reasoning.

Rank
Model
Score
Gemini 4 Argon (High)
71.6
%
Gemini 4 Argon (High)
GPT 6 Astra (Max)
71
%
GPT 6 Astra (Max)
Claude Opus 5.5 (Adaptive/Max)
66.3
%
Claude Opus 5.5 (Adaptive/Max)
GPT 6 Sol (Max)
53.6
%
GPT 6 Sol (Max)
Claude Fable 5.1 (Adaptive/Max)
46.2
%
Claude Fable 5.1 (Adaptive/Max)
GPT 5.6 Sol (Max)
45
%
GPT 5.6 Sol (Max)
Gemini 3.8 Flash (High)
40.9
%
Gemini 3.8 Flash (High)
Gemini 3.7 Flash (High)
40.4
%
Gemini 3.7 Flash (High)
Claude Fable 5 (Adaptive/Max)
34.8
%
Claude Fable 5 (Adaptive/Max)
GPT 5.6 Terra (Max)
34
%
GPT 5.6 Terra (Max)
Gemini 3.6 Flash
34
%
Gemini 3.6 Flash
Muse Spark 1.2 (xHigh)
32.1
%
Muse Spark 1.2 (xHigh)
GPT 5.6 Luna (Max)
31.4
%
GPT 5.6 Luna (Max)
GPT 5.5 (xHigh)
31
%
GPT 5.5 (xHigh)
GPT 5.4 (xHigh)
29.6
%
GPT 5.4 (xHigh)
Qwen 3.8 Max
29.1
%
Qwen 3.8 Max
GPT 6 Luna (Max)
29.1
%
GPT 6 Luna (Max)
Muse Spark 1.3 (xHigh)
27.6
%
Muse Spark 1.3 (xHigh)
Claude Opus 5 (Adaptive/Max)
27.3
%
Claude Opus 5 (Adaptive/Max)
Kimi K3 (Max)
26.6
%
Kimi K3 (Max)
Gemini 3.1 Pro
26.1
%
Gemini 3.1 Pro
Muse Spark 1.1 (xHigh)
24.4
%
Muse Spark 1.1 (xHigh)
Grok 4.6 (xHigh)
20.5
%
Grok 4.6 (xHigh)
Qwen 3.8 Flash
17.4
%
Qwen 3.8 Flash
Muse Glimmer 30B (xHigh)
17.2
%
Muse Glimmer 30B (xHigh)
Grok 4.5
17
%
Grok 4.5
Claude Sonnet 5 (Adaptive/Max)
16.6
%
Claude Sonnet 5 (Adaptive/Max)
Claude Opus 4.7 (Adaptive/Max)
16.5
%
Claude Opus 4.7 (Adaptive/Max)
GLM 5.3 Flash
16.3
%
GLM 5.3 Flash
Claude Opus 4.8 (Adaptive/Max)
15.9
%
Claude Opus 4.8 (Adaptive/Max)
Qwen 3.5 Plus
15.9
%
Qwen 3.5 Plus
Qwen 3.7 Plus
15.8
%
Qwen 3.7 Plus
Grok 4.7 (xHigh)
14.7
%
Grok 4.7 (xHigh)
Grok 4.3 (High)
12.9
%
Grok 4.3 (High)
Kimi K2.5
12.6
%
Kimi K2.5
Kimi K2.6
12.2
%
Kimi K2.6
DeepSeek V4.1 Flash (Max)
11.6
%
DeepSeek V4.1 Flash (Max)
DeepSeek V4 Flash Vision (experimental) (Max)
11.2
%
DeepSeek V4 Flash Vision (experimental) (Max)
Gemini 3.5 Flash-Lite
10.2
%
Gemini 3.5 Flash-Lite
Inkling
10
%
Inkling
Inkling Small (xHigh)
9.1
%
Inkling Small (xHigh)
Mistral Large 3
9
%
Mistral Large 3
Long-Context Agentic Instruction Following

HANDBOOK.md Agents

Can an agent follow a 100-page company handbook inside an enterprise RL environment?

HANDBOOK.md is a benchmark for long-context agentic instruction following, modeled on how professionals follow corporate policy in their day-to-day work. Each task is a unique RL environment with internal tools and external MCP servers, across five enterprise domains.

Rank
Model
Score
Claude Fable 5.1 (Adaptive/Max)
38.8
%
Claude Fable 5.1 (Adaptive/Max)
Claude Fable 5 (Adaptive/Max)
36.2
%
Claude Fable 5 (Adaptive/Max)
Claude Opus 5 (Adaptive/Max)
32.3
%
Claude Opus 5 (Adaptive/Max)
Grok 4.6 (xHigh)
27.3
%
Grok 4.6 (xHigh)
DeepSeek V4 Pro (Max)
26.9
%
DeepSeek V4 Pro (Max)
GPT 5.6 Sol (Max)
23.5
%
GPT 5.6 Sol (Max)
DeepSeek V4 Flash Vision (experimental) (Max)
23.1
%
DeepSeek V4 Flash Vision (experimental) (Max)
Claude Opus 4.8 (Adaptive/Max)
21.9
%
Claude Opus 4.8 (Adaptive/Max)
GPT 5.5 (xHigh)
21.5
%
GPT 5.5 (xHigh)
Qwen 3.8 Max
16.5
%
Qwen 3.8 Max
Grok 4.5 (High)
15.8
%
Grok 4.5 (High)
Muse Spark 1.3 (xHigh)
15.4
%
Muse Spark 1.3 (xHigh)
GLM 5.3
15
%
GLM 5.3
Muse Spark 1.1 (xHigh)
13.5
%
Muse Spark 1.1 (xHigh)
Muse Spark 1.2 (xHigh)
13.1
%
Muse Spark 1.2 (xHigh)
Kimi K3 (Max)
11.9
%
Kimi K3 (Max)
Gemini 3.5 Flash (High)
11.2
%
Gemini 3.5 Flash (High)
Gemini 3.8 Flash (High)
10.8
%
Gemini 3.8 Flash (High)
Gemini 3.7 Flash (High)
10.8
%
Gemini 3.7 Flash (High)
GLM 5.3 Flash
10.4
%
GLM 5.3 Flash
Claude Sonnet 4.6 (Adaptive/Max)
10.4
%
Claude Sonnet 4.6 (Adaptive/Max)
Gemini 3.1 Pro
10
%
Gemini 3.1 Pro
GLM 5.2 (xHigh)
10
%
GLM 5.2 (xHigh)
DeepSeek V4 Pro (preview) (xHigh)
9.2
%
DeepSeek V4 Pro (preview) (xHigh)
Qwen 3.7 Max
8.5
%
Qwen 3.7 Max
Hy3 (High)
7.7
%
Hy3 (High)
DeepSeek V4 Flash (xHigh)
7.3
%
DeepSeek V4 Flash (xHigh)
DeepSeek V4 Flash (preview) (xHigh)
7.3
%
DeepSeek V4 Flash (preview) (xHigh)
Kimi K2.6
6.9
%
Kimi K2.6
Gemini 3.6 Flash (High)
5
%
Gemini 3.6 Flash (High)
Muse Glimmer 30B (xHigh)
3.5
%
Muse Glimmer 30B (xHigh)
Gemini 3.5 Flash-Lite (High)
3.1
%
Gemini 3.5 Flash-Lite (High)
Inkling (xHigh)
2.3
%
Inkling (xHigh)
Grok 4.3 (High)
1.9
%
Grok 4.3 (High)
Nemotron 3 Ultra
1.5
%
Nemotron 3 Ultra
Inkling Small (xHigh)
1.2
%
Inkling Small (xHigh)
Nemotron Lightning 3.5 30B A3B
0
%
Nemotron Lightning 3.5 30B A3B
Antidote / Everyday

Antidote: Everyday Edition

Some leaderboards reward the answers that convince you in two seconds. Antidote rewards the answer that leaves you better off a month later. It's our real-world AI leaderboard, graded by doctors, lawyers, and engineers who read every word, check every citation, and run every line of code.

Rank
Model
elo score (95% ci)
Claude Fable 5
1099
(
1085
-
1112
)
Gemini 3.1 Pro
1091
(
1081
-
1101
)
Gemini 3.8 Flash
1087
(
1070
-
1104
)
Gemini 3.6 Flash
1083
(
1069
-
1097
)
Gemini 3.5 Flash
1079
(
1068
-
1090
)
Kimi K3
1077
(
1063
-
1091
)
Gemini 3.7 Flash
1075
(
1059
-
1090
)
GLM 5.3
1067
(
1051
-
1083
)
Claude Fable 5.1
1065
(
1048
-
1083
)
Claude Opus 5
1055
(
1041
-
1069
)
Qwen 3.7 Max
1051
(
1040
-
1061
)
Kimi K2.6
1045
(
1036
-
1054
)
Claude Opus 4.7
1042
(
1031
-
1053
)
Claude Opus 4.6
1042
(
1032
-
1052
)
GLM 5.2
1040
(
1027
-
1052
)
Claude Opus 4.8
1030
(
1018
-
1043
)
Muse Spark 1.2
1028
(
1011
-
1046
)
GPT 5.6 Sol
1027
(
1014
-
1041
)
Hy3
1022
(
1007
-
1037
)
DeepSeek V4 Pro
1018
(
1003
-
1033
)
DeepSeek V4 Flash Vision (experimental)
1015
(
999
-
1030
)
Muse Spark 1.1
1014
(
998
-
1030
)
GPT 5.5
1013
(
1000
-
1025
)
Claude Sonnet 4.6
1012
(
1000
-
1024
)
DeepSeek V4 Pro (preview)
1007
(
995
-
1019
)
Kimi K2.5
1006
(
995
-
1018
)
Grok 4.20 Beta
1001
(
989
-
1012
)
Qwen 3.5 Plus
999
(
987
-
1011
)
Hy4 Preview
992
(
974
-
1009
)
Gemini 3.5 Flash-Lite
991
(
977
-
1005
)
Inkling
991
(
977
-
1005
)
Qwen 3.8 Max
985
(
970
-
1000
)
Grok 4.5
973
(
960
-
986
)
DeepSeek V4 Flash (preview)
966
(
954
-
978
)
Grok 4.6
966
(
950
-
981
)
DeepSeek V3.2
954
(
940
-
967
)
Grok 4.3
950
(
939
-
960
)
Mistral Large 3
947
(
938
-
957
)
Ernie 5.1
947
(
936
-
958
)
Qwen 3.8 Flash
944
(
927
-
962
)
Claude Haiku 4.5
932
(
919
-
944
)
Muse Glimmer 30B
931
(
915
-
947
)
GPT 5.4 Mini
929
(
917
-
941
)
Gemma 3 12B
893
(
880
-
905
)
Nemotron Lightning 3.5 30B A3B
888
(
871
-
904
)
Ernie 4.5 300B
863
(
850
-
876
)
Nova 2 Pro
819
(
809
-
830
)
Agentic Finance

DAYJOB: Finance

DAYJOB: Finance measures long-horizon finance agents across corporate finance, banking, credit, investing, and real assets. It's built around the messy reality of knowledge work: realistic requests, complex environments, and end-to-end professional judgment.

Rank
Model
Score
Claude Opus 5.5 (Adaptive/Max)
23.9
%
Claude Opus 5.5 (Adaptive/Max)
GPT 6 Astra (Max)
21.5
%
GPT 6 Astra (Max)
Gemini 4 Argon (High)
20.3
%
Gemini 4 Argon (High)
Claude Fable 5.1 (Adaptive/Max)
19.8
%
Claude Fable 5.1 (Adaptive/Max)
Muse Spark 1.3 (Max)
14.8
%
Muse Spark 1.3 (Max)
Grok 4.7 (xHigh)
14.5
%
Grok 4.7 (xHigh)
Grok 4.6 (xHigh)
12.8
%
Grok 4.6 (xHigh)
Claude Opus 5 (Adaptive/Max)
11.3
%
Claude Opus 5 (Adaptive/Max)
GPT 6 Sol (Max)
9.3
%
GPT 6 Sol (Max)
Claude Fable 5 (Adaptive/Max)
9
%
Claude Fable 5 (Adaptive/Max)
Hy4 Preview
7.3
%
Hy4 Preview
GLM 5.3 (Max)
6
%
GLM 5.3 (Max)
GPT 5.6 Sol (Max)
5.5
%
GPT 5.6 Sol (Max)
Kimi K3 (Max)
3.8
%
Kimi K3 (Max)
Qwen 3.8 Max
3
%
Qwen 3.8 Max
Gemini 3.7 Flash (High)
2.8
%
Gemini 3.7 Flash (High)
Gemini 3.8 Flash (High)
2.8
%
Gemini 3.8 Flash (High)
GLM 5.3 Flash (Max)
2.3
%
GLM 5.3 Flash (Max)
Muse Spark 1.2 (xHigh)
2.3
%
Muse Spark 1.2 (xHigh)
Claude Sonnet 5 (Adaptive/Max)
2.3
%
Claude Sonnet 5 (Adaptive/Max)
GPT 6 Luna (Max)
2.3
%
GPT 6 Luna (Max)
GPT 5.6 Luna (Max)
1.8
%
GPT 5.6 Luna (Max)
GPT 5.6 Terra (xHigh)
1.8
%
GPT 5.6 Terra (xHigh)
DeepSeek V4 Pro (Max)
1
%
DeepSeek V4 Pro (Max)
Hy3 (High)
0.3
%
Hy3 (High)
DeepSeek V4 Flash (Max)
0
%
DeepSeek V4 Flash (Max)
Gemini 3.1 Pro
0
%
Gemini 3.1 Pro
Inkling (Max)
0
%
Inkling (Max)
Kimi K2.7 Code (Max)
0
%
Kimi K2.7 Code (Max)
Mistral Large 3
0
%
Mistral Large 3
Muse Glimmer 30B (xHigh)
0
%
Muse Glimmer 30B (xHigh)
Nemotron 3 Ultra
0
%
Nemotron 3 Ultra
Enterprise Instruction Following

ComplexConstraints

A benchmark for professional instruction following, where constraints depend on each other, fire conditionally, and must be inferred from context.

Rank
Model
Score
GPT 6 Astra (Max)
57.7
%
Muse Spark 1.3 (xHigh)
51.9
%
GPT 6 Sol (Max)
50.8
%
GPT 5.6 Sol (Max)
50.5
%
Claude Opus 5.5 (Adaptive/Max)
49.9
%
DeepSeek V4.1 Flash (Max)
49.7
%
GPT 5.5 (xHigh)
49.5
%
Grok 4.7 (xHigh)
48.9
%
Gemini 3.8 Flash (High)
48.4
%
GPT 5.4 (xHigh)
46
%
Hy4 Preview
46
%
Qwen 3.8 Max
45.5
%
Claude Fable 5.1 (Adaptive/Max)
45.1
%
Gemini 3.1 Pro
43.7
%
Qwen 3.8 Flash
43.3
%
Gemini 3.7 Flash (High)
42.4
%
DeepSeek V4 Pro (Max)
42
%
Gemini 3.6 Flash (High)
40
%
Muse Spark 1.2 (xHigh)
39.9
%
DeepSeek V4 Flash Vision (experimental) (Max)
39.9
%
Claude Fable 5 (Max)
38.1
%
Muse Spark 1.1 (xHigh)
38.1
%
Kimi K3
37.9
%
Claude Opus 5 (Adaptive/Max)
37.3
%
Gemini 3.5 Flash
37.1
%
Grok 4.6 (xHigh)
36.5
%
Claude Opus 4.6 (Adaptive/High)
36.3
%
Grok 4.5
35.9
%
Claude Opus 4.8 (Max)
35.6
%
Hy3 (High)
34.3
%
Inkling
34.1
%
Claude Sonnet 4.6 (Adaptive/High)
34
%
Qwen 3.7 Max
33.5
%
Claude Opus 4.7
31.6
%
GLM 5.3
30.5
%
Claude Opus 4.7 (Adaptive/High)
30.3
%
Kimi K2.6
30.1
%
GLM 5.2 (Max)
29.3
%
DeepSeek V4 Pro (preview)
28
%
GLM 5.3 Flash
27.6
%
GPT 6 Luna
27.3
%
Muse Glimmer 30B (xHigh)
23.9
%
Gemini 3.5 Flash-Lite (High)
23.2
%
DeepSeek V4 Flash (preview)
20.9
%
Kimi K2.5
18.2
%
Qwen 3.5 Plus
17.8
%
Nemotron 3 Ultra
17.8
%
Claude Opus 4.6
16.9
%
Grok 4.20 Beta
16.4
%
Ernie 5.1
14.7
%
Claude Sonnet 4.6
10.7
%
Nemotron Lightning 3.5 30B A3B
7.1
%
DeepSeek V3.2
2.2
%
Mistral Large 3
0.4
%
Ernie 4.5 300B
0
%
Nova 2 Pro
0
%
Professional Multimodal Reasoning

GDP.pdf

Can frontier models master the documents that run the world? GDP.pdf is a multimodal and reasoning benchmark that takes real-world prompts and PDFs pulled directly from expert professional workflows.

Rank
Model
Score
GPT 6 Astra (Max)
34.2
%
GPT 6 Astra (Max)
GPT 5.6 Sol (Max)
30.7
%
GPT 5.6 Sol (Max)
Claude Opus 5.5 (Adaptive/Max)
30.6
%
Claude Opus 5.5 (Adaptive/Max)
Claude Fable 5 (Adaptive/Max)
29.8
%
Claude Fable 5 (Adaptive/Max)
Muse Spark 1.3 (xHigh)
27.6
%
Muse Spark 1.3 (xHigh)
Claude Fable 5.1 (Adaptive/Max)
27.6
%
Claude Fable 5.1 (Adaptive/Max)
GPT 6 Sol (Max)
26.4
%
GPT 6 Sol (Max)
GPT 5.5 (xHigh)
26
%
GPT 5.5 (xHigh)
GPT 5.6 Terra
24.7
%
GPT 5.6 Terra
Claude Opus 5 (Adaptive/Max)
24
%
Claude Opus 5 (Adaptive/Max)
Claude Opus 4.8 (Adaptive/Max)
24
%
Claude Opus 4.8 (Adaptive/Max)
Gemini 3.7 Flash (High)
23.8
%
Gemini 3.7 Flash (High)
Gemini 3.8 Flash (High)
23.2
%
Gemini 3.8 Flash (High)
Qwen 3.8 Max
23.2
%
Qwen 3.8 Max
GPT 6 Luna (Max)
23
%
GPT 6 Luna (Max)
Grok 4.7 (xHigh)
22.8
%
Grok 4.7 (xHigh)
GPT 5.6 Luna
22.7
%
GPT 5.6 Luna
Claude Opus 4.7 (Adaptive/Max)
21
%
Claude Opus 4.7 (Adaptive/Max)
DeepSeek V4.1 Flash (Max)
19.8
%
DeepSeek V4.1 Flash (Max)
Kimi K3 (Max)
19
%
Kimi K3 (Max)
Claude Sonnet 4.6 (Adaptive/Max)
18
%
Claude Sonnet 4.6 (Adaptive/Max)
Grok 4.6 (xHigh)
17.2
%
Grok 4.6 (xHigh)
Gemini 3.1 Pro
17
%
Gemini 3.1 Pro
Qwen 3.8 Flash
16.6
%
Qwen 3.8 Flash
DeepSeek V4 Flash Vision (experimental) (Max)
16.2
%
DeepSeek V4 Flash Vision (experimental) (Max)
Muse Spark 1.2 (xHigh)
16
%
Muse Spark 1.2 (xHigh)
Muse Spark 1.1
15
%
Muse Spark 1.1
GLM 5.3 Flash
14
%
GLM 5.3 Flash
Grok 4.5 (High)
14
%
Grok 4.5 (High)
Gemini 3.6 Flash (High)
14
%
Gemini 3.6 Flash (High)
Gemini 3.5 Flash
14
%
Gemini 3.5 Flash
Kimi K2.6
12
%
Kimi K2.6
Muse Glimmer 30B (xHigh)
11.8
%
Muse Glimmer 30B (xHigh)
Gemini 3.5 Flash-Lite (High)
10
%
Gemini 3.5 Flash-Lite (High)
Gemini 3 Flash
10
%
Gemini 3 Flash
Grok 4.3 (High)
8
%
Grok 4.3 (High)
Nova 2 Pro
2
%
Nova 2 Pro
Nemotron 3 Nano Omni
2
%
Nemotron 3 Nano Omni
Mistral Large 3
2
%
Mistral Large 3
Enterprise Agents in Realistic RL Environments

EnterpriseBench: CoreCraft Agents

Stop testing models in tiny, self-contained environments. We built CoreCraft, a large-scale startup world, and deployed AI agents to solve real tasks. Our goal: to move agents beyond the cleanliness of the lab and into the chaos of enterprise reality.

Rank
Model
Score
Claude Fable 5.1 (Adaptive/Max)
77.4
%
Claude Fable 5.1 (Adaptive/Max)
Claude Opus 5.5 (Adaptive/Max)
74.4
%
Claude Opus 5.5 (Adaptive/Max)
Claude Fable 5 (Adaptive/Max)
70.3
%
Claude Fable 5 (Adaptive/Max)
Frontier Research Mathematics

Riemann-bench

We evaluate AI models on advanced mathematical problems requiring deep reasoning and novel synthesis. Our benchmark features cutting-edge problems sourced from leading mathematicians, including Ivy League professors, PhD IMO medalists, and graduate students at the top of their field.

Rank
Model
Score
GPT-5.6 Sol (Max)
74.4
%
GPT-5.6 Sol (Max)
GPT 6 Astra (Max)
72
%
GPT 6 Astra (Max)
Claude Opus 5.5 (Adaptive/Max)
69.6
%
Claude Opus 5.5 (Adaptive/Max)
Claude Opus 5 (Adaptive/Max)
68
%
Claude Opus 5 (Adaptive/Max)
GPT 5.6 Terra (Max)
66.4
%
GPT 5.6 Terra (Max)
Claude Fable 5.1 (Adaptive/Max)
65.6
%
Claude Fable 5.1 (Adaptive/Max)
GPT 6 Sol (Max)
64.8
%
GPT 6 Sol (Max)
Claude Fable 5 (Adaptive/Max)
60
%
Claude Fable 5 (Adaptive/Max)
GPT 6 Luna (Max)
59.2
%
GPT 6 Luna (Max)
GPT 5.5 (xHigh)
55.2
%
GPT 5.5 (xHigh)
Gemini 3.8 Flash (High)
51.2
%
Gemini 3.8 Flash (High)
Claude Opus 4.8 (Adaptive/Max)
47.2
%
Claude Opus 4.8 (Adaptive/Max)
GPT 5.4 (xHigh)
41.6
%
GPT 5.4 (xHigh)
Gemini 3.7 Flash (High)
39.2
%
Gemini 3.7 Flash (High)
Grok 4.5 (High)
38.4
%
Grok 4.5 (High)
DeepSeek V4 Pro (Max)
38.4
%
DeepSeek V4 Pro (Max)
Kimi K3
37.6
%
Kimi K3
GPT-5.2 (xHigh)
37.6
%
GPT-5.2 (xHigh)
Gemini 3.5 Flash (High)
36.8
%
Gemini 3.5 Flash (High)
Gemini 3.1 Pro
33.6
%
Gemini 3.1 Pro
Claude Opus 4.7 (Adaptive/Max)
32.8
%
Claude Opus 4.7 (Adaptive/Max)
Gemini 3.6 Flash (High)
30.4
%
Gemini 3.6 Flash (High)
Muse Spark 1.3 (xHigh)
28
%
Muse Spark 1.3 (xHigh)
Claude Opus 4.6 (Adaptive/Max)
27.2
%
Claude Opus 4.6 (Adaptive/Max)
Hy4 Preview
24
%
Hy4 Preview
DeepSeek V4.1 Flash (Max)
24
%
DeepSeek V4.1 Flash (Max)
Muse Spark 1.2 (xHigh)
23.2
%
Muse Spark 1.2 (xHigh)
DeepSeek V4 Flash Vision (experimental) (High)
23.2
%
DeepSeek V4 Flash Vision (experimental) (High)
Muse Spark 1.1 (xHigh)
20.8
%
Muse Spark 1.1 (xHigh)
Inkling (xHigh)
15.2
%
Inkling (xHigh)
Qwen 3.8 Max
15.2
%
Qwen 3.8 Max
Qwen 3.7 Max
15.2
%
Qwen 3.7 Max
Muse Glimmer 30B (xHigh)
13.6
%
Muse Glimmer 30B (xHigh)
Kimi K2.5
12
%
Kimi K2.5
GLM 5.3 (Max)
12
%
GLM 5.3 (Max)
Gemini 3.5 Flash-Lite (High)
11.2
%
Gemini 3.5 Flash-Lite (High)
Claude Opus 4.5 (Adaptive/Max)
11.2
%
Claude Opus 4.5 (Adaptive/Max)
GLM 5.2 (Max)
10.4
%
GLM 5.2 (Max)
DeepSeek V4 Flash (Max)
10.4
%
DeepSeek V4 Flash (Max)
GLM 5.3 Flash (Max)
8.8
%
GLM 5.3 Flash (Max)
DeepSeek V3.2 (Thinking)
8
%
DeepSeek V3.2 (Thinking)
Kimi K2.6
7.2
%
Kimi K2.6
DeepSeek V4 Pro (preview) (xHigh)
5.6
%
DeepSeek V4 Pro (preview) (xHigh)
MAI Thinking 1
4.8
%
MAI Thinking 1
Creative, Business, and Everyday Writing

Hemingway-bench

Most AI writing benchmarks reward surface-level signals: elaborate metaphors and prose that looks impressive at a glance. Hemingway-bench rewards writing that is actually good.

Our leaderboard is judged by professional writers who evaluate creative writing, business writing, and everyday writing tasks for taste, originality, coherence, and emotional intelligence.

Rank
Model
elo score (95% ci)
Claude Fable 5
1110
(
1091
-
1130
)
Gemini 3.7 Flash
1105
(
1084
-
1126
)
Claude Fable 5.1
1095
(
1071
-
1118
)
Gemini 3.8 Flash
1093
(
1069
-
1117
)
Gemini 3.6 Flash
1075
(
1055
-
1095
)
Gemini 3.1 Pro
1073
(
1057
-
1088
)
Gemini 3.5 Flash
1071
(
1054
-
1088
)
Kimi K3
1069
(
1049
-
1089
)
GPT 5.6 Sol
1057
(
1037
-
1076
)
Gemini 3 Pro
1056
(
1031
-
1081
)
GLM 5.3
1053
(
1030
-
1076
)
Gemini 3 Flash
1052
(
1037
-
1066
)
Claude Opus 5
1047
(
1026
-
1068
)
DeepSeek V4 Pro
1046
(
1025
-
1068
)
Claude Opus 4.6
1045
(
1029
-
1061
)
Claude Opus 4.8 (Adaptive/Default)
1044
(
1026
-
1061
)
Claude Opus 4.7
1042
(
1026
-
1058
)
Claude Opus 4.8
1038
(
1021
-
1055
)
GLM 5.2
1036
(
1019
-
1053
)
GPT 5.5
1035
(
1019
-
1052
)
GLM 5.3 Flash
1034
(
1010
-
1057
)
Kimi K2.6
1021
(
1004
-
1038
)
Claude Opus 4.5
1016
(
995
-
1037
)
Hy4 Preview
1014
(
991
-
1037
)
Muse Spark 1.2
1013
(
992
-
1034
)
Qwen 3.8 Max
1009
(
988
-
1030
)
Claude Sonnet 4.6
1001
(
985
-
1016
)
DeepSeek V4 Pro (preview)
999
(
982
-
1015
)
GPT 5.2 Chat
994
(
978
-
1010
)
Gemini 3.5 Flash-Lite
990
(
970
-
1010
)
Qwen 3.5 Plus
990
(
974
-
1005
)
Kimi K2.5
988
(
973
-
1003
)
Muse Spark 1.1
985
(
966
-
1005
)
GPT 5.4
985
(
969
-
1001
)
Grok 4.6
983
(
962
-
1004
)
DeepSeek V4 Flash Vision (experimental)
981
(
958
-
1003
)
Grok 4.5
972
(
954
-
991
)
Hy3
971
(
949
-
994
)
DeepSeek V4 Flash (preview)
969
(
953
-
986
)
GPT 5.2
954
(
929
-
980
)
Qwen 3 Max
948
(
920
-
975
)
Qwen 3.8 Flash
929
(
905
-
952
)
Grok 4.1 Fast Reasoning
916
(
901
-
931
)
Muse Glimmer 30B
913
(
892
-
935
)
Nemotron Lightning 3.5 30B A3B
888
(
866
-
911
)
Kimi K2 Instruct
883
(
856
-
910
)
Llama 4 Maverick
803
(
781
-
824
)
Nova 2 Pro
749
(
718
-
779
)

Stay Posted