New Frontier data and RL environments, off the shelf

Benchmarks

Greatness isn't accidental. How we measure it shouldn't be either. If we want AGI that builds billion-dollar enterprises and globe-spanning infrastructure, we need benchmarks that test for intelligence and sophistication.

This is our ranking of models, measured by their capacity for rigorous reasoning and real-world mastery.
View by :
Professional Graphical Reasoning

Chartography

Chartography is our benchmark for professional chart understanding. It tests whether frontier models can read the Kaplan-Meier curves, candlestick charts, contour maps, Sankey diagrams, Bode plots, and other specialized graphics that professionals use to make real decisions every day. It evaluates visual perception, domain-aware interpretation, and multi-step graphical reasoning.

Rank
Model
Score
GPT 5.6 Sol (Max)
45
%
GPT 5.6 Sol (Max)
Gemini 3.7 Flash
43
%
Gemini 3.7 Flash
Gemini 3.7 Flash (High)
40.4
%
Gemini 3.7 Flash (High)
GPT 5.6 Sol
39.5
%
GPT 5.6 Sol
Gemini 3.5 Flash
35.9
%
Gemini 3.5 Flash
Claude Fable 5 (Adaptive/Max)
34.8
%
Claude Fable 5 (Adaptive/Max)
GPT 5.6 Terra (Max)
34
%
GPT 5.6 Terra (Max)
Gemini 3.6 Flash
34
%
Gemini 3.6 Flash
Muse Spark 1.2 (xHigh)
32.1
%
Muse Spark 1.2 (xHigh)
Muse Spark 1.2
31.5
%
Muse Spark 1.2
GPT 5.6 Luna (Max)
31.4
%
GPT 5.6 Luna (Max)
GPT 5.5 (xHigh)
31
%
GPT 5.5 (xHigh)
GPT 5.4 (xHigh)
29.6
%
GPT 5.4 (xHigh)
Claude Fable 5
29.5
%
Claude Fable 5
Qwen 3.8 Max
29.1
%
Qwen 3.8 Max
GPT 5.5
28.5
%
GPT 5.5
Claude Opus 5 (Adaptive/Max)
27.3
%
Claude Opus 5 (Adaptive/Max)
GPT 5.6 Terra
26.6
%
GPT 5.6 Terra
Kimi K3 (Max)
26.6
%
Kimi K3 (Max)
Gemini 3.1 Pro
26.1
%
Gemini 3.1 Pro
Claude Opus 5
25.7
%
Claude Opus 5
Muse Spark 1.1 (xHigh)
24.4
%
Muse Spark 1.1 (xHigh)
Muse Spark 1.1
23.6
%
Muse Spark 1.1
Grok 4.6
21.9
%
Grok 4.6
GPT 5.6 Luna
21.4
%
GPT 5.6 Luna
Grok 4.6 (xHigh)
20.5
%
Grok 4.6 (xHigh)
Muse Glimmer 30B
17.9
%
Muse Glimmer 30B
Muse Glimmer 30B (xHigh)
17.2
%
Muse Glimmer 30B (xHigh)
Grok 4.5
17
%
Grok 4.5
Claude Sonnet 5 (Adaptive/Max)
16.6
%
Claude Sonnet 5 (Adaptive/Max)
Claude Opus 4.7 (Adaptive/Max)
16.5
%
Claude Opus 4.7 (Adaptive/Max)
Claude Opus 4.8 (Adaptive/Max)
15.9
%
Claude Opus 4.8 (Adaptive/Max)
Qwen 3.5 Plus
15.9
%
Qwen 3.5 Plus
Qwen 3.7 Plus
15.8
%
Qwen 3.7 Plus
GPT 5.4
14.1
%
GPT 5.4
Claude Opus 4.7
13.6
%
Claude Opus 4.7
Grok 4.3 (High)
12.9
%
Grok 4.3 (High)
Kimi K2.5
12.6
%
Kimi K2.5
Claude Sonnet 5
12.2
%
Claude Sonnet 5
Kimi K2.6
12.2
%
Kimi K2.6
Grok 4.3
11.7
%
Grok 4.3
Claude Opus 4.8
11.3
%
Claude Opus 4.8
Gemini 3.5 Flash-Lite
10.2
%
Gemini 3.5 Flash-Lite
Inkling (xHigh)
10
%
Inkling (xHigh)
Mistral Large 3
9
%
Mistral Large 3
Long-Context Agentic Instruction Following

HANDBOOK.md Agents

Can an agent follow a 100-page company handbook inside an enterprise RL environment?

HANDBOOK.md is a benchmark for long-context agentic instruction following, modeled on how professionals follow corporate policy in their day-to-day work. Each task is a unique RL environment with internal tools and external MCP servers, across five enterprise domains.

Rank
Model
Score
Claude Fable 5 (Adaptive/Max)
36.2
%
Claude Fable 5 (Adaptive/Max)
Claude Fable 5
34.2
%
Claude Fable 5
Claude Opus 5 (Adaptive/Max)
32.3
%
Claude Opus 5 (Adaptive/Max)
Claude Opus 5
29.6
%
Claude Opus 5
Grok 4.6 (xHigh)
27.3
%
Grok 4.6 (xHigh)
DeepSeek V4 Pro (Max)
26.5
%
DeepSeek V4 Pro (Max)
Grok 4.6
25.8
%
Grok 4.6
GPT 5.6 Sol (Max)
23.5
%
GPT 5.6 Sol (Max)
Claude Opus 4.8 (Adaptive/Max)
21.9
%
Claude Opus 4.8 (Adaptive/Max)
GPT 5.6 Sol
21.5
%
GPT 5.6 Sol
GPT 5.5 (xHigh)
21.5
%
GPT 5.5 (xHigh)
GPT 5.5
21.5
%
GPT 5.5
DeepSeek V4 Pro
19.6
%
DeepSeek V4 Pro
Claude Opus 4.8
18.9
%
Claude Opus 4.8
Qwen 3.8 Max
16.5
%
Qwen 3.8 Max
Grok 4.5 (High)
15.8
%
Grok 4.5 (High)
Muse Spark 1.1 (xHigh)
13.5
%
Muse Spark 1.1 (xHigh)
Muse Spark 1.2 (xHigh)
13.1
%
Muse Spark 1.2 (xHigh)
GLM 5.2
12.7
%
GLM 5.2
Kimi K3 (Max)
11.9
%
Kimi K3 (Max)
Gemini 3.7 Flash
11.9
%
Gemini 3.7 Flash
Muse Spark 1.2
11.2
%
Muse Spark 1.2
Gemini 3.5 Flash (High)
11.2
%
Gemini 3.5 Flash (High)
Gemini 3.7 Flash (High)
10.8
%
Gemini 3.7 Flash (High)
Claude Sonnet 4.6 (Adaptive/Max)
10.4
%
Claude Sonnet 4.6 (Adaptive/Max)
Gemini 3.1 Pro
10
%
Gemini 3.1 Pro
GLM 5.2 (xHigh)
10
%
GLM 5.2 (xHigh)
Gemini 3.5 Flash
9.2
%
Gemini 3.5 Flash
DeepSeek V4 Pro (preview) (xHigh)
9.2
%
DeepSeek V4 Pro (preview) (xHigh)
Qwen 3.7 Max
8.5
%
Qwen 3.7 Max
Claude Sonnet 4.6
7.7
%
Claude Sonnet 4.6
DeepSeek V4 Flash (xHigh)
7.3
%
DeepSeek V4 Flash (xHigh)
DeepSeek V4 Flash (preview) (xHigh)
7.3
%
DeepSeek V4 Flash (preview) (xHigh)
DeepSeek V4 Flash (preview)
7.3
%
DeepSeek V4 Flash (preview)
DeepSeek V4 Flash
7.3
%
DeepSeek V4 Flash
Kimi K2.6
6.9
%
Kimi K2.6
DeepSeek V4 Pro (preview)
6.9
%
DeepSeek V4 Pro (preview)
Gemini 3.6 Flash (High)
5
%
Gemini 3.6 Flash (High)
Muse Glimmer 30B (xHigh)
3.5
%
Muse Glimmer 30B (xHigh)
Muse Glimmer 30B
3.5
%
Muse Glimmer 30B
Gemini 3.5 Flash-Lite (High)
3.1
%
Gemini 3.5 Flash-Lite (High)
Inkling (xHigh)
2.3
%
Inkling (xHigh)
Grok 4.3 (High)
1.9
%
Grok 4.3 (High)
Nemotron 3 Ultra
1.5
%
Nemotron 3 Ultra
Grok 4.3
0.8
%
Grok 4.3
Antidote / Everyday

Antidote: Everyday Edition

Some leaderboards reward the answers that convince you in two seconds. Antidote rewards the answer that leaves you better off a month later. It's our real-world AI leaderboard, graded by doctors, lawyers, and engineers who read every word, check every citation, and run every line of code.

Rank
Model
elo score (95% ci)
Claude Fable 5
1110
(
1094
-
1125
)
Gemini 3.6 Flash
1087
(
1071
-
1104
)
Gemini 3.1 Pro
1086
(
1076
-
1096
)
Gemini 3.5 Flash
1083
(
1070
-
1095
)
Kimi K3
1074
(
1058
-
1089
)
Claude Opus 5
1052
(
1035
-
1069
)
Qwen 3.7 Max
1051
(
1040
-
1063
)
GLM 5.2
1045
(
1031
-
1059
)
Claude Opus 4.7
1043
(
1031
-
1055
)
Kimi K2.6
1042
(
1033
-
1052
)
Claude Opus 4.6
1038
(
1027
-
1048
)
Claude Opus 4.8
1035
(
1022
-
1049
)
Muse Spark 1.2
1030
(
1013
-
1046
)
GPT 5.6 Sol
1022
(
1007
-
1038
)
Muse Spark 1.1
1015
(
1000
-
1031
)
Claude Sonnet 4.6
1012
(
1000
-
1025
)
GPT 5.5
1010
(
997
-
1023
)
DeepSeek V4 Pro
1006
(
993
-
1019
)
Kimi K2.5
1006
(
994
-
1017
)
Grok 4.20 Beta
1000
(
989
-
1012
)
Qwen 3.5 Plus
999
(
987
-
1010
)
Gemini 3.5 Flash-Lite
995
(
979
-
1011
)
Qwen 3.8 Max
990
(
972
-
1008
)
Inkling
986
(
970
-
1001
)
Grok 4.5
977
(
963
-
992
)
DeepSeek V4 Flash (preview)
968
(
956
-
980
)
DeepSeek V3.2
954
(
941
-
967
)
Grok 4.3
952
(
941
-
963
)
Mistral Large 3
948
(
937
-
958
)
Ernie 5.1
945
(
933
-
957
)
Claude Haiku 4.5
937
(
924
-
951
)
GPT 5.4 Mini
936
(
923
-
949
)
Gemma 3 12B
896
(
883
-
910
)
Ernie 4.5 300B
866
(
853
-
879
)
Nova 2 Pro
824
(
813
-
835
)
Creative, Business, and Everyday Writing

Hemingway-bench

Most AI writing benchmarks reward surface-level signals: elaborate metaphors and prose that looks impressive at a glance. Hemingway-bench rewards writing that is actually good.

Our leaderboard is judged by professional writers who evaluate creative writing, business writing, and everyday writing tasks for taste, originality, coherence, and emotional intelligence.

Rank
Model
elo score (95% ci)
Claude Fable 5
1118
(
1097
-
1139
)
Gemini 3.6 Flash
1075
(
1053
-
1097
)
Gemini 3.1 Pro
1069
(
1053
-
1085
)
Kimi K3
1066
(
1044
-
1088
)
Gemini 3.5 Flash
1066
(
1048
-
1084
)
GPT 5.6 Sol
1061
(
1040
-
1082
)
Claude Opus 5
1055
(
1032
-
1079
)
Gemini 3 Pro
1054
(
1030
-
1078
)
Gemini 3 Flash
1053
(
1038
-
1068
)
Claude Opus 4.6
1046
(
1029
-
1062
)
GPT 5.5
1038
(
1020
-
1055
)
Claude Opus 4.7
1038
(
1021
-
1056
)
Claude Opus 4.8
1038
(
1019
-
1056
)
GLM 5.2
1030
(
1012
-
1049
)
Kimi K2.6
1026
(
1007
-
1044
)
Claude Opus 4.5
1015
(
995
-
1035
)
Muse Spark 1.2
1014
(
990
-
1037
)
Qwen 3.8 Max
1002
(
979
-
1026
)
DeepSeek V4 Pro
1000
(
982
-
1017
)
Claude Sonnet 4.6
997
(
982
-
1013
)
Kimi K2.5
993
(
977
-
1008
)
GPT 5.2 Chat
991
(
975
-
1007
)
Qwen 3.5 Plus
988
(
972
-
1004
)
Gemini 3.5 Flash-Lite
983
(
960
-
1005
)
Muse Spark 1.1
980
(
958
-
1001
)
GPT 5.4
980
(
964
-
997
)
DeepSeek V4 Flash (preview)
978
(
961
-
996
)
Grok 4.5
973
(
952
-
993
)
GPT 5.2
955
(
930
-
980
)
Qwen 3 Max
948
(
921
-
975
)
Grok 4.1 Fast Reasoning
915
(
900
-
931
)
Kimi K2 Instruct
885
(
859
-
911
)
Llama 4 Maverick
806
(
785
-
827
)
Nova 2 Pro
754
(
724
-
784
)
Enterprise Instruction Following

ComplexConstraints

A benchmark for professional instruction following, where constraints depend on each other, fire conditionally, and must be inferred from context.

Rank
Model
Score
GPT 5.6 Sol (Max)
50.5
%
GPT 5.5 (xHigh)
49.5
%
GPT 5.5 (High)
48.9
%
GPT 5.4 (xHigh)
46
%
Qwen 3.8 Max
45.5
%
GPT 5.5
44.4
%
Gemini 3.1 Pro
43.7
%
GPT 5.6 Sol
43.7
%
Gemini 3.7 Flash (High)
42.4
%
DeepSeek V4 Pro
42.1
%
GPT 5.4 (High)
42.1
%
DeepSeek V4 Pro (Max)
42
%
Gemini 3.6 Flash (High)
40
%
Muse Spark 1.2 (xHigh)
39.9
%
Claude Fable 5 (Max)
38.1
%
Muse Spark 1.1 (xHigh)
38.1
%
Kimi K3
37.9
%
Claude Opus 5 (Adaptive/Max)
37.3
%
Gemini 3.5 Flash
37.1
%
Claude Fable 5 (High)
36.9
%
Muse Spark 1.1 (High)
36.9
%
Gemini 3.7 Flash
36.5
%
Grok 4.6 (xHigh)
36.5
%
Claude Opus 4.6 (Adaptive/High)
36.3
%
Grok 4.5
35.9
%
Claude Opus 5 (Adaptive/High)
35.7
%
Claude Opus 4.8 (Max)
35.6
%
Gemini 3.6 Flash
35.5
%
Grok 4.6
35.1
%
Claude Opus 4.8 (Adaptive/High)
34.8
%
Muse Spark 1.2
34.8
%
Claude Opus 4.8
34.2
%
Inkling
34.1
%
Claude Sonnet 4.6 (Adaptive/High)
34
%
Qwen 3.7 Max
33.5
%
Claude Opus 4.7
31.6
%
Claude Opus 4.7 (Adaptive/High)
30.3
%
Kimi K2.6
30.1
%
GLM 5.2 (Max)
29.3
%
DeepSeek V4 Pro (preview)
28
%
Muse Glimmer 30B (xHigh)
23.9
%
Muse Glimmer 30B
23.7
%
Gemini 3.5 Flash-Lite (High)
23.2
%
GLM 5.2
22.8
%
DeepSeek V4 Flash (preview)
20.9
%
Kimi K2.5
18.2
%
Qwen 3.5 Plus
17.8
%
Nemotron 3 Ultra
17.8
%
Claude Opus 4.6
16.9
%
Grok 4.20 Beta
16.4
%
Ernie 5.1
14.7
%
Claude Sonnet 4.6
10.7
%
DeepSeek V3.2
2.2
%
Gemini 3.5 Flash-Lite
1.3
%
Mistral Large 3
0.4
%
Ernie 4.5 300B
0
%
Nova 2 Pro
0
%
Professional Multimodal Reasoning

GDP.pdf

Can frontier models master the documents that run the world? GDP.pdf is a multimodal and reasoning benchmark that takes real-world prompts and PDFs pulled directly from expert professional workflows.

Rank
Model
Score
GPT 5.6 Sol (Max)
30.7
%
GPT 5.6 Sol (Max)
Claude Fable 5 (Adaptive/Max)
29.8
%
Claude Fable 5 (Adaptive/Max)
GPT 5.5 (xHigh)
26
%
GPT 5.5 (xHigh)
GPT 5.6 Terra
24.7
%
GPT 5.6 Terra
Claude Opus 5 (Adaptive/Max)
24
%
Claude Opus 5 (Adaptive/Max)
Claude Opus 4.8 (Adaptive/Max)
24
%
Claude Opus 4.8 (Adaptive/Max)
GPT 5.6 Luna
22.7
%
GPT 5.6 Luna
Claude Opus 4.7 (Adaptive/Max)
21
%
Claude Opus 4.7 (Adaptive/Max)
Kimi K3 (Max)
19
%
Kimi K3 (Max)
Claude Sonnet 4.6 (Adaptive/Max)
18
%
Claude Sonnet 4.6 (Adaptive/Max)
Gemini 3.1 Pro
17
%
Gemini 3.1 Pro
Muse Spark 1.2 (xHigh)
16
%
Muse Spark 1.2 (xHigh)
Muse Spark 1.1
15
%
Muse Spark 1.1
Grok 4.5 (High)
14
%
Grok 4.5 (High)
Gemini 3.6 Flash (High)
14
%
Gemini 3.6 Flash (High)
Gemini 3.5 Flash
14
%
Gemini 3.5 Flash
Muse Spark 1.2
12
%
Muse Spark 1.2
Kimi K2.6
12
%
Kimi K2.6
Gemini 3.5 Flash-Lite (High)
10
%
Gemini 3.5 Flash-Lite (High)
Gemini 3 Flash
10
%
Gemini 3 Flash
Grok 4.3 (High)
8
%
Grok 4.3 (High)
Nova 2 Pro
2
%
Nova 2 Pro
Nemotron 3 Nano Omni
2
%
Nemotron 3 Nano Omni
Mistral Large 3
2
%
Mistral Large 3
Enterprise Agents in Realistic RL Environments

EnterpriseBench: CoreCraft Agents

Stop testing models in tiny, self-contained environments. We built CoreCraft, a large-scale startup world, and deployed AI agents to solve real tasks. Our goal: to move agents beyond the cleanliness of the lab and into the chaos of enterprise reality.

Rank
Model
Score
Claude Fable 5 (Adaptive/Max)
70.3
%
Claude Fable 5 (Adaptive/Max)
Claude Opus 5 (Adaptive/Max)
68.7
%
Claude Opus 5 (Adaptive/Max)
Grok 4.6 (xHigh)
65.6
%
Grok 4.6 (xHigh)
Frontier Research Mathematics

Riemann-bench

We evaluate AI models on advanced mathematical problems requiring deep reasoning and novel synthesis. Our benchmark features cutting-edge problems sourced from leading mathematicians, including Ivy League professors, PhD IMO medalists, and graduate students at the top of their field.

Rank
Model
Score
GPT-5.6 Sol (Max)
74.4
%
GPT-5.6 Sol (Max)
Claude Opus 5 (Adaptive/Max)
68
%
Claude Opus 5 (Adaptive/Max)
Claude Fable 5 (Adaptive/Max)
60
%
Claude Fable 5 (Adaptive/Max)
GPT 5.5 (xHigh)
55.2
%
GPT 5.5 (xHigh)
Claude Opus 4.8 (Adaptive/Max)
47.2
%
Claude Opus 4.8 (Adaptive/Max)
GPT 5.4 (xHigh)
41.6
%
GPT 5.4 (xHigh)
Gemini 3.7 Flash (High)
39.2
%
Gemini 3.7 Flash (High)
Grok 4.5 (High)
38.4
%
Grok 4.5 (High)
DeepSeek V4 Pro 0813
38.4
%
DeepSeek V4 Pro 0813
Kimi K3
37.6
%
Kimi K3
GPT-5.2 (xHigh)
37.6
%
GPT-5.2 (xHigh)
Gemini 3.5 Flash (High)
36.8
%
Gemini 3.5 Flash (High)
Gemini 3.1 Pro
33.6
%
Gemini 3.1 Pro
Claude Opus 4.7 (Adaptive/Max)
32.8
%
Claude Opus 4.7 (Adaptive/Max)
Gemini 3.6 Flash (High)
30.4
%
Gemini 3.6 Flash (High)
Claude Opus 4.6 (Adaptive/Max)
27.2
%
Claude Opus 4.6 (Adaptive/Max)
Muse Spark 1.2 (xHigh)
23.2
%
Muse Spark 1.2 (xHigh)
Muse Spark 1.1 (xHigh)
20.8
%
Muse Spark 1.1 (xHigh)
Inkling (xHigh)
15.2
%
Inkling (xHigh)
Qwen 3.8 Max
15.2
%
Qwen 3.8 Max
Qwen 3.7 Max
15.2
%
Qwen 3.7 Max
Muse Glimmer 30B
13.6
%
Muse Glimmer 30B
Kimi K2.5
12
%
Kimi K2.5
Gemini 3.5 Flash-Lite (High)
11.2
%
Gemini 3.5 Flash-Lite (High)
Claude Opus 4.5 (Adaptive/Max)
11.2
%
Claude Opus 4.5 (Adaptive/Max)
GLM 5.2
10.4
%
GLM 5.2
DeepSeek V4 Flash
10.4
%
DeepSeek V4 Flash
DeepSeek V3.2 (Thinking)
8
%
DeepSeek V3.2 (Thinking)
Kimi K2.6
7.2
%
Kimi K2.6
DeepSeek V4 Pro
5.6
%
DeepSeek V4 Pro

Stay Posted on New Benchmarks