New Frontier data and RL environments, off the shelf

Benchmarks

Greatness isn't accidental. How we measure it shouldn't be either. If we want AGI that builds billion-dollar enterprises and globe-spanning infrastructure, we need benchmarks that test for intelligence and sophistication.

This is our ranking of models, measured by their capacity for rigorous reasoning and real-world mastery.
View by :
Professional Graphical Reasoning

Chartography

Chartography is our benchmark for professional chart understanding. It tests whether frontier models can read the Kaplan-Meier curves, candlestick charts, contour maps, Sankey diagrams, Bode plots, and other specialized graphics that professionals use to make real decisions every day. It evaluates visual perception, domain-aware interpretation, and multi-step graphical reasoning.

Rank
Model
Score
GPT 5.6 Sol (Max)
45
%
GPT 5.6 Sol (Max)
GPT 5.6 Sol
39.5
%
GPT 5.6 Sol
Gemini 3.5 Flash
35.9
%
Gemini 3.5 Flash
Claude Fable 5 (Adaptive/Max)
34.8
%
Claude Fable 5 (Adaptive/Max)
GPT 5.6 Terra (Max)
34
%
GPT 5.6 Terra (Max)
Gemini 3.6 Flash
34
%
Gemini 3.6 Flash
GPT 5.6 Luna (Max)
31.4
%
GPT 5.6 Luna (Max)
GPT 5.5 (xHigh)
31
%
GPT 5.5 (xHigh)
GPT 5.4 (xHigh)
29.6
%
GPT 5.4 (xHigh)
Claude Fable 5
29.5
%
Claude Fable 5
GPT 5.5
28.5
%
GPT 5.5
Claude Opus 5 (Adaptive/Max)
27.3
%
Claude Opus 5 (Adaptive/Max)
GPT 5.6 Terra
26.6
%
GPT 5.6 Terra
Kimi K3
26.6
%
Kimi K3
Gemini 3.1 Pro
26.1
%
Gemini 3.1 Pro
Claude Opus 5
25.7
%
Claude Opus 5
Muse Spark 1.1 (xHigh)
24.4
%
Muse Spark 1.1 (xHigh)
Muse Spark 1.1
23.6
%
Muse Spark 1.1
GPT 5.6 Luna
21.4
%
GPT 5.6 Luna
Grok 4.5
17.3
%
Grok 4.5
Grok 4.5 (High)
16.7
%
Grok 4.5 (High)
Claude Sonnet 5 (Adaptive/Max)
16.6
%
Claude Sonnet 5 (Adaptive/Max)
Claude Opus 4.7 (Adaptive/Max)
16.5
%
Claude Opus 4.7 (Adaptive/Max)
Qwen 3.5 Plus
15.9
%
Qwen 3.5 Plus
Claude Opus 4.8 (Adaptive/Max)
15.9
%
Claude Opus 4.8 (Adaptive/Max)
Qwen 3.7 Plus
15.8
%
Qwen 3.7 Plus
GPT 5.4
14.1
%
GPT 5.4
Claude Opus 4.7
13.6
%
Claude Opus 4.7
Grok 4.3 (High)
12.9
%
Grok 4.3 (High)
Kimi K2.5
12.6
%
Kimi K2.5
Claude Sonnet 5
12.2
%
Claude Sonnet 5
Kimi K2.6
12.2
%
Kimi K2.6
Grok 4.3
11.7
%
Grok 4.3
Claude Opus 4.8
11.3
%
Claude Opus 4.8
Gemini 3.5 Flash Lite
10.2
%
Gemini 3.5 Flash Lite
Mistral Large 3
9
%
Mistral Large 3
Inkling
3.4
%
Inkling
Long-Context Agentic Instruction Following

HANDBOOK.md Agents

Can an agent follow a 100-page company handbook inside an enterprise RL environment?

HANDBOOK.md is a benchmark for long-context agentic instruction following, modeled on how professionals follow corporate policy in their day-to-day work. Each task is a unique RL environment with internal tools and external MCP servers, across five enterprise domains.

Rank
Model
Score
Claude Fable 5 (Adaptive/Max)
36.2
%
Claude Fable 5 (Adaptive/Max)
Claude Fable 5
34.2
%
Claude Fable 5
Claude Opus 5 (Adaptive/Max)
32.3
%
Claude Opus 5 (Adaptive/Max)
Claude Opus 5
29.6
%
Claude Opus 5
GPT 5.6 Sol (Max)
23.5
%
GPT 5.6 Sol (Max)
Claude Opus 4.8 (Adaptive/Max)
21.9
%
Claude Opus 4.8 (Adaptive/Max)
GPT 5.5 (xHigh)
21.5
%
GPT 5.5 (xHigh)
GPT 5.5
21.5
%
GPT 5.5
GPT 5.6 Sol
21.5
%
GPT 5.6 Sol
Claude Opus 4.8
18.9
%
Claude Opus 4.8
Grok 4.5 (High)
15.8
%
Grok 4.5 (High)
Muse Spark 1.1 (xHigh)
13.5
%
Muse Spark 1.1 (xHigh)
GLM 5.2
12.7
%
GLM 5.2
Kimi K3 (Max)
11.9
%
Kimi K3 (Max)
Gemini 3.5 Flash (High)
11.2
%
Gemini 3.5 Flash (High)
Claude Sonnet 4.6 (Adaptive/Max)
10.4
%
Claude Sonnet 4.6 (Adaptive/Max)
GLM 5.2 (xHigh)
10
%
GLM 5.2 (xHigh)
Gemini 3.1 Pro
10
%
Gemini 3.1 Pro
Gemini 3.5 Flash
9.2
%
Gemini 3.5 Flash
DeepSeek V4 Pro (xHigh)
9.2
%
DeepSeek V4 Pro (xHigh)
Qwen 3.7 Max
8.5
%
Qwen 3.7 Max
Hy3 (High)
7.7
%
Hy3 (High)
Claude Sonnet 4.6
7.7
%
Claude Sonnet 4.6
DeepSeek V4 Flash (xHigh)
7.3
%
DeepSeek V4 Flash (xHigh)
DeepSeek V4 Flash
7.3
%
DeepSeek V4 Flash
Kimi K2.6
6.9
%
Kimi K2.6
DeepSeek V4 Pro
6.9
%
DeepSeek V4 Pro
Gemini 3.6 Flash (High)
5
%
Gemini 3.6 Flash (High)
Gemini 3.5 Flash-Lite (High)
3.1
%
Gemini 3.5 Flash-Lite (High)
Inkling (Max)
1.9
%
Inkling (Max)
Grok 4.3 (High)
1.9
%
Grok 4.3 (High)
Nemotron 3 Ultra
1.5
%
Nemotron 3 Ultra
Grok 4.3
0.8
%
Grok 4.3
Antidote / Everyday

Antidote: Everyday Edition

Some leaderboards reward the answers that convince you in two seconds; they optimize for confidence, decoration, and flattery. Antidote rewards the answer that leaves you better off a month later; it optimizes for you.

It's our real-world AI leaderboard, graded by doctors, lawyers, and engineers who read every word, check every citation, and run every line of code.

Today's release covers everyday chatbot use, with agentic and enterprise editions coming soon.

Rank
Model
elo score (95% ci)
Claude Fable 5
1113
(
1098
-
1129
)
Gemini 3.6 Flash
1087
(
1070
-
1104
)
Gemini 3.1 Pro
1086
(
1076
-
1096
)
Gemini 3.5 Flash
1082
(
1069
-
1094
)
Kimi K3
1075
(
1059
-
1091
)
Qwen 3.7 Max
1051
(
1040
-
1063
)
Claude Opus 5
1049
(
1031
-
1067
)
GLM 5.2
1045
(
1031
-
1059
)
Claude Opus 4.7
1041
(
1029
-
1053
)
Kimi K2.6
1040
(
1031
-
1050
)
Claude Opus 4.6
1039
(
1028
-
1049
)
Claude Opus 4.8
1036
(
1022
-
1050
)
GPT 5.6 Sol
1021
(
1005
-
1037
)
Muse Spark 1.1
1017
(
1001
-
1033
)
Claude Sonnet 4.6
1013
(
1000
-
1025
)
GPT 5.5
1009
(
996
-
1022
)
Kimi K2.5
1005
(
994
-
1017
)
DeepSeek V4 Pro
1004
(
991
-
1017
)
Grok 4.20 Beta
1000
(
989
-
1011
)
Qwen 3.5 Plus
998
(
987
-
1010
)
Gemini 3.5 Flash-Lite
997
(
981
-
1014
)
Inkling
987
(
971
-
1003
)
Grok 4.5
977
(
962
-
992
)
DeepSeek V4 Flash
968
(
956
-
980
)
DeepSeek V3.2
954
(
941
-
967
)
Grok 4.3
953
(
942
-
964
)
Mistral Large 3
947
(
937
-
957
)
Ernie 5.1
946
(
934
-
958
)
Claude Haiku 4.5
937
(
924
-
951
)
GPT 5.4 Mini
934
(
921
-
947
)
Gemma 3 12B
897
(
883
-
911
)
Ernie 4.5 300B
866
(
853
-
879
)
Nova 2 Pro
824
(
813
-
835
)
Creative, Business, and Everyday Writing

Hemingway-bench

Most AI writing benchmarks reward surface-level signals: elaborate metaphors and prose that looks impressive at a glance. Hemingway-bench rewards writing that is actually good.

Our leaderboard is judged by professional writers who evaluate creative writing, business writing, and everyday writing tasks for taste, originality, coherence, and emotional intelligence.

Rank
Model
elo score (95% ci)
Claude Fable 5
1118
(
1096
-
1139
)
Gemini 3.6 Flash
1076
(
1053
-
1099
)
Gemini 3.1 Pro
1069
(
1053
-
1085
)
Kimi K3
1065
(
1042
-
1087
)
Gemini 3.5 Flash
1064
(
1046
-
1082
)
GPT 5.6 Sol
1060
(
1039
-
1081
)
Claude Opus 5
1057
(
1033
-
1081
)
Gemini 3 Pro
1053
(
1029
-
1077
)
Gemini 3 Flash
1051
(
1036
-
1066
)
Claude Opus 4.6
1043
(
1027
-
1060
)
Claude Opus 4.7
1039
(
1022
-
1056
)
Claude Opus 4.8 (Adaptive/Default)
1039
(
1020
-
1058
)
Claude Opus 4.8
1038
(
1019
-
1057
)
GPT 5.5
1036
(
1018
-
1053
)
GLM 5.2
1032
(
1013
-
1051
)
Kimi K2.6
1026
(
1008
-
1045
)
Claude Opus 4.5
1014
(
994
-
1034
)
DeepSeek V4 Pro
998
(
981
-
1016
)
Claude Sonnet 4.6
996
(
980
-
1012
)
Kimi K2.5
991
(
976
-
1007
)
Qwen 3.5 Plus
989
(
973
-
1004
)
GPT 5.2 Chat
989
(
973
-
1005
)
Gemini 3.5 Flash-Lite
984
(
961
-
1007
)
GPT 5.4
980
(
964
-
997
)
Muse Spark 1.1
978
(
956
-
1000
)
DeepSeek V4 Flash
976
(
959
-
994
)
Grok 4.5
973
(
952
-
994
)
GPT 5.2
954
(
929
-
979
)
Qwen 3 Max
948
(
921
-
974
)
Grok 4.1 Fast Reasoning
918
(
902
-
933
)
Kimi K2 Instruct
885
(
859
-
911
)
Llama 4 Maverick
806
(
785
-
827
)
Nova 2 Pro
754
(
724
-
784
)
Enterprise Instruction Following

ComplexConstraints

A benchmark for the kind of instruction following professional work demands — where constraints depend on each other, fire conditionally, and must be inferred from context.

Rank
Model
Score
GPT 5.6 Sol (Max)
50.5
%
GPT 5.5 (xHigh)
49.5
%
GPT 5.5 (High)
48.9
%
GPT 5.4 (xHigh)
46
%
GPT 5.5
44.4
%
Gemini 3.1 Pro
43.7
%
GPT 5.6 Sol
43.7
%
GPT 5.4 (High)
42.1
%
Gemini 3.6 Flash (High)
40
%
Claude Fable 5 (Max)
38.1
%
Muse Spark 1.1 (xHigh)
38.1
%
Kimi K3
37.9
%
Claude Opus 5 (Adaptive/Max)
37.3
%
Gemini 3.5 Flash
37.1
%
Claude Fable 5 (High)
36.9
%
Muse Spark 1.1 (High)
36.9
%
Claude Opus 4.6 (High)
36.3
%
Grok 4.5
35.9
%
Claude Opus 5 (Adaptive/High)
35.7
%
Claude Opus 4.8 (Max)
35.6
%
Gemini 3.6 Flash
35.5
%
Claude Opus 4.8 (High)
34.8
%
Claude Opus 4.8
34.2
%
Claude Sonnet 4.6 (High)
34
%
Qwen 3.7 Max
33.5
%
Claude Opus 4.7
31.6
%
Claude Opus 4.7 (High)
30.3
%
Kimi K2.6
30.1
%
GLM 5.2 (Max)
29.3
%
DeepSeek V4 Pro
28
%
Gemini 3.5 Flash-Lite (High)
23.2
%
GLM 5.2
22.8
%
DeepSeek V4 Flash
20.9
%
Kimi K2.5
18.2
%
Qwen 3.5 Plus
17.8
%
Nemotron 3 Ultra
17.8
%
Claude Opus 4.6
16.9
%
Grok 4.20 Beta
16.4
%
Ernie 5.1
14.7
%
Claude Sonnet 4.6
10.7
%
DeepSeek V3.2
2.2
%
Gemini 3.5 Flash-Lite
1.3
%
Mistral Large 3
0.4
%
Inkling
0.3
%
Ernie 4.5 300B
0
%
Nova 2 Pro
0
%
Professional Multimodal Reasoning

GDP.pdf

Can frontier models master the documents that run the world? GDP.pdf is a multimodal and reasoning benchmark that takes real-world prompts and PDFs pulled directly from expert professional workflows.

Rank
Model
Score
GPT 5.6 Sol
30.7
%
GPT 5.6 Sol
Claude Fable 5 (Adaptive/Max)
29.8
%
Claude Fable 5 (Adaptive/Max)
GPT 5.5 (xHigh)
26
%
GPT 5.5 (xHigh)
GPT 5.6 Terra
24.7
%
GPT 5.6 Terra
Claude Opus 5 (Adaptive/Max)
24
%
Claude Opus 5 (Adaptive/Max)
Claude Opus 4.8 (Adaptive/Max)
24
%
Claude Opus 4.8 (Adaptive/Max)
GPT 5.6 Luna
22.7
%
GPT 5.6 Luna
Claude Opus 4.7 (Adaptive/Max)
21
%
Claude Opus 4.7 (Adaptive/Max)
Kimi K3
19
%
Kimi K3
Claude Sonnet 4.6 (Adaptive Max)
18
%
Claude Sonnet 4.6 (Adaptive Max)
Gemini 3.1 Pro
17
%
Gemini 3.1 Pro
Muse Spark 1.1
15
%
Muse Spark 1.1
Grok 4.5 (High)
14
%
Grok 4.5 (High)
Gemini 3.5 Flash
14
%
Gemini 3.5 Flash
Gemini 3.6 Flash (High)
14
%
Gemini 3.6 Flash (High)
Kimi K2.6
12
%
Kimi K2.6
Gemini 3 Flash
10
%
Gemini 3 Flash
Gemini 3.5 Flash-Lite (High)
10
%
Gemini 3.5 Flash-Lite (High)
Grok 4.3 (High)
8
%
Grok 4.3 (High)
Nova 2 (Pro)
2
%
Nova 2 (Pro)
NVIDIA Nemotron 3 Nano Omni
2
%
NVIDIA Nemotron 3 Nano Omni
Mistral Large 3
2
%
Mistral Large 3
Enterprise Agents in Realistic RL Environments

EnterpriseBench: CoreCraft Agents

Stop testing models in tiny, self-contained environments. We built CoreCraft, a large-scale startup world, and deployed AI agents to solve real tasks. Our goal: to move agents beyond the cleanliness of the lab and into the chaos of enterprise reality.

Rank
Model
Score
Fable 5 (Max reasoning)
70.3
%
Fable 5 (Max reasoning)
GPT-5.5
52.8
%
GPT-5.5
Claude Opus 4.8 (Max reasoning)
52.3
%
Claude Opus 4.8 (Max reasoning)
Frontier Research Mathematics

Riemann-bench

We evaluate AI models on advanced mathematical problems requiring deep reasoning and novel synthesis. Our benchmark features problems from cutting-edge mathematics, sourced from leading mathematicians – Ivy League professors, PhD IMO medalists, graduate students at the top of their field – in the course of their research.

Rank
Model
Score
GPT-5.6 Sol (Max)
74.4
%
GPT-5.6 Sol (Max)
Claude Opus 5 (Adaptive/Max)
68
%
Claude Opus 5 (Adaptive/Max)
Claude Fable 5 (Adaptive/Max)
60
%
Claude Fable 5 (Adaptive/Max)
GPT 5.5 (xHigh)
55.2
%
GPT 5.5 (xHigh)
Claude Opus 4.8 (Adaptive/Max)
47.2
%
Claude Opus 4.8 (Adaptive/Max)
GPT 5.4 (xHigh)
41.6
%
GPT 5.4 (xHigh)
Grok 4.5 (High)
38.4
%
Grok 4.5 (High)
Kimi K3
37.6
%
Kimi K3
GPT-5.2 (xHigh)
37.6
%
GPT-5.2 (xHigh)
Gemini 3.5 Flash (High)
36.8
%
Gemini 3.5 Flash (High)
Gemini 3.1 Pro
33.6
%
Gemini 3.1 Pro
Claude Opus 4.7 (Adaptive/Max)
32.8
%
Claude Opus 4.7 (Adaptive/Max)
Gemini 3.6 Flash (High)
30.4
%
Gemini 3.6 Flash (High)
Claude Opus 4.6 (Adaptive/Max)
27.2
%
Claude Opus 4.6 (Adaptive/Max)
Muse Spark 1.1 (xHigh)
20.8
%
Muse Spark 1.1 (xHigh)
Inkling
15.2
%
Inkling
Qwen 3.7 Max
15.2
%
Qwen 3.7 Max
Kimi K2.5
12
%
Kimi K2.5
Gemini 3.5 Flash-Lite (High)
11.2
%
Gemini 3.5 Flash-Lite (High)
Claude Opus 4.5 (Adaptive/Max)
11.2
%
Claude Opus 4.5 (Adaptive/Max)
GLM 5.2
10.4
%
GLM 5.2
DeepSeek V4 Flash
10.4
%
DeepSeek V4 Flash
DeepSeek V3.2 (Thinking)
8
%
DeepSeek V3.2 (Thinking)
Kimi K2.6
7.2
%
Kimi K2.6
DeepSeek V4 Pro
5.6
%
DeepSeek V4 Pro

Stay Posted on New Benchmarks