We are releasing HLE-Diamond, a refined subset of Humanity’s Last Exam (HLE) question collection, following a year-long process of cleaning and refinement with input from research communities.
HLE-Diamond consists of 1,000 questions.
Main results. We compare the performance of current models on HLE-Diamond without tools.
GPT-6 Astra66.2%
Claude Opus 5.561.2%
Claude Fable 5.154.1%
GPT-6 Sol43.8%
GPT-5.6 Sol41.6%
Claude Opus 541.3%
Muse Spark 1.334.8%
Gemini 3.8 Flash34.3%
Grok 4.725.3%
Models are evaluated using highest reasoning effort available. or
Dataset. HLE-Diamond consists of 500 reasoning and 500 knowledge questions, measuring reasoning and expert knowledge, respectively.
GPT-6 Astra
Reasoning81.6%Knowledge50.8%Claude Opus 5.5
Reasoning69.8%Knowledge52.6%Claude Fable 5.1
Reasoning62.4%Knowledge45.8%GPT-6 Sol
Reasoning58.4%Knowledge29.2%GPT-5.6 Sol
Reasoning54.4%Knowledge28.8%Claude Opus 5
Reasoning47.8%Knowledge34.8%Muse Spark 1.3
Reasoning45.0%Knowledge24.6%Gemini 3.8 Flash
Reasoning38.6%Knowledge30.0%Grok 4.7
Reasoning34.6%Knowledge16.0%
Models are evaluated using highest reasoning effort available.
Evaluation with tools. HLE-Diamond questions are designed to be answerable in a closed-book setting, testing both reasoning and expert knowledge. Since HLE is also used to evaluate the capabilities of agentic systems, we outline our recommended settings for evaluating HLE-Diamond with tools here.
GPT-6 Astra
Without tools60.6%With tools (web+code)82.9%Claude Opus 5.5
Without tools55.0%With tools (web+code)73.9%Claude Fable 5.1
Without tools51.3%With tools (web+code)72.4%Claude Opus 5
Without tools38.6%With tools (web+code)69.1%Gemini 3.8 Flash
Without tools34.3%With tools (web+code)60.3%GPT-6 Sol
Without tools33.8%With tools (web+code)64.9%GPT-5.6 Sol
Without tools31.2%With tools (web+code)56.5%Muse Spark 1.3
Without tools25.4%With tools (web+code)55.5%
All models are evaluated with reasoning high.
Acknowledgement. We extend our deepest gratitude to all participating question contributors and experts involved in creating and refining the dataset, and to the researchers whose inputs across HLE-Rolling shaped HLE-Diamond.
For any inquiries, please contact agibenchmark@safe.ai.
Citation
@article{phan2025lastexam,
title = {A benchmark of expert-level academic questions to assess {AI} capabilities},
author = {{Center for AI Safety} and {Scale AI} and {HLE Contributors Consortium}},
journal = {Nature},
volume = {649},
pages = {1139--1146},
year = {2026},
doi = {10.1038/s41586-025-09962-4},
eprint = {2501.14249},
archivePrefix = {arXiv},
primaryClass = {cs.LG},
url = {https://arxiv.org/abs/2501.14249}
}