Humanity's Last Exam
HLE-Diamond Logo

Introducing HLE-Diamond

Center for AI Safety&Scale AI
Hugging FaceDatasetload_dataset("cais/hle-diamond")

Center for AI Safety and Scale AI

We are releasing HLE-Diamond, a refined subset of Humanity’s Last Exam (HLE) question collection, following a year-long process of cleaning and refinement with input from research communities.

HLE-Diamond consists of 1,000 questions.

Main results. We compare the performance of current models on HLE-Diamond without tools.

HLE-Diamond
  • GPT-6 Astra66.2%
  • Claude Opus 5.561.2%
  • Claude Fable 5.154.1%
  • GPT-6 Sol43.8%
  • GPT-5.6 Sol41.6%
  • Claude Opus 541.3%
  • Muse Spark 1.334.8%
  • Gemini 3.8 Flash34.3%
  • Grok 4.725.3%

Models are evaluated using highest reasoning effort available. or

Dataset. HLE-Diamond consists of 500 reasoning and 500 knowledge questions, measuring reasoning and expert knowledge, respectively.

Reasoning and Knowledge Partitions
ReasoningKnowledge
  • GPT-6 Astra
    Reasoning81.6%
    Knowledge50.8%
  • Claude Opus 5.5
    Reasoning69.8%
    Knowledge52.6%
  • Claude Fable 5.1
    Reasoning62.4%
    Knowledge45.8%
  • GPT-6 Sol
    Reasoning58.4%
    Knowledge29.2%
  • GPT-5.6 Sol
    Reasoning54.4%
    Knowledge28.8%
  • Claude Opus 5
    Reasoning47.8%
    Knowledge34.8%
  • Muse Spark 1.3
    Reasoning45.0%
    Knowledge24.6%
  • Gemini 3.8 Flash
    Reasoning38.6%
    Knowledge30.0%
  • Grok 4.7
    Reasoning34.6%
    Knowledge16.0%

Models are evaluated using highest reasoning effort available.

Evaluation with tools. HLE-Diamond questions are designed to be answerable in a closed-book setting, testing both reasoning and expert knowledge. Since HLE is also used to evaluate the capabilities of agentic systems, we outline our recommended settings for evaluating HLE-Diamond with tools here.

HLE-Diamond with tools
Without toolsWith tools (web+code)
  • GPT-6 Astra
    Without tools60.6%
    With tools (web+code)82.9%
  • Claude Opus 5.5
    Without tools55.0%
    With tools (web+code)73.9%
  • Claude Fable 5.1
    Without tools51.3%
    With tools (web+code)72.4%
  • Claude Opus 5
    Without tools38.6%
    With tools (web+code)69.1%
  • Gemini 3.8 Flash
    Without tools34.3%
    With tools (web+code)60.3%
  • GPT-6 Sol
    Without tools33.8%
    With tools (web+code)64.9%
  • GPT-5.6 Sol
    Without tools31.2%
    With tools (web+code)56.5%
  • Muse Spark 1.3
    Without tools25.4%
    With tools (web+code)55.5%

All models are evaluated with reasoning high.

Acknowledgement. We extend our deepest gratitude to all participating question contributors and experts involved in creating and refining the dataset, and to the researchers whose inputs across HLE-Rolling shaped HLE-Diamond.

For any inquiries, please contact agibenchmark@safe.ai.

Citation

@article{phan2025lastexam,
      title = {A benchmark of expert-level academic questions to assess {AI} capabilities},
      author = {{Center for AI Safety} and {Scale AI} and {HLE Contributors Consortium}},
      journal = {Nature},
      volume = {649},
      pages = {1139--1146},
      year = {2026},
      doi = {10.1038/s41586-025-09962-4},
      eprint = {2501.14249},
      archivePrefix = {arXiv},
      primaryClass = {cs.LG},
      url = {https://arxiv.org/abs/2501.14249}
}