---
title: "Humanity's Last Exam (benchmark)"
source: "https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)"
wiki: "systems-analysis.info/eng"
article: "Humanity's_Last_Exam_(benchmark)"
language: "en"
categories:
  - "Category:English"
  - "Category:Large language models"
  - "Category:LLM benchmarks"
  - "Category:Machine learning"
  - "Category:Technology"
revision_id: 183
wiki_created_at: 2026-09-06T22:18:34Z
wiki_modified_at: 2026-09-06T22:18:34Z
downloaded_at: 2026-09-07T22:21:39Z
---

# Humanity's Last Exam (benchmark)

**Humanity's Last Exam** (**HLE**) is a benchmark designed to evaluate the capabilities of advanced artificial intelligence (AI) systems on closed-ended academic tasks requiring knowledge and reasoning at the level of top human experts. The benchmark was developed in 2024–2025 by the non-profit organization *Center for AI Safety (CAIS)* in collaboration with the data company *Scale AI*<sup>[\[1\]](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_note-hle_paper-1)</sup>. A peer-reviewed version was later published in the journal *Nature* in January 2026 under the title "A benchmark of expert-level academic questions to assess AI capabilities"<sup>[\[2\]](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_note-nature_hle-2)</sup>.

HLE is conceived as a "final academic exam" for AI models—an exceptionally difficult, closed-ended test intended to measure how close modern models are to an expert level and where the gaps in their abilities remain<sup>[\[1\]](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_note-hle_paper-1)</sup>. The [benchmark](https://systems-analysis.info/eng/LLM_benchmarks "LLM benchmarks") consists of a publicly released set of 2,500 questions covering over one hundred different disciplines, supplemented by a private holdout set of questions used to detect overfitting<sup>[\[1\]](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_note-hle_paper-1)</sup>.

## Background

By the mid-2020s, leading language models such as GPT-4 and Claude had reached such high performance on popular test suites (for example, MMLU) that many benchmarks ceased to be a reliable measure of progress. Where these tests had once been a challenging frontier, frontier models now exceeded 90% accuracy on them, making it difficult to objectively assess further improvements<sup>[\[3\]](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_note-reuters_stump-3)</sup>. Stanford HAI's *AI Index 2025* report later cited Humanity's Last Exam as one of the "more challenging benchmarks" developed specifically because the popular tests had reached saturation<sup>[\[4\]](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_note-ai_index-4)</sup>.

In this context, **Dan Hendrycks**—director of CAIS and the creator of the earlier [MMLU benchmark](https://systems-analysis.info/eng/MMLU_Benchmark "MMLU Benchmark")—proposed the idea of a "last exam": a set of maximally difficult questions that could distinguish AI capabilities from the level of a true expert. According to Hendrycks, the direct impetus came from a conversation with entrepreneur Elon Musk, who considered existing tests too easy, remarking that the [MMLU](https://systems-analysis.info/eng/MMLU_Benchmark "MMLU Benchmark") questions were "undergrad level" and that he wanted problems "a world-class expert could do"<sup>[\[5\]](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_note-nyt_origin-5)</sup>. The project was initially given a working title along the lines of "The Last Stand," which was changed before the public launch to "Humanity's Last Exam" as being less apocalyptic while better matching the closed-form, exam-style format of the questions<sup>[\[5\]](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_note-nyt_origin-5)</sup>.

## Creation of the benchmark

To implement the idea, CAIS partnered with Scale AI, whose chief executive Alexandr Wang and research director Summer Yue helped lead the effort alongside Hendrycks. On September 16, 2024, a global call for the most difficult questions for the future exam was announced. Scientists and specialists worldwide were invited to submit problems capable of stumping the most advanced AI models, and a prize fund of \$500,000 was established to attract contributions<sup>[\[3\]](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_note-reuters_stump-3)</sup>. The reward structure paid \$5,000 for each of the top 50 questions and \$500 for the next 500; authors of accepted questions were also offered co-authorship on the resulting paper<sup>[\[6\]](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_note-wiki_hle-6)</sup>. The submission deadline, initially set for November 1, 2024, was later extended to November 15, 2024.

The selection of problems proceeded in several stages. First, submissions were filtered using frontier AI models: if the models answered a question correctly (or, for multiple-choice items, did better than random guessing), the question was discarded as not difficult enough. Questions that the models failed were then passed to human experts, who reviewed them in two rounds to verify correctness and ensure a single, unambiguous, verifiable answer that could not be obtained through a simple internet search. From an initial pool of tens of thousands of submissions, nearly 1,000 experts from over 500 institutions across 50 countries—mostly professors, researchers, and graduate-degree holders—contributed to the final dataset<sup>[\[7\]](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_note-hle_site-7)</sup>.

The benchmark was first announced in late January 2025, with the arXiv preprint submitted on January 24, 2025<sup>[\[1\]](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_note-hle_paper-1)</sup>. After release, a public "community feedback bug bounty" program ran until March 21, 2025 to identify and remove erroneous or searchable questions; on April 3, 2025 the dataset was finalized at **2,500 publicly released questions**, with flagged items removed and replaced<sup>[\[7\]](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_note-hle_site-7)</sup>. A portion of the questions is withheld in a private holdout set for control testing and to prevent models from overfitting to a fixed set<sup>[\[6\]](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_note-wiki_hle-6)</sup>. The peer-reviewed version of the work appeared in *Nature* on January 28, 2026<sup>[\[2\]](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_note-nature_hle-2)</sup>.

## Structure and content of the benchmark

The HLE question set covers a wide range of academic disciplines. The questions are distributed by subject area as follows<sup>[\[2\]](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_note-nature_hle-2)</sup>:

- **Mathematics**: ~41%
- **Biology and Medicine**: ~11%
- **Computer Science and AI**: ~10%
- **Physics**: ~9%
- **Humanities and Social Sciences**: ~9%
- **Chemistry**: ~7%
- **Engineering**: ~4%
- **Other fields**: ~9%

Approximately **14%** of all tasks are **[multimodal](https://systems-analysis.info/eng/Multimodal_reasoning "Multimodal reasoning")**, requiring the analysis of an accompanying image (a diagram, figure, or inscription) in addition to text<sup>[\[2\]](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_note-nature_hle-2)</sup>. About **76%** of the questions are **exact-match (short-answer)** items, where the model must independently produce a precise answer (a number, term, or string); the remaining **24%** are **multiple-choice** questions with five or more options<sup>[\[2\]](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_note-nature_hle-2)</sup>. The subject matter is deliberately specialized—examples range from translating ancient Palmyrene inscriptions to identifying microanatomical structures in birds or analyzing features of Biblical Hebrew pronunciation<sup>[\[2\]](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_note-nature_hle-2)</sup>.

All tasks in HLE share common properties:

- **Extremely high difficulty**: each problem requires knowledge and skill comparable to that of a qualified specialist in the field<sup>[\[8\]](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_note-techradar_pass-8)</sup>.
- **Verifiable answer**: each question has a specific, provably correct answer suitable for automated grading.
- **Resistance to search**: the tasks are designed so that the answer cannot be found with a simple search query; success requires deep understanding and reasoning<sup>[\[1\]](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_note-hle_paper-1)</sup>.

For automated grading, the benchmark's maintainers use a separate model (o3-mini) as an extractor and judge to compare a model's response against the ground-truth answer, and they note that results can vary slightly with different judge models and prompts<sup>[\[7\]](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_note-hle_site-7)</sup>.

## Model performance results

Humanity's Last Exam immediately confirmed its reputation as an extremely challenging test: at release, **no contemporary AI model came close to expert-level performance**, and the gap has narrowed only gradually as models have improved. Because scores depend heavily on the date and on whether external tools are allowed, the figures below should be read as a time series rather than a fixed ranking.

- **Initial release (early 2025, no tools).** Leading models scored in the single digits. GPT-4o reached roughly 3%, while the strongest reasoning models of the moment—OpenAI's o1 and DeepSeek-R1—answered only about 9% of questions correctly<sup>[\[7\]](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_note-hle_site-7)</sup>.
- **Tool-augmented agents (February 2025).** OpenAI's experimental **Deep Research** agent, built on the o3 model and allowed to perform automatic web searches, correctly solved **26.6%** of the tasks—roughly three times the best tool-free score at the time, though still far from a passing grade. Its access to search makes direct comparison with tool-free models uneven<sup>[\[9\]](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_note-hindustan_times_26-9)</sup>.
- **Through 2025 (no tools).** Newer reasoning models pushed tool-free accuracy into the low 20s. Google's Gemini 2.5 Pro was reported at about 18.8% at its March 2025 launch and later in the low-20s, while xAI's Grok 4 reached roughly 25% by mid-year<sup>[\[7\]](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_note-hle_site-7)</sup>.
- **Current state (2026).** The frontier has continued to climb steeply. By early 2026, top models on the official leaderboard reached the mid-30s; by mid-2026, public leaderboards placed the leading systems in the mid-40s and above. Reported text-only figures included the mid-40s for Google's Gemini 3.x Pro line and the low-to-mid 40s for the strongest GPT-5 variants, while on the Artificial Analysis board Anthropic's Claude Fable 5 and Claude Opus 4.8 held the top positions (about 53% and 46% respectively), among the first public-leaderboard results above the 50% mark<sup>[\[7\]](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_note-hle_site-7)[\[10\]](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_note-aa_leaderboard-10)</sup>.

A notable secondary finding concerns **calibration**. At initial publication, models combined low accuracy (under 10%) with very high stated confidence (calibration error above 80%)—strong evidence that the systems were confabulating rather than recognizing the limits of their own knowledge<sup>[\[7\]](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_note-hle_site-7)</sup>.

## Criticism and limitations

Despite its influence, HLE has drawn substantive criticism.

- **Trivia versus intelligence.** Some contributors and observers have questioned whether answering highly specialized, graduate-level questions is a meaningful measure of general intelligence. Kevin Zhou, a theoretical-physics researcher at UC Berkeley who contributed questions, noted the large gap between exam performance and genuine research capability<sup>[\[6\]](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_note-wiki_hle-6)</sup>. (The peer-reviewed *Nature* version used the more descriptive title "A benchmark of expert-level academic questions to assess AI capabilities.")
- **Answer accuracy.** In July 2025, the AI research organization FutureHouse published an investigation reporting that roughly **29%** (95% confidence interval ±3.7 percentage points) of the **321 text-only biology/health and chemistry questions** it audited had answers that conflicted with the peer-reviewed literature. The audit deliberately covered only those subsets rather than the benchmark as a whole, and used FutureHouse's PaperQA2 research agent to cross-check each answer's rationale against published evidence<sup>[\[11\]](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_note-futurehouse-11)</sup>. In a follow-up, the HLE team conducted its own targeted re-review of a biology, chemistry, and health subset and reported an expert-disagreement rate of about **18%**, arguing that part of the gap reflects genuine disagreement among experts on very hard questions rather than outright errors; FutureHouse separately released a curated "HLE Bio/Chem Gold" subset of validated items<sup>[\[12\]](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_note-hle_followup-12)</sup>.
- **Systematic revision.** By 2026 these reliability concerns had prompted systematic re-verification of the dataset. The 2026 "HLE-Verified" effort re-audited the public set through expert review and model-based cross-checks, certifying 1,811 of the 2,500 questions (either verified as correct or repaired under preserved evaluation intent) and releasing the remaining 689 as a documented "uncertain" set; the authors report that verification and repair measurably shift downstream model scores, indicating that a portion of measured performance had reflected annotation artifacts rather than genuine capability differences<sup>[\[13\]](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_note-hle_verified-13)</sup>.
- **Compensation and process.** Some contributors reported unclear payment structures and shifting timelines during development, with several PhD-level experts expressing frustration over ambiguous expectations around the \$500,000 prize pool<sup>[\[6\]](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_note-wiki_hle-6)</sup>.
- **Not a test of general intelligence.** By design, HLE measures structured, closed-ended academic knowledge and reasoning. It does not evaluate creativity, initiative, open-ended research ability, or the skill of posing new scientific questions, so even a perfect score would not by itself indicate artificial general intelligence (AGI)<sup>[\[7\]](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_note-hle_site-7)</sup>.

## Significance and outlook

The emergence of HLE was a significant event in the AI community, as the benchmark filled a pressing need for a new, more challenging measure of progress.

- **A common baseline.** HLE offers researchers and policymakers a standardized—if methodology-sensitive—reference point for assessing AI capabilities, allowing them to track improvement over time and gauge how close machines are to the level of human experts.
- **A tool to inform policy.** A shared, standardized reference test supports more substantive discussion of AI development trajectories, potential risks, and possible governance measures.
- **The final frontier of academic testing.** The name "Last Exam" reflects the idea that this set of problems could be the final closed-book exam needed to evaluate AI. Passing HLE would mean that, in terms of formal knowledge and rigorously verifiable reasoning, a machine had reached the level of the best human experts<sup>[\[7\]](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_note-hle_site-7)</sup>.

The authors predicted that, given the pace of progress, models might exceed 50% accuracy on HLE by the end of 2025<sup>[\[7\]](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_note-hle_site-7)</sup>. In practice this threshold was not reached on that timeline—at the close of 2025 the strongest models stood roughly in the mid-20s to high-30s on the official leaderboard—and public leaderboards did not show frontier systems clearly crossing the 50% mark until 2026, roughly half a year later than projected. Crossing it means that machines have come very close to expert level on a narrow but important measure of academic knowledge, while still leaving open-ended scientific and creative work as a separate, unmeasured frontier.

## Model results by source

HLE scores are not absolute: they depend on the evaluation date, the judge model and prompt used for grading, whether external tools (e.g. web search) are allowed, and whether the multimodal questions are included or only the text-only subset. Different leaderboards therefore report different numbers for the same model, and figures are best compared *within* a single source rather than across sources. The two leaderboards below—the benchmark's own (Scale AI / SEAL) and an independent evaluator (Artificial Analysis)—illustrate this: Gemini 3.1 Pro Preview is reported at 47.3% on the Scale text-only board and 44.7% by Artificial Analysis, and the two boards do not even agree on the current leader. Both tables are point-in-time snapshots that change frequently; each is dated below, and for current standings the live leaderboards should be consulted.

### Scale AI / SEAL — official leaderboard (text-only)

The benchmark's own leaderboard, run with the Center for AI Safety. Models are evaluated on the text-only subset (about 86% of the dataset) at temperature 0; grading uses o3-mini as an automatic extractor and judge. Rank is the statistical upper bound (a model is ranked above another only when its lower 95% confidence bound exceeds the other's upper bound), so models can share a rank. "Calib. err." is the RMS calibration error—how far a model's stated confidence is from its actual accuracy; high values indicate overconfidence. The standings below are a mid-2026 snapshot and change frequently<sup>[\[14\]](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_note-scale_leaderboard-14)</sup>.

| Rank (UB)                                                      | Model                                  | Accuracy, % (±95% CI) | Calib. err. |
|----------------------------------------------------------------|----------------------------------------|-----------------------|-------------|
| 1                                                              | Gemini 3.1 Pro Preview (thinking high) | 47.3 ±2.1             | 50          |
| 1                                                              | GPT-5.4 Pro                            | 45.3 ±2.1             | 37          |
| 3                                                              | Muse Spark                             | 40.9 ±2.1             | 51          |
| 3                                                              | Gemini 3 Pro Preview                   | 37.7 ±2.0             | 57          |
| 4                                                              | GPT-5.4 (xhigh thinking)               | 36.5 ±2.0             | 42          |
| 4                                                              | Claude Opus 4.6 Thinking Max           | 36.2 ±2.0             | 46          |
| 5                                                              | GPT-5 Pro                              | 33.3 ±2.0             | 49          |
| 8                                                              | GPT-5.2                                | 28.5 ±1.9             | 45          |
| 8                                                              | Claude Opus 4.5 Thinking               | 26.3 ±1.9             | 55          |
| 8                                                              | GPT-5                                  | 26.3 ±1.9             | 50          |
| 9                                                              | GPT-5.1 Thinking                       | 24.7 ±1.8             | 54          |
| 11                                                             | Gemini 2.5 Pro Preview (06-05)         | 22.1 ±1.8             | 72          |
| *For reference — initial-era models (late 2024 / early 2025):* |                                        |                       |             |
| 34                                                             | o1 (December 2024)                     | 7.8 ±1.1              | 84          |
| 48                                                             | Claude 3.5 Sonnet (October 2024)       | 4.3 ±0.9              | 83          |
| 59                                                             | GPT-4o (November 2024)                 | 2.3 ±0.6              | 88          |

### Artificial Analysis — independent leaderboard

An independent evaluator that runs its own harness on the **text-only subset** of HLE (2,158 questions), excluding the multimodal questions for cross-model comparability. Its figures are not directly comparable to the Scale board above because of differences in methodology, judge model, and question set. The top ten below are read directly from the Artificial Analysis leaderboard as of 1 July 2026; standings change frequently<sup>[\[10\]](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_note-aa_leaderboard-10)</sup>.

| \#  | Model                          | Accuracy, % |
|-----|--------------------------------|-------------|
| 1   | Claude Fable 5 (with fallback) | 53.3        |
| 2   | Claude Opus 4.8 (max)          | 45.7        |
| 3   | Gemini 3.1 Pro Preview         | 44.7        |
| 4   | GPT-5.5 (xhigh)                | 44.3        |
| 5   | GPT-5.5 (high)                 | 43.0        |
| 6   | Gemini 3.5 Flash               | 41.0        |
| 7   | GLM-5.2 (max)                  | 40.1        |
| 8   | Muse Spark                     | 39.9        |
| 9   | Claude Sonnet 5 (max)          | 39.6        |
| 10  | Qwen3.7 Max                    | 38.1        |

## External links

- <a href="https://agi.safe.ai/" class="external text" rel="nofollow">Official website of Humanity's Last Exam</a>
- <a href="https://www.nature.com/articles/s41586-025-09962-4" class="external text" rel="nofollow">Peer-reviewed paper in <em>Nature</em> (2026)</a>
- <a href="https://arxiv.org/abs/2501.14249" class="external text" rel="nofollow">Research paper presenting the benchmark (arXiv)</a>
- <a href="https://en.wikipedia.org/wiki/Humanity%27s_Last_Exam" class="external text" rel="nofollow">Humanity's Last Exam — Wikipedia</a>
- <a href="https://github.com/centerforaisafety/hle" class="external text" rel="nofollow">Humanity's Last Exam — Github</a>
- <a href="https://artificialanalysis.ai/evaluations/humanitys-last-exam" class="external text" rel="nofollow">Humanity's Last Exam Leaderboard — artificialanalysis.ai</a>
- <a href="https://labs.scale.com/leaderboard/humanitys_last_exam" class="external text" rel="nofollow">HHumanity's Last Exam - Scale Labs Leaderboard</a>

## Literature

- Liang, P. et al. (2022). *Holistic Evaluation of Language Models (HELM)*. <a href="https://arxiv.org/abs/2211.09110" class="external text" rel="nofollow">arXiv:2211.09110</a>.
- Chang, Y. et al. (2023). *A Survey on Evaluation of Large Language Models*. <a href="https://arxiv.org/abs/2307.03109" class="external text" rel="nofollow">arXiv:2307.03109</a>.
- Ni, S. et al. (2025). *A Survey on Large Language Model Benchmarks*. <a href="https://arxiv.org/abs/2508.15361" class="external text" rel="nofollow">arXiv:2508.15361</a>.
- Biderman, S. et al. (2024). *The Language Model Evaluation Harness (lm-eval): Guidance and Lessons Learned*. <a href="https://arxiv.org/abs/2405.14782" class="external text" rel="nofollow">arXiv:2405.14782</a>.
- Kiela, D. et al. (2021). *Dynabench: Rethinking Benchmarking in NLP*. <a href="https://arxiv.org/abs/2104.14337" class="external text" rel="nofollow">arXiv:2104.14337</a>.
- Ma, Z. et al. (2021). *Dynaboard: An Evaluation‑As‑A‑Service Platform for Holistic Next‑Generation Benchmarking*. <a href="https://arxiv.org/abs/2106.06052" class="external text" rel="nofollow">arXiv:2106.06052</a>.
- Goel, K. et al. (2021). *Robustness Gym: Unifying the NLP Evaluation Landscape*. <a href="https://arxiv.org/abs/2101.04840" class="external text" rel="nofollow">arXiv:2101.04840</a>.
- Xu, C. et al. (2024). *Benchmark Data Contamination of Large Language Models: A Survey*. <a href="https://arxiv.org/abs/2406.04244" class="external text" rel="nofollow">arXiv:2406.04244</a>.
- Liu, S. et al. (2025). *A Comprehensive Survey on Safety Evaluation of LLMs*. <a href="https://arxiv.org/abs/2506.11094" class="external text" rel="nofollow">arXiv:2506.11094</a>.
- Chiang, W.-L. et al. (2024). *Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference*. <a href="https://arxiv.org/abs/2403.04132" class="external text" rel="nofollow">arXiv:2403.04132</a>.
- Boubdir, M. et al. (2023). *Elo Uncovered: Robustness and Best Practices in Language Model Evaluation*. <a href="https://arxiv.org/abs/2311.17295" class="external text" rel="nofollow">arXiv:2311.17295</a>.
- Huang, L. et al. (2023). *A Survey on Hallucination in Large Language Models*. <a href="https://arxiv.org/abs/2311.05232" class="external text" rel="nofollow">arXiv:2311.05232</a>.

## References

1.  <span id="cite_note-hle_paper-1">↑ <sup>[1.0](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_ref-hle_paper_1-0)</sup> <sup>[1.1](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_ref-hle_paper_1-1)</sup> <sup>[1.2](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_ref-hle_paper_1-2)</sup> <sup>[1.3](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_ref-hle_paper_1-3)</sup> <sup>[1.4](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_ref-hle_paper_1-4)</sup> Phan, L., Gatti, A., Han, Z. et al. "Humanity's Last Exam". *arXiv:2501.14249*, 2025. <a href="https://arxiv.org/abs/2501.14249" class="external autonumber" rel="nofollow">[1]</a></span>
2.  <span id="cite_note-nature_hle-2">↑ <sup>[2.0](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_ref-nature_hle_2-0)</sup> <sup>[2.1](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_ref-nature_hle_2-1)</sup> <sup>[2.2](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_ref-nature_hle_2-2)</sup> <sup>[2.3](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_ref-nature_hle_2-3)</sup> <sup>[2.4](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_ref-nature_hle_2-4)</sup> <sup>[2.5](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_ref-nature_hle_2-5)</sup> "A benchmark of expert-level academic questions to assess AI capabilities". *Nature*, vol. 649, pp. 1139–1146, 2026. DOI: 10.1038/s41586-025-09962-4. <a href="https://www.nature.com/articles/s41586-025-09962-4" class="external autonumber" rel="nofollow">[2]</a></span>
3.  <span id="cite_note-reuters_stump-3">↑ <sup>[3.0](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_ref-reuters_stump_3-0)</sup> <sup>[3.1](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_ref-reuters_stump_3-1)</sup> Dastin, J. & Paul, K. "AI experts ready 'Humanity's Last Exam' to stump powerful tech". *Reuters*, 2024. <a href="https://www.reuters.com/technology/artificial-intelligence/ai-experts-ready-humanitys-last-exam-stump-powerful-tech-2024-09-16/" class="external autonumber" rel="nofollow">[3]</a></span>
4.  <span id="cite_note-ai_index-4">[↑](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_ref-ai_index_4-0) Maslej, N. et al. *The AI Index 2025 Annual Report*. Stanford Institute for Human-Centered AI, April 2025, pp. 141–142.</span>
5.  <span id="cite_note-nyt_origin-5">↑ <sup>[5.0](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_ref-nyt_origin_5-0)</sup> <sup>[5.1](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_ref-nyt_origin_5-1)</sup> Roose, K. "When A.I. Passes This Test, Look Out". *The New York Times*, 23 January 2025.</span>
6.  <span id="cite_note-wiki_hle-6">↑ <sup>[6.0](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_ref-wiki_hle_6-0)</sup> <sup>[6.1](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_ref-wiki_hle_6-1)</sup> <sup>[6.2](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_ref-wiki_hle_6-2)</sup> <sup>[6.3](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_ref-wiki_hle_6-3)</sup> "Humanity's Last Exam". In *Wikipedia*. <a href="https://en.wikipedia.org/wiki/Humanity%27s_Last_Exam" class="external autonumber" rel="nofollow">[4]</a></span>
7.  <span id="cite_note-hle_site-7">↑ <sup>[7.00](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_ref-hle_site_7-0)</sup> <sup>[7.01](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_ref-hle_site_7-1)</sup> <sup>[7.02](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_ref-hle_site_7-2)</sup> <sup>[7.03](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_ref-hle_site_7-3)</sup> <sup>[7.04](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_ref-hle_site_7-4)</sup> <sup>[7.05](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_ref-hle_site_7-5)</sup> <sup>[7.06](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_ref-hle_site_7-6)</sup> <sup>[7.07](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_ref-hle_site_7-7)</sup> <sup>[7.08](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_ref-hle_site_7-8)</sup> <sup>[7.09](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_ref-hle_site_7-9)</sup> "Humanity's Last Exam". *Center for AI Safety*. <a href="https://agi.safe.ai/" class="external autonumber" rel="nofollow">[5]</a></span>
8.  <span id="cite_note-techradar_pass-8">[↑](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_ref-techradar_pass_8-0) "Could you pass 'Humanity's Last Exam'? Probably not, but neither can AI". *TechRadar*. <a href="https://www.techradar.com/computing/artificial-intelligence/could-you-pass-humanitys-last-exam-probably-not-but-neither-can-ai" class="external autonumber" rel="nofollow">[6]</a></span>
9.  <span id="cite_note-hindustan_times_26-9">[↑](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_ref-hindustan_times_26_9-0) "OpenAI's deep research can complete 26% of 'Humanity's Last Exam': What is it and what does it mean?". *Hindustan Times*. <a href="https://www.hindustantimes.com/technology/openais-deep-research-can-complete-26-of-humanity-s-last-exam-what-is-it-and-what-does-it-mean-101739355881687.html" class="external autonumber" rel="nofollow">[7]</a></span>
10. <span id="cite_note-aa_leaderboard-10">↑ <sup>[10.0](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_ref-aa_leaderboard_10-0)</sup> <sup>[10.1](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_ref-aa_leaderboard_10-1)</sup> "Humanity's Last Exam Benchmark Leaderboard". *Artificial Analysis*. <a href="https://artificialanalysis.ai/evaluations/humanitys-last-exam" class="external autonumber" rel="nofollow">[8]</a></span>
11. <span id="cite_note-futurehouse-11">[↑](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_ref-futurehouse_11-0) FutureHouse. "About 30% of Humanity's Last Exam chemistry/biology answers are likely wrong", 2025. <a href="https://www.futurehouse.org/research/hle-exam" class="external autonumber" rel="nofollow">[9]</a></span>
12. <span id="cite_note-hle_followup-12">[↑](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_ref-hle_followup_12-0) The HLE organizing team revised its preprint in response to the FutureHouse audit and, in a September 2025 re-review of a biology/chemistry/health subset, reported an expert-disagreement rate of roughly 18%. See arXiv:2501.14249 (revised) and the FutureHouse update at <a href="https://www.futurehouse.org/research/hle-exam" class="external autonumber" rel="nofollow">[10]</a>.</span>
13. <span id="cite_note-hle_verified-13">[↑](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_ref-hle_verified_13-0) Zhai, W., Wang, Z., Wang, J. et al. "HLE-Verified: A Systematic Verification and Structured Revision of Humanity's Last Exam". *arXiv:2602.13964*, 2026. <a href="https://arxiv.org/abs/2602.13964" class="external autonumber" rel="nofollow">[11]</a></span>
14. <span id="cite_note-scale_leaderboard-14">[↑](https://systems-analysis.info/eng/Humanity's_Last_Exam_(benchmark)#cite_ref-scale_leaderboard_14-0) "Humanity's Last Exam (Text Only)". *Scale AI / SEAL Leaderboards*. <a href="https://scale.com/leaderboard/humanitys_last_exam_text_only" class="external autonumber" rel="nofollow">[12]</a></span>
