---
title: "Humanity's Last Exam (benchmark) (TH)"
source: "https://systems-analysis.info/int/Humanity's_Last_Exam_(benchmark)_(TH)"
wiki: "systems-analysis.info/int"
article: "Humanity's_Last_Exam_(benchmark)_(TH)"
language: "th"
categories:
  - "Category:Large language models"
  - "Category:LLM benchmarks"
  - "Category:Machine learning"
  - "Category:Thai"
revision_id: 3033
wiki_created_at: 2026-09-06T23:14:37Z
wiki_modified_at: 2026-09-06T23:14:37Z
downloaded_at: 2026-09-07T22:54:33Z
---

# Humanity's Last Exam (benchmark) (TH)

**Humanity's Last Exam** (**HLE**, ไทย: «การสอบครั้งสุดท้ายของมวลมนุษยชาติ») — คือ benchmark ทดสอบแบบครอบคลุม ที่ออกแบบมาเพื่อประเมินความสามารถของระบบปัญญาประดิษฐ์ (AI) ขั้นสูงในงานที่ต้องใช้ความรู้และทักษะการใช้เหตุผลในระดับที่เทียบเท่ากับผู้เชี่ยวชาญมนุษย์ชั้นนำ benchmark นี้ถูกพัฒนาขึ้นในช่วงปี 2024–2025 โดยองค์กรไม่แสวงหาผลกำไร Center for AI Safety (CAIS) ร่วมกับบริษัท Scale AI<sup>[\[1\]](https://systems-analysis.info/int/Humanity's_Last_Exam_(benchmark)_(TH)#cite_note-hle_paper-1)</sup>

โครงการ HLE ถูกออกแบบให้เป็น «การสอบวิชาการครั้งสุดท้าย» สำหรับโมเดล AI — การทดสอบที่ยากเย็นอย่างถึงที่สุด ซึ่งจะช่วยระบุว่าโมเดลในปัจจุบันกำลังเข้าใกล้ระดับผู้เชี่ยวชาญหรือไม่ และยังมีช่องว่างใดเหลืออยู่ในความสามารถของโมเดลเหล่านั้น<sup>[\[1\]](https://systems-analysis.info/int/Humanity's_Last_Exam_(benchmark)_(TH)#cite_note-hle_paper-1)</sup> benchmark ประกอบด้วยคำถามที่ยากอย่างยิ่งจำนวน 2,500 ข้อ ครอบคลุมสาขาวิชามากกว่าหนึ่งร้อยสาขา<sup>[\[2\]](https://systems-analysis.info/int/Humanity's_Last_Exam_(benchmark)_(TH)#cite_note-wiki_hle-2)</sup>

## ประวัติการสร้าง

เมื่อถึงกลางทศวรรษ 2020 โมเดลภาษาขนาดใหญ่อย่าง GPT-4 และ Claude ได้แสดงผลลัพธ์สูงเป็นอย่างมากในชุดทดสอบยอดนิยม (เช่น MMLU) จนทำให้ benchmark หลายรายการไม่สามารถใช้เป็นมาตรวัดความก้าวหน้าที่เชื่อถือได้อีกต่อไป การสอบในระดับปริญญาตรีมาตรฐานถูกโมเดลทำได้อย่าง «ถล่มทลาย» จนทำให้การประเมินการปรับปรุงในขั้นต่อไปอย่างเป็นกลางกลายเป็นสิ่งที่เป็นไปไม่ได้<sup>[\[3\]](https://systems-analysis.info/int/Humanity's_Last_Exam_(benchmark)_(TH)#cite_note-reuters_stump-3)</sup>

ในสถานการณ์เช่นนี้ **Dan Hendrycks** ผู้อำนวยการ CAIS และนักวิจัย AI ที่มีชื่อเสียง ได้เสนอแนวคิด «การสอบครั้งสุดท้ายของมวลมนุษยชาติ» — ชุดคำถามที่มีความซับซ้อนสูงสุด ซึ่งสามารถแยกแยะความสามารถของ AI ออกจากระดับของผู้เชี่ยวชาญที่แท้จริงได้ แรงบันดาลใจมาจากการสนทนากับนักธุรกิจ Elon Musk ซึ่งแสดงความเห็นว่าการทดสอบที่มีอยู่ในปัจจุบันนั้นง่ายเกินไปแล้ว<sup>[\[2\]](https://systems-analysis.info/int/Humanity's_Last_Exam_(benchmark)_(TH)#cite_note-wiki_hle-2)</sup>

เพื่อให้แนวคิดนี้เป็นจริง CAIS ได้ร่วมมือกับ Scale AI โดยเมื่อวันที่ 15 กันยายน 2024 มีการประกาศอย่างเป็นทางการถึงการรวบรวมคำถามที่ยากที่สุดจากทั่วโลกสำหรับการสอบในอนาคต ผู้จัดงานได้เชิญชวนนักวิทยาศาสตร์และผู้เชี่ยวชาญทั่วโลกให้ส่งโจทย์ที่สามารถทำให้แม้แต่โมเดล AI ที่ก้าวหน้าที่สุดต้องหยุดชะงัก เพื่อสร้างแรงจูงใจให้ผู้เข้าร่วม จึงมีการจัดตั้งกองทุนรางวัลมูลค่า 500,000 ดอลลาร์<sup>[\[3\]](https://systems-analysis.info/int/Humanity's_Last_Exam_(benchmark)_(TH)#cite_note-reuters_stump-3)</sup>

การคัดเลือกโจทย์เกิดขึ้นเป็นหลายขั้นตอน ในตอนแรก คำถามที่ส่งมาจะถูกกรองผ่านโมเดล AI ขั้นสูง หากอัลกอริทึมสามารถแก้โจทย์ได้อย่างมั่นใจ โจทย์นั้นจะถูกคัดออกว่าไม่ยากเพียงพอ โจทย์ที่ AI ไม่สามารถแก้ได้จะผ่านการตรวจสอบโดยผู้เชี่ยวชาญเพื่อประเมินความถูกต้องและการมีคำตอบที่ถูกต้องเพียงหนึ่งเดียว ในที่สุด ผู้เชี่ยวชาญเกือบ 1,000 คนจากสถาบันการศึกษาและวิจัยมากกว่า 500 แห่งได้มีส่วนร่วมในการสร้างชุดโจทย์นี้<sup>[\[4\]](https://systems-analysis.info/int/Humanity's_Last_Exam_(benchmark)_(TH)#cite_note-hle_site-4)</sup>

เวอร์ชันสุดท้ายของ benchmark ซึ่งประกอบด้วย **2,500 คำถาม** ได้ถูกนำเสนอในต้นปี 2025 โจทย์บางส่วนถูกเก็บไว้ในคลังสำรองแบบปิดเพื่อการทดสอบควบคุมและป้องกันการปรับแต่งโมเดลให้เข้ากับชุดคำถามที่กำหนดไว้<sup>[\[2\]](https://systems-analysis.info/int/Humanity's_Last_Exam_(benchmark)_(TH)#cite_note-wiki_hle-2)</sup>

## โครงสร้างและเนื้อหาของ Benchmark

ชุดคำถาม HLE ครอบคลุมสาขาวิชาความรู้เชิงวิชาการในวงกว้างอย่างมาก โจทย์ถูกแบ่งตามหัวข้อดังนี้:

- **คณิตศาสตร์**: ~41%
- **ชีววิทยาและการแพทย์**: ~11%
- **วิทยาการคอมพิวเตอร์และ AI**: ~10%
- **ฟิสิกส์**: ~9%
- **มนุษยศาสตร์และสังคมศาสตร์**: ~9%
- **เคมี**: ~7%
- **วิศวกรรมศาสตร์**: ~4%
- **สาขาอื่น ๆ**: ~9%

ประมาณ **14%** ของโจทย์ทั้งหมดเป็น **โจทย์แบบ multimodal** กล่าวคือ การแก้โจทย์เหล่านี้ต้องอาศัยการวิเคราะห์ภาพ (รูปวาด ไดอะแกรม ข้อความในภาพ)<sup>[\[2\]](https://systems-analysis.info/int/Humanity's_Last_Exam_(benchmark)_(TH)#cite_note-wiki_hle-2)</sup> โจทย์ส่วนใหญ่ (ประมาณ 3/4) เป็น **คำถามปลายเปิดที่ต้องตอบสั้น ๆ** ซึ่งโมเดลต้องสร้างคำตอบที่แม่นยำด้วยตนเอง (ตัวเลข คำศัพท์ ชื่อ) ส่วนที่เหลือเป็นคำถามแบบเลือกตอบ

โจทย์ทุกข้อใน HLE มีคุณสมบัติร่วมกัน ดังนี้:

- **ความยากสูงอย่างยิ่ง**: ปัญหาแต่ละข้อต้องการความรู้และทักษะในระดับที่เทียบเท่ากับผู้เชี่ยวชาญที่มีคุณสมบัติในสาขานั้น ๆ<sup>[\[5\]](https://systems-analysis.info/int/Humanity's_Last_Exam_(benchmark)_(TH)#cite_note-techradar_pass-5)</sup>
- **คำตอบที่ตรวจสอบได้**: คำถามแต่ละข้อมีคำตอบที่ถูกต้องซึ่งกำหนดไว้ชัดเจนและพิสูจน์ได้
- **ทนทานต่อการค้นหา**: โจทย์ถูกคัดเลือกมาเพื่อให้ไม่สามารถหาคำตอบได้ด้วยการค้นหาแบบง่าย ๆ การทำโจทย์ให้สำเร็จต้องอาศัยความเข้าใจเชิงลึกในเนื้อหาและการใช้เหตุผล<sup>[\[1\]](https://systems-analysis.info/int/Humanity's_Last_Exam_(benchmark)_(TH)#cite_note-hle_paper-1)</sup>

## ผลการทดสอบโมเดล

Humanity's Last Exam พิสูจน์ชื่อเสียงว่าเป็นการทดสอบที่ยากเย็นอย่างสุดขีดทันที: **ไม่มีโมเดล AI ใดในปัจจุบันที่สามารถแสดงผลลัพธ์ใกล้เคียงกับระดับมนุษย์บน benchmark นี้ได้** โมเดลภาษาชั้นนำในปี 2025 แสดงความแม่นยำในระดับที่ต่ำมาก

- **GPT-4** หลายเวอร์ชันจาก OpenAI และ **Claude** จาก Anthropic แสดงผลลัพธ์ **ต่ำกว่า 10%**<sup>[\[4\]](https://systems-analysis.info/int/Humanity's_Last_Exam_(benchmark)_(TH)#cite_note-hle_site-4)</sup>
- ผลลัพธ์สูงสุดในบรรดา LLM มาตรฐานเป็นของโมเดล **Gemini 2.5 Pro** (Google DeepMind) ด้วยความแม่นยำประมาณ **21.6%**<sup>[\[4\]](https://systems-analysis.info/int/Humanity's_Last_Exam_(benchmark)_(TH)#cite_note-hle_site-4)</sup>
- แม้แต่โมเดลที่ดีที่สุดก็ตอบคำถาม HLE ผิดประมาณ 4/5 ของคำถามทั้งหมด ซึ่งเน้นย้ำให้เห็นถึงขนาดของช่องว่างระหว่างความสามารถปัจจุบันของ AI กับระดับผู้เชี่ยวชาญมนุษย์<sup>[\[1\]](https://systems-analysis.info/int/Humanity's_Last_Exam_(benchmark)_(TH)#cite_note-hle_paper-1)</sup>

ผลลัพธ์ที่น่าสนใจเป็นพิเศษคือผลของ agent ทดลอง **ChatGPT Deep Research** จาก OpenAI ซึ่งได้รับอนุญาตให้ทำการค้นหาโดยอัตโนมัติ โดยการจำลองการทำงานของนักวิจัย agent นี้สามารถแก้โจทย์ได้อย่างถูกต้องถึง **26.6%** — ผลลัพธ์ที่สูงกว่าโมเดลใด ๆ ที่ไม่มีเครื่องมือดังกล่าวมากกว่า 2 เท่า แต่ยังคงห่างไกลจากคะแนนผ่านอย่างมาก<sup>[\[6\]](https://systems-analysis.info/int/Humanity's_Last_Exam_(benchmark)_(TH)#cite_note-hindustan_times_26-6)</sup>

## ความสำคัญและแนวโน้ม

การปรากฏขึ้นของ HLE ถือเป็นเหตุการณ์สำคัญในชุมชน AI เนื่องจาก benchmark นี้ได้เติมเต็มความต้องการเร่งด่วนในการมีมาตรวัดความก้าวหน้าแบบใหม่ที่ซับซ้อนยิ่งขึ้น

- **จุดอ้างอิงร่วมกัน** HLE มอบเครื่องมือเชิงวัตถุวิสัยสำหรับนักวิจัยและผู้กำหนดนโยบายในการประเมินความสามารถของ AI ช่วยให้สามารถติดตามพลวัตของการปรับปรุงและเข้าใจว่าเครื่องจักรกำลังเข้าใกล้ระดับมนุษย์มากเพียงใด
- **เครื่องมือสำหรับสนับสนุนนโยบาย** การมีการทดสอบอ้างอิงเช่นนี้ส่งเสริมการอภิปรายที่เป็นรูปธรรมมากขึ้นเกี่ยวกับทิศทางการพัฒนา AI ความเสี่ยงที่อาจเกิดขึ้น และมาตรการกำกับดูแลที่จำเป็น
- **เส้นแบ่งสุดท้ายของการทดสอบเชิงวิชาการ** ชื่อ «การสอบครั้งสุดท้าย» สะท้อนแนวคิดที่ว่าชุดโจทย์นี้อาจเป็นการสอบแบบปิดครั้งสุดท้ายสำหรับการประเมิน AI การผ่าน HLE ได้อย่างมั่นใจจะหมายความว่า ในแง่ของความรู้เชิงวิชาการและทักษะการใช้เหตุผลที่ตรวจสอบได้อย่างเข้มงวด เครื่องจักรได้บรรลุระดับของผู้เชี่ยวชาญมนุษย์ชั้นนำแล้ว<sup>[\[4\]](https://systems-analysis.info/int/Humanity's_Last_Exam_(benchmark)_(TH)#cite_note-hle_site-4)</sup>

สิ่งสำคัญที่ต้องสังเกตคือ แม้แต่การผ่าน HLE ได้อย่างสมบูรณ์ก็ไม่ได้หมายความว่าบรรลุปัญญาประดิษฐ์ทั่วไป (AGI) เนื่องจากการทดสอบไม่ได้ตรวจสอบความสามารถด้านความคิดสร้างสรรค์ ความคิดริเริ่ม หรือความสามารถในการตั้งคำถามทางวิทยาศาสตร์ใหม่ ๆ<sup>[\[4\]](https://systems-analysis.info/int/Humanity's_Last_Exam_(benchmark)_(TH)#cite_note-hle_site-4)</sup>

ด้วยความก้าวหน้าอย่างรวดเร็ว นักวิจัยคาดการณ์ว่าโมเดลอาจมีความแม่นยำเกิน 50% บน HLE ภายในสิ้นปี 2025 ซึ่งจะหมายความว่าเครื่องจักรได้เข้าใกล้ระดับมนุษย์อย่างใกล้ชิดในด้านความรู้เชิงวิชาการซึ่งเป็นตัวชี้วัดที่แคบแต่มีความสำคัญ<sup>[\[4\]](https://systems-analysis.info/int/Humanity's_Last_Exam_(benchmark)_(TH)#cite_note-hle_site-4)</sup>

## ลิงก์

- เว็บไซต์อย่างเป็นทางการของ Humanity's Last Exam
- บทความวิชาการที่นำเสนอ benchmark

## วรรณกรรม

- Liang, P. et al. (2022). *Holistic Evaluation of Language Models (HELM)*. arXiv:2211.09110.
- Chang, Y. et al. (2023). *A Survey on Evaluation of Large Language Models*. arXiv:2307.03109.
- Ni, S. et al. (2025). *A Survey on Large Language Model Benchmarks*. arXiv:2508.15361.
- Biderman, S. et al. (2024). *The Language Model Evaluation Harness (lm-eval): Guidance and Lessons Learned*. arXiv:2405.14782.
- Kiela, D. et al. (2021). *Dynabench: Rethinking Benchmarking in NLP*. arXiv:2104.14337.
- Ma, Z. et al. (2021). *Dynaboard: An Evaluation‑As‑A‑Service Platform for Holistic Next‑Generation Benchmarking*. arXiv:2106.06052.
- Goel, K. et al. (2021). *Robustness Gym: Unifying the NLP Evaluation Landscape*. arXiv:2101.04840.
- Xu, C. et al. (2024). *Benchmark Data Contamination of Large Language Models: A Survey*. arXiv:2406.04244.
- Liu, S. et al. (2025). *A Comprehensive Survey on Safety Evaluation of LLMs*. arXiv:2506.11094.
- Chiang, W.-L. et al. (2024). *Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference*. arXiv:2403.04132.
- Boubdir, M. et al. (2023). *Elo Uncovered: Robustness and Best Practices in Language Model Evaluation*. arXiv:2311.17295.
- Huang, L. et al. (2023). *A Survey on Hallucination in Large Language Models*. arXiv:2311.05232.

## หมายเหตุ

1.  <span id="cite_note-hle_paper-1">↑ <sup>[1.0](https://systems-analysis.info/int/Humanity's_Last_Exam_(benchmark)_(TH)#cite_ref-hle_paper_1-0)</sup> <sup>[1.1](https://systems-analysis.info/int/Humanity's_Last_Exam_(benchmark)_(TH)#cite_ref-hle_paper_1-1)</sup> <sup>[1.2](https://systems-analysis.info/int/Humanity's_Last_Exam_(benchmark)_(TH)#cite_ref-hle_paper_1-2)</sup> <sup>[1.3](https://systems-analysis.info/int/Humanity's_Last_Exam_(benchmark)_(TH)#cite_ref-hle_paper_1-3)</sup> Fan, L. et al. «Humanity's Last Exam: A New Benchmark for AI Alignment». *arXiv:2501.14249*, 2025. <a href="https://arxiv.org/abs/2501.14249" class="external autonumber" rel="nofollow">[1]</a></span>
2.  <span id="cite_note-wiki_hle-2">↑ <sup>[2.0](https://systems-analysis.info/int/Humanity's_Last_Exam_(benchmark)_(TH)#cite_ref-wiki_hle_2-0)</sup> <sup>[2.1](https://systems-analysis.info/int/Humanity's_Last_Exam_(benchmark)_(TH)#cite_ref-wiki_hle_2-1)</sup> <sup>[2.2](https://systems-analysis.info/int/Humanity's_Last_Exam_(benchmark)_(TH)#cite_ref-wiki_hle_2-2)</sup> <sup>[2.3](https://systems-analysis.info/int/Humanity's_Last_Exam_(benchmark)_(TH)#cite_ref-wiki_hle_2-3)</sup> «Humanity's Last Exam». In *Wikipedia*. <a href="https://en.wikipedia.org/wiki/Humanity%27s_Last_Exam" class="external autonumber" rel="nofollow">[2]</a></span>
3.  <span id="cite_note-reuters_stump-3">↑ <sup>[3.0](https://systems-analysis.info/int/Humanity's_Last_Exam_(benchmark)_(TH)#cite_ref-reuters_stump_3-0)</sup> <sup>[3.1](https://systems-analysis.info/int/Humanity's_Last_Exam_(benchmark)_(TH)#cite_ref-reuters_stump_3-1)</sup> Dastin, J. & Paul, K. «AI experts ready 'Humanity's Last Exam' to stump powerful tech». *Reuters*, 2024. <a href="https://www.reuters.com/technology/artificial-intelligence/ai-experts-ready-humanitys-last-exam-stump-powerful-tech-2024-09-16/" class="external autonumber" rel="nofollow">[3]</a></span>
4.  <span id="cite_note-hle_site-4">↑ <sup>[4.0](https://systems-analysis.info/int/Humanity's_Last_Exam_(benchmark)_(TH)#cite_ref-hle_site_4-0)</sup> <sup>[4.1](https://systems-analysis.info/int/Humanity's_Last_Exam_(benchmark)_(TH)#cite_ref-hle_site_4-1)</sup> <sup>[4.2](https://systems-analysis.info/int/Humanity's_Last_Exam_(benchmark)_(TH)#cite_ref-hle_site_4-2)</sup> <sup>[4.3](https://systems-analysis.info/int/Humanity's_Last_Exam_(benchmark)_(TH)#cite_ref-hle_site_4-3)</sup> <sup>[4.4](https://systems-analysis.info/int/Humanity's_Last_Exam_(benchmark)_(TH)#cite_ref-hle_site_4-4)</sup> <sup>[4.5](https://systems-analysis.info/int/Humanity's_Last_Exam_(benchmark)_(TH)#cite_ref-hle_site_4-5)</sup> «Humanity's Last Exam». *Center for AI Safety*. <a href="https://agi.safe.ai/" class="external autonumber" rel="nofollow">[4]</a></span>
5.  <span id="cite_note-techradar_pass-5">[↑](https://systems-analysis.info/int/Humanity's_Last_Exam_(benchmark)_(TH)#cite_ref-techradar_pass_5-0) «Could you pass 'Humanity's Last Exam'? Probably not, but neither can AI». *TechRadar*. <a href="https://www.techradar.com/computing/artificial-intelligence/could-you-pass-humanitys-last-exam-probably-not-but-neither-can-ai" class="external autonumber" rel="nofollow">[5]</a></span>
6.  <span id="cite_note-hindustan_times_26-6">[↑](https://systems-analysis.info/int/Humanity's_Last_Exam_(benchmark)_(TH)#cite_ref-hindustan_times_26_6-0) «OpenAI's deep research can complete 26% of 'Humanity's Last Exam': What is it and what does it mean?». *Hindustan Times*. <a href="https://www.hindustantimes.com/technology/openais-deep-research-can-complete-26-of-humanity-s-last-exam-what-is-it-and-what-does-it-mean-101739355881687.html" class="external autonumber" rel="nofollow">[6]</a></span>
