---
title: "LLM benchmarks — เกณฑ์มาตรฐาน LLM"
source: "https://systems-analysis.info/int/LLM_benchmarks_%E2%80%94_%E0%B9%80%E0%B8%81%E0%B8%93%E0%B8%91%E0%B9%8C%E0%B8%A1%E0%B8%B2%E0%B8%95%E0%B8%A3%E0%B8%90%E0%B8%B2%E0%B8%99_LLM"
wiki: "systems-analysis.info/int"
article: "LLM_benchmarks_—_เกณฑ์มาตรฐาน_LLM"
language: "th"
categories:
  - "Category:Large language models"
  - "Category:LLM benchmarks"
  - "Category:Machine learning"
  - "Category:Thai"
revision_id: 3594
wiki_created_at: 2026-09-06T23:23:08Z
wiki_modified_at: 2026-09-06T23:23:08Z
downloaded_at: 2026-09-07T22:57:46Z
---

# LLM benchmarks — เกณฑ์มาตรฐาน LLM

**เกณฑ์มาตรฐาน (benchmark) ของโมเดลภาษาขนาดใหญ่** — คือชุดทดสอบที่ได้มาตรฐาน ซึ่งออกแบบมาเพื่อวัด เปรียบเทียบ และประเมินคุณภาพและความสามารถของโมเดลภาษาขนาดใหญ่ (LLM)<sup>[\[1\]](https://systems-analysis.info/int/LLM_benchmarks_%E2%80%94_%E0%B9%80%E0%B8%81%E0%B8%93%E0%B8%91%E0%B9%8C%E0%B8%A1%E0%B8%B2%E0%B8%95%E0%B8%A3%E0%B8%90%E0%B8%B2%E0%B8%99_LLM#cite_note-ibm-benchmarks-1)</sup> โดยทั่วไป benchmark แต่ละชุดจะประกอบด้วยชุดงานที่กำหนดตายตัว (เช่น คำถาม ข้อความ หรือคำสั่ง) ซึ่งทราบคำตอบที่ถูกต้องหรือเกณฑ์การประเมินไว้ล่วงหน้า แนวทางนี้ช่วยให้เปรียบเทียบโมเดลต่าง ๆ ได้อย่างเป็นกลางในเงื่อนไขเดียวกัน ทำให้สามารถติดตามความก้าวหน้าในสาขานี้และระบุจุดแข็งจุดอ่อนของโมเดลได้<sup>[\[2\]](https://systems-analysis.info/int/LLM_benchmarks_%E2%80%94_%E0%B9%80%E0%B8%81%E0%B8%93%E0%B8%91%E0%B9%8C%E0%B8%A1%E0%B8%B2%E0%B8%95%E0%B8%A3%E0%B8%90%E0%B8%B2%E0%B8%99_LLM#cite_note-evidently-guide-2)</sup>

การใช้ benchmark อย่างสม่ำเสมอมีบทบาทสำคัญในการพัฒนา LLM โดยกระตุ้นให้นักพัฒนาปรับปรุงโมเดลและสร้างความโปร่งใสและความสามารถในการเปรียบเทียบผลลัพธ์ภายในชุมชนวิทยาศาสตร์ วิวัฒนาการของ benchmark สะท้อนให้เห็นพัฒนาการของ LLM เอง ตั้งแต่งานด้านความเข้าใจภาษาอย่างง่ายไปจนถึงการทดสอบที่ซับซ้อนซึ่งตรวจสอบการให้เหตุผลหลายขั้นตอน สามัญสำนึก จริยธรรม และความปลอดภัย<sup>[\[3\]](https://systems-analysis.info/int/LLM_benchmarks_%E2%80%94_%E0%B9%80%E0%B8%81%E0%B8%93%E0%B8%91%E0%B9%8C%E0%B8%A1%E0%B8%B2%E0%B8%95%E0%B8%A3%E0%B8%90%E0%B8%B2%E0%B8%99_LLM#cite_note-habr-popular-llm-3)</sup>

## หมวดหมู่หลักและตัวอย่าง

Benchmark ของ LLM ครอบคลุมทักษะและสาขาการประยุกต์ใช้ที่หลากหลาย ด้านล่างนี้จะกล่าวถึงหมวดหมู่หลักและชุดงานที่เป็นที่รู้จักมากที่สุดในแต่ละหมวด

### ความเข้าใจภาษาทั่วไป

หมวดหมู่นี้ประเมินความสามารถพื้นฐานของโมเดลในการทำความเข้าใจและตีความภาษาธรรมชาติ

- **GLUE** (General Language Understanding Evaluation, 2019) — หนึ่งใน benchmark แบบครบวงจรชุดแรก ๆ ที่ประกอบด้วยงานหลากหลายประเภท ตั้งแต่การตรวจจับอารมณ์ความรู้สึกไปจนถึงการประเมินความสอดคล้องเชิงตรรกะของข้อความ ผลลัพธ์จากงานทั้งหมดจะถูกรวมเป็นคะแนนเดียว ทำให้สามารถเปรียบเทียบโมเดลยุคแรก ๆ ตามประสิทธิภาพโดยรวมได้<sup>[\[4\]](https://systems-analysis.info/int/LLM_benchmarks_%E2%80%94_%E0%B9%80%E0%B8%81%E0%B8%93%E0%B8%91%E0%B9%8C%E0%B8%A1%E0%B8%B2%E0%B8%95%E0%B8%A3%E0%B8%90%E0%B8%B2%E0%B8%99_LLM#cite_note-wang2019glue-4)</sup>
- **SuperGLUE** (2019) — ผู้สืบทอดที่ "เข้มข้นขึ้น" ของ GLUE ซึ่งพัฒนาขึ้นเพื่อรับมือกับการที่โมเดลต่าง ๆ บรรลุระดับใกล้เคียงมนุษย์บน GLUE ได้อย่างรวดเร็ว SuperGLUE ประกอบด้วยงานที่ยากขึ้น ซึ่งต้องการความเข้าใจบริบทในเชิงลึกและความสามารถในการสรุปความ<sup>[\[5\]](https://systems-analysis.info/int/LLM_benchmarks_%E2%80%94_%E0%B9%80%E0%B8%81%E0%B8%93%E0%B8%91%E0%B9%8C%E0%B8%A1%E0%B8%B2%E0%B8%95%E0%B8%A3%E0%B8%90%E0%B8%B2%E0%B8%99_LLM#cite_note-wang2019superglue-5)</sup>
- **WinoGrande** (2019) — รูปแบบที่ขยายจากปริศนา Winograd Schema ประกอบด้วยงาน 44,000 ข้อสำหรับการแก้ไขความกำกวมของสรรพนามในประโยคที่ต้องใช้สามัญสำนึกในการเลือกการตีความที่ถูกต้อง<sup>[\[6\]](https://systems-analysis.info/int/LLM_benchmarks_%E2%80%94_%E0%B9%80%E0%B8%81%E0%B8%93%E0%B8%91%E0%B9%8C%E0%B8%A1%E0%B8%B2%E0%B8%95%E0%B8%A3%E0%B8%90%E0%B8%B2%E0%B8%99_LLM#cite_note-sakaguchi2019-6)</sup>

### Benchmark แบบหลายงานและแบบครบวงจร

ชุดทดสอบเหล่านี้ตรวจสอบโมเดลในช่วงกว้างของความรู้และทักษะ โดยก้าวข้ามขอบเขตของงานทางภาษาศาสตร์ล้วน ๆ

- **MMLU** (Massive Multitask Language Understanding, 2020) — ชุดงานในรูปแบบแบบทดสอบที่ครอบคลุม 57 สาขาวิชา ตั้งแต่วิชาระดับโรงเรียนไปจนถึงความรู้เฉพาะทางระดับวิชาชีพ (นิติศาสตร์ การแพทย์) MMLU วัดความกว้างของความรู้รอบด้านของโมเดล<sup>[\[7\]](https://systems-analysis.info/int/LLM_benchmarks_%E2%80%94_%E0%B9%80%E0%B8%81%E0%B8%93%E0%B8%91%E0%B9%8C%E0%B8%A1%E0%B8%B2%E0%B8%95%E0%B8%A3%E0%B8%90%E0%B8%B2%E0%B8%99_LLM#cite_note-hendrycks2020mmlu-7)</sup>
- **BIG-bench** (Beyond the Imitation Game Benchmark, 2022) — benchmark แบบร่วมมือที่ใหญ่ที่สุดในช่วงเวลาที่สร้างขึ้น โดยพัฒนาโดยผู้เขียนมากกว่า 400 คน ประกอบด้วยงานมากกว่า 200 รายการในหัวข้อที่หลากหลาย ตั้งแต่ภาษาศาสตร์ไปจนถึงฟิสิกส์ เพื่อทดสอบโมเดลนอกเหนือจากการจับคู่รูปแบบ และเปิดเผยข้อจำกัดของโมเดลในสถานการณ์ที่ไม่เป็นมาตรฐาน<sup>[\[8\]](https://systems-analysis.info/int/LLM_benchmarks_%E2%80%94_%E0%B9%80%E0%B8%81%E0%B8%93%E0%B8%91%E0%B9%8C%E0%B8%A1%E0%B8%B2%E0%B8%95%E0%B8%A3%E0%B8%90%E0%B8%B2%E0%B8%99_LLM#cite_note-srivastava2022bigbench-8)</sup>

### สามัญสำนึกและความน่าเชื่อถือ

Benchmark เหล่านี้ประเมินความสามารถของโมเดลในการสรุปความเชิงตรรกะเกี่ยวกับสถานการณ์ในชีวิตประจำวันและหลีกเลี่ยงการเผยแพร่ข้อมูลเท็จ

- **HellaSwag** (2019) — ทดสอบสามัญสำนึกผ่านงานเลือกการต่อเนื้อหาที่น่าเชื่อถือที่สุดสำหรับคำอธิบายสถานการณ์ ลักษณะพิเศษของ benchmark นี้คือการมี "กับดัก": คำตอบผิดถูกสร้างขึ้นโดยอัตโนมัติและดูสมจริงมาก ซึ่งต้องการให้โมเดลมีความเข้าใจบริบทในเชิงลึก<sup>[\[9\]](https://systems-analysis.info/int/LLM_benchmarks_%E2%80%94_%E0%B9%80%E0%B8%81%E0%B8%93%E0%B8%91%E0%B9%8C%E0%B8%A1%E0%B8%B2%E0%B8%95%E0%B8%A3%E0%B8%90%E0%B8%B2%E0%B8%99_LLM#cite_note-zellers2019hellaswag-9)</sup>
- **TruthfulQA** (2021) — วัดแนวโน้มของโมเดลในการเผยแพร่ตำนานและความเชื่อผิด ๆ ที่แพร่หลาย ประกอบด้วยคำถามที่คำตอบที่นิยมบนอินเทอร์เน็ตนั้นผิด (เช่น "วัคซีนทำให้เกิดออทิซึมหรือไม่?") โมเดลต้องไม่ยอมรับแบบแผนเท็จและให้คำตอบที่ถูกต้องตามข้อเท็จจริง<sup>[\[10\]](https://systems-analysis.info/int/LLM_benchmarks_%E2%80%94_%E0%B9%80%E0%B8%81%E0%B8%93%E0%B8%91%E0%B9%8C%E0%B8%A1%E0%B8%B2%E0%B8%95%E0%B8%A3%E0%B8%90%E0%B8%B2%E0%B8%99_LLM#cite_note-lin2021truthfulqa-10)</sup>

### โจทย์คณิตศาสตร์

- **GSM8K** (2021) — ประกอบด้วยโจทย์คณิตศาสตร์แบบอธิบายความหลายพันข้อในระดับชั้นประถมศึกษา แต่ละโจทย์ต้องการการดำเนินการตามลำดับ 2–8 ขั้นตอนทางเลขคณิตเพื่อให้ได้คำตอบ ซึ่งทดสอบความสามารถของโมเดลในการให้เหตุผลหลายขั้นตอน<sup>[\[11\]](https://systems-analysis.info/int/LLM_benchmarks_%E2%80%94_%E0%B9%80%E0%B8%81%E0%B8%93%E0%B8%91%E0%B9%8C%E0%B8%A1%E0%B8%B2%E0%B8%95%E0%B8%A3%E0%B8%90%E0%B8%B2%E0%B8%99_LLM#cite_note-cobbe2021gsm8k-11)</sup>
- **MATH** (2021) — ชุดงานที่ยากขึ้น ประกอบด้วยโจทย์จากการแข่งขันและโอลิมปิกคณิตศาสตร์ ครอบคลุมเนื้อหาพีชคณิต เรขาคณิต และทฤษฎีจำนวน ซึ่งต้องการให้โมเดลมีความเชี่ยวชาญในวิธีการแก้ปัญหาที่ไม่ธรรมดา<sup>[\[12\]](https://systems-analysis.info/int/LLM_benchmarks_%E2%80%94_%E0%B9%80%E0%B8%81%E0%B8%93%E0%B8%91%E0%B9%8C%E0%B8%A1%E0%B8%B2%E0%B8%95%E0%B8%A3%E0%B8%90%E0%B8%B2%E0%B8%99_LLM#cite_note-hendrycks2021math-12)</sup>

### การสร้างโค้ดโปรแกรม

- **HumanEval** (2021) — การทดสอบมาตรฐานสำหรับประเมินความสามารถของ LLM ในการเขียนโค้ด ประกอบด้วยงานการเขียนโปรแกรม 164 ข้อ ซึ่งโมเดลต้องสร้างโค้ด Python ที่ถูกต้องตามคำอธิบายที่กำหนด ความถูกต้องจะถูกประเมินด้วย unit test<sup>[\[13\]](https://systems-analysis.info/int/LLM_benchmarks_%E2%80%94_%E0%B9%80%E0%B8%81%E0%B8%93%E0%B8%91%E0%B9%8C%E0%B8%A1%E0%B8%B2%E0%B8%95%E0%B8%A3%E0%B8%90%E0%B8%B2%E0%B8%99_LLM#cite_note-chen2021humaneval-13)</sup>
- **SWE-bench** (2023) — benchmark ที่สมจริงยิ่งขึ้น โดยรวบรวมคำอธิบายปัญหาจริง (*issues*) จาก GitHub โมเดลต้องสร้าง patch (ส่วนของโค้ด) ที่แก้ไขปัญหาดังกล่าว ซึ่งต้องการความเข้าใจโค้ดของผู้อื่นจำนวนมากและการให้เหตุผลแบบขั้นตอนที่ซับซ้อน<sup>[\[14\]](https://systems-analysis.info/int/LLM_benchmarks_%E2%80%94_%E0%B9%80%E0%B8%81%E0%B8%93%E0%B8%91%E0%B9%8C%E0%B8%A1%E0%B8%B2%E0%B8%95%E0%B8%A3%E0%B8%90%E0%B8%B2%E0%B8%99_LLM#cite_note-jimenez2023swebench-14)</sup>

### การประเมินโมเดลสนทนา

- **Chatbot Arena** (2024) — แพลตฟอร์มออนไลน์แบบเปิดที่โมเดลนิรนามสองตัวเข้าร่วมในการสนทนาคู่กับผู้ใช้ หลังจากการสนทนา ผู้ใช้จะโหวตว่าคำตอบของใครดีกว่า จากการ "ดวล" หลายพันครั้ง จะได้คะแนน Elo ตามความชอบของผู้ใช้ ซึ่งสะท้อนคุณภาพของโมเดลในการสนทนาจริง<sup>[\[15\]](https://systems-analysis.info/int/LLM_benchmarks_%E2%80%94_%E0%B9%80%E0%B8%81%E0%B8%93%E0%B8%91%E0%B9%8C%E0%B8%A1%E0%B8%B2%E0%B8%95%E0%B8%A3%E0%B8%90%E0%B8%B2%E0%B8%99_LLM#cite_note-chiang2024chatbot-15)</sup>
- **MT-Bench** (2023) — benchmark แบบอัตโนมัติสำหรับการทดสอบความเครียดของทักษะการสนทนา ประกอบด้วยคู่คำถาม 80 คู่ที่จำลองการสนทนาหลายรอบ คำตอบของโมเดลจะถูกประเมินโดย LLM ตัวอื่นที่มีประสิทธิภาพสูงกว่า ("LLM-as-a-judge" เช่น GPT-4) ตามเกณฑ์ที่กำหนดไว้ล่วงหน้า<sup>[\[16\]](https://systems-analysis.info/int/LLM_benchmarks_%E2%80%94_%E0%B9%80%E0%B8%81%E0%B8%93%E0%B8%91%E0%B9%8C%E0%B8%A1%E0%B8%B2%E0%B8%95%E0%B8%A3%E0%B8%90%E0%B8%B2%E0%B8%99_LLM#cite_note-zheng2023mtbench-16)</sup>

### ความปลอดภัยและความน่าเชื่อถือ

- **AgentHarm** (2024) — benchmark ที่ประเมินแนวโน้มของ LLM agent ในการดำเนินการตามคำสั่งที่เป็นอันตราย ประกอบด้วย 110 สถานการณ์ที่แสดงถึงงานที่มีเจตนาร้าย (ตั้งแต่การฉ้อโกงไปจนถึงอาชญากรรมทางไซเบอร์) โมเดลที่ดีควรปฏิเสธการดำเนินการตามคำขอเหล่านั้น<sup>[\[17\]](https://systems-analysis.info/int/LLM_benchmarks_%E2%80%94_%E0%B9%80%E0%B8%81%E0%B8%93%E0%B8%91%E0%B9%8C%E0%B8%A1%E0%B8%B2%E0%B8%95%E0%B8%A3%E0%B8%90%E0%B8%B2%E0%B8%99_LLM#cite_note-andriushchenko2024agentharm-17)</sup>
- **SafetyBench** (2023) — ชุดงานกว้างที่มีคำถามมากกว่า 11,000 ข้อ ตรวจสอบว่าโมเดลหลีกเลี่ยงการสร้างเนื้อหาที่ไม่เหมาะสมและคำแนะนำที่เป็นอันตรายได้อย่างสม่ำเสมอเพียงใด รวมถึงการตอบสนองต่อคำขอที่ยั่วยุ<sup>[\[18\]](https://systems-analysis.info/int/LLM_benchmarks_%E2%80%94_%E0%B9%80%E0%B8%81%E0%B8%93%E0%B8%91%E0%B9%8C%E0%B8%A1%E0%B8%B2%E0%B8%95%E0%B8%A3%E0%B8%90%E0%B8%B2%E0%B8%99_LLM#cite_note-zhang2023safetybench-18)</sup>

## ข้อจำกัดและปัญหาที่เกิดขึ้นในปัจจุบัน

- **การปนเปื้อนของข้อมูล**: ภัยคุกคามหลักต่อความน่าเชื่อถือของการประเมิน — การรั่วไหลของข้อมูลทดสอบไปยังชุดข้อมูลการฝึก โมเดลอาจจำคำตอบได้ ซึ่งทำให้ผลลัพธ์ถูกบวมเกินจริงอย่างไม่เป็นธรรมชาติ<sup>[\[2\]](https://systems-analysis.info/int/LLM_benchmarks_%E2%80%94_%E0%B9%80%E0%B8%81%E0%B8%93%E0%B8%91%E0%B9%8C%E0%B8%A1%E0%B8%B2%E0%B8%95%E0%B8%A3%E0%B8%90%E0%B8%B2%E0%B8%99_LLM#cite_note-evidently-guide-2)</sup>
- **การอิ่มตัวของ benchmark**: เมื่อโมเดลพัฒนาขึ้น ประสิทธิภาพบน benchmark เก่า (เช่น GLUE) จะถึงเพดาน และการทดสอบก็ไม่เป็นประโยชน์ในการแยกแยะโมเดลใหม่ที่มีประสิทธิภาพสูงกว่าอีกต่อไป สิ่งนี้ต้องการการพัฒนาเกณฑ์มาตรฐานที่ซับซ้อนยิ่งขึ้นอย่างต่อเนื่อง<sup>[\[2\]](https://systems-analysis.info/int/LLM_benchmarks_%E2%80%94_%E0%B9%80%E0%B8%81%E0%B8%93%E0%B8%91%E0%B9%8C%E0%B8%A1%E0%B8%B2%E0%B8%95%E0%B8%A3%E0%B8%90%E0%B8%B2%E0%B8%99_LLM#cite_note-evidently-guide-2)</sup>
- **ช่องว่างกับความเป็นจริง**: ผลลัพธ์ที่สูงบน benchmark ไม่ได้รับประกันการทำงานที่เชื่อถือได้ของโมเดลในสถานการณ์จริงที่ไม่มีโครงสร้างเสมอไป สภาพแวดล้อมจริงมักมีความหลากหลายและคาดเดาไม่ได้มากกว่าชุดงานที่กำหนดตายตัวใด ๆ<sup>[\[1\]](https://systems-analysis.info/int/LLM_benchmarks_%E2%80%94_%E0%B9%80%E0%B8%81%E0%B8%93%E0%B8%91%E0%B9%8C%E0%B8%A1%E0%B8%B2%E0%B8%95%E0%B8%A3%E0%B8%90%E0%B8%B2%E0%B8%99_LLM#cite_note-ibm-benchmarks-1)</sup>

## ลิงก์ภายนอก

- Open LLM Leaderboard — อันดับโมเดลแบบเปิดจากชุมชน Hugging Face
- Chatbot Arena Leaderboard — อันดับโมเดลสนทนาตามความชอบของผู้ใช้

## หมายเหตุ

1.  <span id="cite_note-ibm-benchmarks-1">↑ <sup>[1.0](https://systems-analysis.info/int/LLM_benchmarks_%E2%80%94_%E0%B9%80%E0%B8%81%E0%B8%93%E0%B8%91%E0%B9%8C%E0%B8%A1%E0%B8%B2%E0%B8%95%E0%B8%A3%E0%B8%90%E0%B8%B2%E0%B8%99_LLM#cite_ref-ibm-benchmarks_1-0)</sup> <sup>[1.1](https://systems-analysis.info/int/LLM_benchmarks_%E2%80%94_%E0%B9%80%E0%B8%81%E0%B8%93%E0%B8%91%E0%B9%8C%E0%B8%A1%E0%B8%B2%E0%B8%95%E0%B8%A3%E0%B8%90%E0%B8%B2%E0%B8%99_LLM#cite_ref-ibm-benchmarks_1-1)</sup> «What Are LLM Benchmarks?». *IBM*. <a href="https://www.ibm.com/think/topics/llm-benchmarks" class="external autonumber" rel="nofollow">[1]</a></span>
2.  <span id="cite_note-evidently-guide-2">↑ <sup>[2.0](https://systems-analysis.info/int/LLM_benchmarks_%E2%80%94_%E0%B9%80%E0%B8%81%E0%B8%93%E0%B8%91%E0%B9%8C%E0%B8%A1%E0%B8%B2%E0%B8%95%E0%B8%A3%E0%B8%90%E0%B8%B2%E0%B8%99_LLM#cite_ref-evidently-guide_2-0)</sup> <sup>[2.1](https://systems-analysis.info/int/LLM_benchmarks_%E2%80%94_%E0%B9%80%E0%B8%81%E0%B8%93%E0%B8%91%E0%B9%8C%E0%B8%A1%E0%B8%B2%E0%B8%95%E0%B8%A3%E0%B8%90%E0%B8%B2%E0%B8%99_LLM#cite_ref-evidently-guide_2-1)</sup> <sup>[2.2](https://systems-analysis.info/int/LLM_benchmarks_%E2%80%94_%E0%B9%80%E0%B8%81%E0%B8%93%E0%B8%91%E0%B9%8C%E0%B8%A1%E0%B8%B2%E0%B8%95%E0%B8%A3%E0%B8%90%E0%B8%B2%E0%B8%99_LLM#cite_ref-evidently-guide_2-2)</sup> «20 LLM evaluation benchmarks and how they work». *Evidently AI*. <a href="https://www.evidentlyai.com/llm-guide/llm-benchmarks" class="external autonumber" rel="nofollow">[2]</a></span>
3.  <span id="cite_note-habr-popular-llm-3">[↑](https://systems-analysis.info/int/LLM_benchmarks_%E2%80%94_%E0%B9%80%E0%B8%81%E0%B8%93%E0%B8%91%E0%B9%8C%E0%B8%A1%E0%B8%B2%E0%B8%95%E0%B8%A3%E0%B8%90%E0%B8%B2%E0%B8%99_LLM#cite_ref-habr-popular-llm_3-0) «Самые популярные LLM бенчмарки». *Хабр*. <a href="https://habr.com/ru/articles/844974/" class="external autonumber" rel="nofollow">[3]</a></span>
4.  <span id="cite_note-wang2019glue-4">[↑](https://systems-analysis.info/int/LLM_benchmarks_%E2%80%94_%E0%B9%80%E0%B8%81%E0%B8%93%E0%B8%91%E0%B9%8C%E0%B8%A1%E0%B8%B2%E0%B8%95%E0%B8%A3%E0%B8%90%E0%B8%B2%E0%B8%99_LLM#cite_ref-wang2019glue_4-0) Wang, Alex; Singh, Amanpreet; Michael, Julian; et al. «GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding». *arXiv*. <a href="https://arxiv.org/abs/1804.07461" class="external autonumber" rel="nofollow">[4]</a></span>
5.  <span id="cite_note-wang2019superglue-5">[↑](https://systems-analysis.info/int/LLM_benchmarks_%E2%80%94_%E0%B9%80%E0%B8%81%E0%B8%93%E0%B8%91%E0%B9%8C%E0%B8%A1%E0%B8%B2%E0%B8%95%E0%B8%A3%E0%B8%90%E0%B8%B2%E0%B8%99_LLM#cite_ref-wang2019superglue_5-0) Wang, Alex; Pruksachatkun, Yada; Nangia, Nikita; et al. «SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems». *arXiv*. <a href="https://arxiv.org/abs/1905.00537" class="external autonumber" rel="nofollow">[5]</a></span>
6.  <span id="cite_note-sakaguchi2019-6">[↑](https://systems-analysis.info/int/LLM_benchmarks_%E2%80%94_%E0%B9%80%E0%B8%81%E0%B8%93%E0%B8%91%E0%B9%8C%E0%B8%A1%E0%B8%B2%E0%B8%95%E0%B8%A3%E0%B8%90%E0%B8%B2%E0%B8%99_LLM#cite_ref-sakaguchi2019_6-0) Sakaguchi, Keisuke; Le Bras, Ronan; Bhagavatula, Chandra; Choi, Yejin. «WinoGrande: An Adversarial Winograd Schema Challenge at Scale». *arXiv*. <a href="https://arxiv.org/abs/1907.10641" class="external autonumber" rel="nofollow">[6]</a></span>
7.  <span id="cite_note-hendrycks2020mmlu-7">[↑](https://systems-analysis.info/int/LLM_benchmarks_%E2%80%94_%E0%B9%80%E0%B8%81%E0%B8%93%E0%B8%91%E0%B9%8C%E0%B8%A1%E0%B8%B2%E0%B8%95%E0%B8%A3%E0%B8%90%E0%B8%B2%E0%B8%99_LLM#cite_ref-hendrycks2020mmlu_7-0) Hendrycks, Dan; Burns, Collin; Basart, Steven; et al. «Measuring Massive Multitask Language Understanding». *arXiv*. <a href="https://arxiv.org/abs/2009.03300" class="external autonumber" rel="nofollow">[7]</a></span>
8.  <span id="cite_note-srivastava2022bigbench-8">[↑](https://systems-analysis.info/int/LLM_benchmarks_%E2%80%94_%E0%B9%80%E0%B8%81%E0%B8%93%E0%B8%91%E0%B9%8C%E0%B8%A1%E0%B8%B2%E0%B8%95%E0%B8%A3%E0%B8%90%E0%B8%B2%E0%B8%99_LLM#cite_ref-srivastava2022bigbench_8-0) Srivastava, Aarohi; et al. «Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models». *arXiv*. <a href="https://arxiv.org/abs/2206.04615" class="external autonumber" rel="nofollow">[8]</a></span>
9.  <span id="cite_note-zellers2019hellaswag-9">[↑](https://systems-analysis.info/int/LLM_benchmarks_%E2%80%94_%E0%B9%80%E0%B8%81%E0%B8%93%E0%B8%91%E0%B9%8C%E0%B8%A1%E0%B8%B2%E0%B8%95%E0%B8%A3%E0%B8%90%E0%B8%B2%E0%B8%99_LLM#cite_ref-zellers2019hellaswag_9-0) Zellers, Rowan; Holtzman, Ari; Bisk, Yonatan; et al. «HellaSwag: Can a Machine Really Finish Your Sentence?». *arXiv*. <a href="https://arxiv.org/abs/1905.07830" class="external autonumber" rel="nofollow">[9]</a></span>
10. <span id="cite_note-lin2021truthfulqa-10">[↑](https://systems-analysis.info/int/LLM_benchmarks_%E2%80%94_%E0%B9%80%E0%B8%81%E0%B8%93%E0%B8%91%E0%B9%8C%E0%B8%A1%E0%B8%B2%E0%B8%95%E0%B8%A3%E0%B8%90%E0%B8%B2%E0%B8%99_LLM#cite_ref-lin2021truthfulqa_10-0) Lin, Stephanie; Hilton, Jacob; Evans, Owain. «TruthfulQA: Measuring How Models Mimic Human Falsehoods». *arXiv*. <a href="https://arxiv.org/abs/2109.07958" class="external autonumber" rel="nofollow">[10]</a></span>
11. <span id="cite_note-cobbe2021gsm8k-11">[↑](https://systems-analysis.info/int/LLM_benchmarks_%E2%80%94_%E0%B9%80%E0%B8%81%E0%B8%93%E0%B8%91%E0%B9%8C%E0%B8%A1%E0%B8%B2%E0%B8%95%E0%B8%A3%E0%B8%90%E0%B8%B2%E0%B8%99_LLM#cite_ref-cobbe2021gsm8k_11-0) Cobbe, Karl; Kosaraju, Vineet; Bavarian, Mohammad; et al. «Training Verifiers to Solve Math Word Problems». *arXiv*. <a href="https://arxiv.org/abs/2110.14168" class="external autonumber" rel="nofollow">[11]</a></span>
12. <span id="cite_note-hendrycks2021math-12">[↑](https://systems-analysis.info/int/LLM_benchmarks_%E2%80%94_%E0%B9%80%E0%B8%81%E0%B8%93%E0%B8%91%E0%B9%8C%E0%B8%A1%E0%B8%B2%E0%B8%95%E0%B8%A3%E0%B8%90%E0%B8%B2%E0%B8%99_LLM#cite_ref-hendrycks2021math_12-0) Hendrycks, Dan; Burns, Collin; Saund, Saurav; et al. «Measuring Mathematical Problem Solving With the MATH Dataset». *arXiv*. <a href="https://arxiv.org/abs/2103.03874" class="external autonumber" rel="nofollow">[12]</a></span>
13. <span id="cite_note-chen2021humaneval-13">[↑](https://systems-analysis.info/int/LLM_benchmarks_%E2%80%94_%E0%B9%80%E0%B8%81%E0%B8%93%E0%B8%91%E0%B9%8C%E0%B8%A1%E0%B8%B2%E0%B8%95%E0%B8%A3%E0%B8%90%E0%B8%B2%E0%B8%99_LLM#cite_ref-chen2021humaneval_13-0) Chen, Mark; Tworek, Jerry; Jun, Heewoo; et al. «Evaluating Large Language Models Trained on Code». *arXiv*. <a href="https://arxiv.org/abs/2107.03374" class="external autonumber" rel="nofollow">[13]</a></span>
14. <span id="cite_note-jimenez2023swebench-14">[↑](https://systems-analysis.info/int/LLM_benchmarks_%E2%80%94_%E0%B9%80%E0%B8%81%E0%B8%93%E0%B8%91%E0%B9%8C%E0%B8%A1%E0%B8%B2%E0%B8%95%E0%B8%A3%E0%B8%90%E0%B8%B2%E0%B8%99_LLM#cite_ref-jimenez2023swebench_14-0) Jimenez, Carlos E.; et al. «SWE-bench: Can Language Models Resolve Real-World GitHub Issues?». *arXiv*. <a href="https://arxiv.org/abs/2310.06770" class="external autonumber" rel="nofollow">[14]</a></span>
15. <span id="cite_note-chiang2024chatbot-15">[↑](https://systems-analysis.info/int/LLM_benchmarks_%E2%80%94_%E0%B9%80%E0%B8%81%E0%B8%93%E0%B8%91%E0%B9%8C%E0%B8%A1%E0%B8%B2%E0%B8%95%E0%B8%A3%E0%B8%90%E0%B8%B2%E0%B8%99_LLM#cite_ref-chiang2024chatbot_15-0) Chiang, Wei-Lin; et al. «Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preferences». *lmsys.org*. <a href="https://lmsys.org/blog/2023-05-03-arena/" class="external autonumber" rel="nofollow">[15]</a></span>
16. <span id="cite_note-zheng2023mtbench-16">[↑](https://systems-analysis.info/int/LLM_benchmarks_%E2%80%94_%E0%B9%80%E0%B8%81%E0%B8%93%E0%B8%91%E0%B9%8C%E0%B8%A1%E0%B8%B2%E0%B8%95%E0%B8%A3%E0%B8%90%E0%B8%B2%E0%B8%99_LLM#cite_ref-zheng2023mtbench_16-0) Zheng, Lianmin; et al. «Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena». *arXiv*. <a href="https://arxiv.org/abs/2306.05685" class="external autonumber" rel="nofollow">[16]</a></span>
17. <span id="cite_note-andriushchenko2024agentharm-17">[↑](https://systems-analysis.info/int/LLM_benchmarks_%E2%80%94_%E0%B9%80%E0%B8%81%E0%B8%93%E0%B8%91%E0%B9%8C%E0%B8%A1%E0%B8%B2%E0%B8%95%E0%B8%A3%E0%B8%90%E0%B8%B2%E0%B8%99_LLM#cite_ref-andriushchenko2024agentharm_17-0) Andriushchenko, Maksym; et al. «AgentHarm: A Benchmark for Asessing Agentic AI Harm». *arXiv*. <a href="https://arxiv.org/abs/2402.12249" class="external autonumber" rel="nofollow">[17]</a></span>
18. <span id="cite_note-zhang2023safetybench-18">[↑](https://systems-analysis.info/int/LLM_benchmarks_%E2%80%94_%E0%B9%80%E0%B8%81%E0%B8%93%E0%B8%91%E0%B9%8C%E0%B8%A1%E0%B8%B2%E0%B8%95%E0%B8%A3%E0%B8%90%E0%B8%B2%E0%B8%99_LLM#cite_ref-zhang2023safetybench_18-0) Zhang, Zhexin; et al. «SafetyBench: A Comprehensive Benchmark for Evaluating the Safety of Large Language Models». *arXiv*. <a href="https://arxiv.org/abs/2309.07045" class="external autonumber" rel="nofollow">[18]</a></span>
