---
title: "SuperGLUE (benchmark) (TH)"
source: "https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)"
wiki: "systems-analysis.info/int"
article: "SuperGLUE_(benchmark)_(TH)"
language: "th"
categories:
  - "Category:Large language models"
  - "Category:LLM benchmarks"
  - "Category:Machine learning"
  - "Category:Thai"
revision_id: 7016
wiki_created_at: 2026-09-07T00:14:37Z
wiki_modified_at: 2026-09-07T00:14:37Z
downloaded_at: 2026-09-07T23:16:57Z
---

# SuperGLUE (benchmark) (TH)

**SuperGLUE** — คือ **benchmark** ที่ครอบคลุม (ชุดงานทดสอบ) สำหรับประเมินระบบประมวลผลภาษาธรรมชาติ โดยเฉพาะ **โมเดลภาษาขนาดใหญ่** (Large Language Models)<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup> ได้รับการเสนอในปี 2019 โดยกลุ่มนักวิจัยภายใต้การนำของ Alex Wang จากมหาวิทยาลัยนิวยอร์ก โดยมี Facebook AI Research และองค์กรอื่น ๆ ร่วมด้วย<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup>

การสร้าง SuperGLUE มีสาเหตุมาจากที่ว่า ในช่วงกลางปี 2019 benchmark รุ่นก่อนอย่าง GLUE กลายเป็น "งานง่าย" สำหรับโมเดลสมัยใหม่ โดยคะแนนรวมของโมเดลที่ดีที่สุดบน GLUE สูงถึง 88.4 ซึ่งเกินกว่าระดับเฉลี่ยของมนุษย์ (87.1)<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup> ส่งผลให้ช่องว่างสำหรับความก้าวหน้าในอนาคตลดลง<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup> ด้วยเหตุนี้ ผู้เขียนจึงพัฒนา SuperGLUE ขึ้นเป็นทางเลือกที่ยากกว่า ซึ่งสามารถให้การทดสอบความเข้าใจภาษาของโมเดลได้อย่างเข้มงวดยิ่งขึ้น<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup> เป้าหมายของ SuperGLUE คือการจัดเตรียมตัวชี้วัดความก้าวหน้าที่เป็นกลางและยาก "ต่อการเรียนรู้" สำหรับการทำความเข้าใจภาษาอังกฤษโดยรวม<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup> คาดว่าการปรับปรุงผลลัพธ์อย่างเห็นได้ชัดบน SuperGLUE จะต้องอาศัยนวัตกรรมที่สำคัญในวิธี Machine Learning เช่น การเรียนรู้จากตัวอย่างขนาดเล็กที่มีประสิทธิภาพมากขึ้น การเรียนรู้แบบหลายงาน และการเรียนรู้แบบ self-supervised<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup> กล่าวอีกนัยหนึ่ง SuperGLUE รวมงานที่ง่ายสำหรับมนุษย์แต่ยากสำหรับปัญญาประดิษฐ์<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup> เพื่อกระตุ้นการพัฒนาโมเดลที่มีความเข้าใจภาษาอย่างแท้จริงและลึกซึ้ง

## ลักษณะเฉพาะและความแตกต่างจาก GLUE

SuperGLUE ยึดตามรูปแบบของ GLUE เป็นส่วนใหญ่ โดยเสนอ **ตัวชี้วัดคุณภาพรวม** เพียงตัวเดียวสำหรับงานทั้งหมด **ลีดเดอร์บอร์ด** สาธารณะ และ **ชุดเครื่องมือ** สำหรับการวิเคราะห์โมเดล<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup> อย่างไรก็ตาม SuperGLUE นำมาซึ่งการปรับปรุงและนวัตกรรมหลายประการเมื่อเทียบกับรุ่นก่อน<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup>:

- **งานที่ยากขึ้น**: SuperGLUE คัดเลือก **แปดงานที่ยากที่สุด**<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup> สองในนั้นสืบทอดมาจาก GLUE (ในฐานะงานที่ยากที่สุดในนั้น) ส่วนที่เหลือถูกเลือกจากผู้สมัครใหม่โดยพิจารณาจากความยากสำหรับโมเดล NLP สมัยใหม่<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup> ดังนั้น benchmark จึงมุ่งเน้นที่ด้านของความเข้าใจที่โมเดลเคยแสดงผลลัพธ์แย่ที่สุด
- **ความหลากหลายของรูปแบบ**: หากใน GLUE งานทั้งหมดล้วนเป็นการจำแนกประโยคหรือคู่ประโยค SuperGLUE จะรวม **ช่วงรูปแบบที่กว้างขึ้น**<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup> นอกจากการจำแนกแล้ว ยังเพิ่มงาน **การแก้ไข coreference** และ **การตอบคำถาม** ซึ่งต้องการให้โมเดลเข้าใจข้อความที่เชื่อมโยงกันและ **การอนุมานเชิงตรรกะ**<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup>
- **การประเมินของมนุษย์สำหรับทุกงาน**: สำหรับแต่ละงานใน SuperGLUE มีการคำนวณ **ระดับประสิทธิภาพพื้นฐานของมนุษย์** (non-expert)<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup> ซึ่งยืนยันว่าแม้แต่โมเดลที่แข็งแกร่งอย่าง BERT ก็ยังด้อยกว่ามนุษย์อย่างมีนัยสำคัญในช่วงเวลาที่เปิดตัว benchmark<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup> การมี **เกณฑ์อ้างอิงของมนุษย์** (~90% โดยรวม) ช่วยให้มีช่องว่างสำหรับการเติบโตของโมเดลและทำหน้าที่เป็นเป้าหมาย<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup>
- **กฎและเครื่องมือที่โปร่งใส**: มีการทบทวนกฎการโพสต์ผลลัพธ์บนลีดเดอร์บอร์ด (เพื่อให้การเปรียบเทียบเป็นธรรมและระบุผลงานของผู้เขียน dataset)<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup> นอกจากนี้ยังเผยแพร่ชุดเครื่องมือแบบเปิดใหม่สำหรับการ fine-tuning และการเรียนรู้แบบหลายงานของโมเดลบนข้อมูล SuperGLUE ได้อย่างสะดวก<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup>

โดยรวมแล้ว มาตรการเหล่านี้ทำให้ SuperGLUE เป็นการทดสอบที่น่าเชื่อถือมากขึ้นสำหรับ **ความสามารถทางภาษาโดยรวม** ของโมเดล โดยไม่สามารถบรรลุผลลัพธ์สูงด้วยการโกงแบบแคบ ๆ หรือการปรับแต่งให้เข้ากับรูปแบบเฉพาะของ GLUE รุ่นเดิม<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup>

## ชุดงานของ SuperGLUE

SuperGLUE ประกอบด้วย **แปดงาน** ที่ครอบคลุมด้านต่าง ๆ ของการทำความเข้าใจข้อความ

- **BoolQ** (Boolean Questions): งานประเภท **ถาม-ตอบ (QA)** โดยแต่ละตัวอย่างมีข้อความสั้น ๆ (ข้อความจาก Wikipedia) และคำถามที่ต้องตอบว่า "ใช่" หรือ "ไม่"<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup> คำถามถูกสร้างโดยผู้ใช้ (จากการค้นหาใน Google) และต้องการการดึงข้อเท็จจริงที่ชัดเจนหรือโดยนัยจากข้อความ โดยตัวชี้วัดคุณภาพคือความแม่นยำ (accuracy)<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup>
- **CB** (CommitmentBank): งาน **การอนุมานเชิงตรรกะ** (textual entailment) ที่มีสามคลาส<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup> dataset ประกอบด้วยข้อความสั้นที่มีประโยคซับซ้อน โดยต้องพิจารณาว่าผู้เขียนข้อความ **มุ่งมั่นต่อความจริง** ของคำกล่าวที่ฝังอยู่มากน้อยเพียงใด<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup> โดยแท้จริงแล้ว นี่คือการตรวจสอบว่าคำกล่าวสามารถสรุปได้จากบริบทที่ให้มาหรือไม่ งานนี้ยากเนื่องจากขนาดตัวอย่างที่เล็ก (ประมาณ 250 ตัวอย่าง) และความไม่สมดุลของคลาส โดยประเมินคุณภาพด้วยความแม่นยำและค่า F1 เฉลี่ยตามคลาส<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup>
- **COPA** (Choice of Plausible Alternatives): งาน **การใช้เหตุผลเชิงเหตุและผล**<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup> โมเดลได้รับประโยคเดียวเป็นสมมติฐาน และต้องเลือกสาเหตุหรือผลลัพธ์ที่ถูกต้องจากสองตัวเลือก<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup> ตัวอย่างทั้งหมดของ COPA ถูกสร้างด้วยมือและต้องการ **สามัญสำนึก** เพื่อสร้างความสัมพันธ์เชิงเหตุและผล เนื้อหาครอบคลุมสถานการณ์จากบล็อกและสารานุกรมเฉพาะทาง โดยตัวชี้วัดคือความแม่นยำ (สัดส่วนของการเลือกที่ถูกต้อง)<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup> ตัวอย่าง: กำหนดประโยค "เด็กได้รับภูมิคุ้มกันต่อโรค" และคำถาม "สาเหตุคืออะไร?" — มนุษย์เข้าใจทันทีว่าคำตอบที่ถูกต้องคือ "เขาได้รับวัคซีน" ในขณะที่โมเดลต้องคาดเดาความสัมพันธ์เชิงเหตุผล<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup>
- **MultiRC** (Multi-Sentence Reading Comprehension): งาน **การทำความเข้าใจข้อความแบบหลายประโยค** ที่มีองค์ประกอบของการเลือกหลายตัวเลือก<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup> โมเดลได้รับย่อหน้าข้อความ คำถามเกี่ยวกับเนื้อหา และรายการคำตอบที่เป็นไปได้ โดยต้องระบุว่าคำตอบใดถูกต้อง (แต่ละคำถามอาจมีคำตอบที่ถูกต้องหลายข้อ)<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup> ลักษณะพิเศษ: การตอบคำถามโดยทั่วไปต้องรวมข้อมูลจากหลายประโยคของข้อความ ซึ่งทดสอบความสามารถของโมเดลใน **การเชื่อมโยงข้อเท็จจริง**<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup> คุณภาพวัดด้วยสองตัวชี้วัด ได้แก่ F1 ตามคำตอบ (คำนึงถึงชุดที่ถูกต้องบางส่วน) และ Exact Match คือสัดส่วนของคำถามที่ได้รับชุดคำตอบที่ถูกต้องครบถ้วน<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup>
- **ReCoRD** (Reading Comprehension with Commonsense Reasoning Dataset): งาน **การอ่านเพื่อทำความเข้าใจโดยใช้ความรู้**<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup> เป็นการทดสอบแบบ Cloze ที่ดัดแปลง: กำหนดข้อความข่าว (บทความจาก CNN/Daily Mail) และประโยคที่มีคำนามว่างเปล่า โดยโมเดลต้องเลือกว่าหน่วยงานใดจากข้อความเหมาะกับช่องว่าง<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup> ตัวเลือกคำตอบถูกกำหนดให้เป็นหน่วยงานทั้งหมดที่กล่าวถึงในบทความ ซึ่งอาจมีความหมายที่ตรงกัน<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup> การแก้ปัญหาสำเร็จต้องการการทำความเข้าใจบริบทและสามัญสำนึก โดยตัวชี้วัดคือ token-level F1 สูงสุด และ Exact Match (การจับคู่แบบตรงทั้งหมด) ของคำตอบที่ทำนาย<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup>
- **RTE** (Recognizing Textual Entailment): งานการจำแนกแบบไบนารีสำหรับ **การอนุมานจากข้อความ** (entailment vs. not entailment)<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup> ชุดข้อมูลรวมตัวอย่างจากการแข่งขันหลายรายการในการรู้จำการอนุมานจากข้อความ (ซีรีส์ RTE 1-5)<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup> แต่ละงานมีคู่ข้อความ (premise-hypothesis) โดยโมเดลต้องพิจารณาว่า hypothesis ตามมาจากข้อความหรือไม่ ต่างจาก dataset ขนาดใหญ่อื่น ๆ RTE ค่อนข้างเล็ก (ประมาณ 2,500 ตัวอย่างการฝึกอบรม) แต่แสดงให้เห็นถึงประโยชน์อย่างมีนัยสำคัญจาก transfer learning โดยความแม่นยำเพิ่มขึ้นจาก ~56% (ระดับการเดาสุ่ม) ไปเป็น ~86% ด้วยการปรากฏของโมเดลประเภท BERT<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup> อย่างไรก็ตาม ณ เวลาที่เปิดตัว SuperGLUE ความแม่นยำของโมเดลยังคงล้าหลังมนุษย์ประมาณ 8 เปอร์เซ็นต์พอยต์<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup> ดังนั้น RTE จึงถูกรวมไว้ในฐานะหนึ่งในงานที่ยังคงมีช่องว่างกับระดับของมนุษย์
- **WiC** (Word-in-Context): งาน **การแก้ไขความคลุมเครือของความหมายของคำในบริบท** (WSD)<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup> กำหนดสองประโยคอิสระซึ่งแต่ละประโยคมีคำที่มีความหมายหลายนัยเดียวกัน โดยต้องพิจารณาว่าคำนั้น **ใช้ในความหมายเดียวกัน** ในทั้งสองกรณีหรือไม่<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup> ข้อมูลนำมาจากแหล่งทรัพยากรพจนานุกรม (WordNet, VerbNet, Wiktionary) จึงครอบคลุมคำและความหมายที่หลากหลาย<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup> งานนี้ถูกจัดให้เป็นการจำแนกแบบไบนารีและประเมินด้วยสัดส่วนคำตอบที่ถูกต้อง WiC ต้องการให้โมเดลเข้าใจความแตกต่างทางความหมายที่ละเอียดอ่อน ซึ่งแท้จริงแล้วคือการทดสอบ **ความหมายทางคำศัพท์**
- **WSC** (Winograd Schema Challenge): งาน **การแก้ไข coreference โดยใช้สามัญสำนึก**<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup> แต่ละงานประกอบด้วยประโยคเดียวที่มีสรรพนาม และรายการหน่วยงาน (คำนาม) สองหน่วยจากประโยคเดียวกัน<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup> ต้องพิจารณาว่า **สรรพนามที่กำหนดอ้างถึง** คำนามใดในสองคำนั้น<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup> ตัวอย่างประโยค winograd แบบคลาสสิก: "ถ้วยรางวัลไม่พอดีกับกระเป๋าเดินทาง เพราะมันเล็กเกินไป" — มนุษย์เข้าใจว่า "มัน" อ้างถึงกระเป๋าเดินทาง (กระเป๋าเดินทางเล็กเกินไป) ตัวอย่างเช่นนี้ไม่สามารถแก้ไขได้หากปราศจาก **ความรู้ในชีวิตประจำวันและบริบท**<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup> ใน GLUE มีรูปแบบที่ง่ายกว่าของงานนี้อยู่แล้ว (WNLI) แต่โมเดลต่าง ๆ ไม่สามารถเอาชนะแม้กระทั่งระดับความสุ่มได้เป็นเวลานาน<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup> เฉพาะเทคนิคพิเศษ เช่น การเพิ่มข้อมูลภายนอกที่มีตัวอย่างคล้ายกัน เท่านั้นที่ยกระดับคุณภาพของโมเดลบน WSC ไปถึง ~90% ภายในปี 2019<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup> อย่างไรก็ตาม มนุษย์แก้งาน WSC ได้แทบไม่มีข้อผิดพลาด (~96-100% ของคำตอบที่ถูกต้อง)<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup> ใน SuperGLUE รวมเวอร์ชันดั้งเดิมของ WSC ในรูปแบบการจำแนกแบบไบนารี (สำหรับแต่ละคู่ "สรรพนาม-หน่วยงาน" โมเดลตอบว่าอ้างถึงกันหรือไม่)<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup> งานนี้ยังคงเป็นหนึ่งในการทดสอบที่ยากที่สุดซึ่งต้องการการใช้เหตุผลเชิง commonsense

การทดสอบทั้งหมดของ SuperGLUE มี **ชุดทดสอบแบบปิด** โดยคำตอบที่ไม่เป็นที่รู้จักสำหรับนักพัฒนา<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup> โมเดลส่งการทำนายไปยังเซิร์ฟเวอร์ซึ่งคำนวณคะแนนรวม — ความแม่นยำเฉลี่ยตามงาน (สำหรับงานที่มีหลายตัวชี้วัด จะเฉลี่ยตัวชี้วัดภายในก่อน)<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup> **คะแนน SuperGLUE** เพียงตัวเดียวนี้ช่วยให้เปรียบเทียบโมเดลตามระดับปัญญาทางภาษาโดยรวมได้ง่ายขึ้น

## ผลลัพธ์และความก้าวหน้าของโมเดล

เมื่อเปิดตัว SuperGLUE ผู้เขียนได้อ้างอิงผลลัพธ์ของโมเดลพื้นฐานที่แข็งแกร่ง (BERT ที่ปรับปรุงแล้ว) เป็นจุดอ้างอิง และผลลัพธ์เหล่านั้น **ต่ำกว่าของมนุษย์อย่างมีนัยสำคัญ** ในทุกงาน<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup> โดยเฉลี่ยแล้ว โมเดลที่ดีที่สุดในขณะนั้นทำคะแนนได้ **ต่ำกว่าประมาณ 20 คะแนน** เมื่อเทียบกับมนุษย์ตามตัวชี้วัดรวม<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup> ในงานบางอย่าง ช่องว่างนั้นใหญ่เป็นพิเศษ: ตัวอย่างเช่น ในงาน WSC โมเดลแทบจะไปถึง ~65% ของความแม่นยำเมื่อเทียบกับ 100% ของมนุษย์ (ล้าหลัง ~35 คะแนน)<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup> แม้แต่ในงานที่ดูเหมือน "ง่ายกว่า" (BoolQ, CB, RTE, WiC) ระบบอัตโนมัติก็ยังล้าหลังระดับมนุษย์ประมาณ 10 คะแนน<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup> ความแตกต่างเหล่านี้ยืนยันว่า SuperGLUE ท้าทายเทคโนโลยีปัจจุบันอย่างจริงจังและไม่สามารถแก้ไขได้อย่างเรียบง่าย

อย่างไรก็ตาม เพียงไม่กี่เดือนหลังจากการปรากฏตัวของ SuperGLUE ก็เริ่มมี **ความก้าวหน้าอย่างรวดเร็ว**<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup> ในปลายปี 2019 นักวิจัยของ Google ได้นำเสนอโมเดล **T5** (Text-To-Text Transfer Transformer) ที่มี 11 พันล้านพารามิเตอร์ ซึ่งบรรลุผลลัพธ์รวม 88.9 ใกล้เคียงกับระดับมนุษย์ที่ ~89.8<sup>[\[2\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-reddit-t5-2)</sup> ในทางปฏิบัติ T5 ปรับปรุงสถิติก่อนหน้าบน SuperGLUE ถึง 4.3 คะแนนและลดสัดส่วนข้อผิดพลาดลงเกือบหนึ่งในสาม<sup>[\[2\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-reddit-t5-2)</sup> เหลือเพียงช่องว่างขั้นต่ำ **0.9 คะแนน** จากตัวชี้วัดของมนุษย์<sup>[\[2\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-reddit-t5-2)</sup> นักพัฒนาสังเกตว่า SuperGLUE ถูกออกแบบมาโดยเจตนาเพื่อให้งานง่ายสำหรับมนุษย์ ดังนั้นการที่โมเดลไปถึงระดับ ~89% จึงเป็นเหตุการณ์สำคัญ<sup>[\[2\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-reddit-t5-2)</sup>

โมเดลแรกที่ **สามารถเอาชนะคุณภาพเฉลี่ยของมนุษย์** ได้คือโมเดลจาก Microsoft ชื่อ **DeBERTa** (Decoding-enhanced BERT with disentangled attention)<sup>[\[3\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-microsoft-deberta-3)</sup> ในเดือนมกราคม 2021 นักวิจัยรายงานว่า DeBERTa เวอร์ชันที่มี 1.5 พันล้านพารามิเตอร์ทำคะแนนได้ **89.9 คะแนน** ซึ่งสูงกว่าเกณฑ์อ้างอิงของมนุษย์เล็กน้อยที่ 89.8<sup>[\[3\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-microsoft-deberta-3)</sup> นี่เป็น **ครั้งแรก** ที่โมเดลเดี่ยวสามารถเอาชนะมนุษย์ตามตัวชี้วัด SuperGLUE<sup>[\[3\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-microsoft-deberta-3)</sup> นอกจากนี้ ensemble ของโมเดล DeBERTa หลายตัวยังเพิ่มสถิติไปถึง ~90.3 คะแนน<sup>[\[3\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-microsoft-deberta-3)</sup> โมเดล DeBERTa แซงหน้าผู้นำเดิม (Google T5) ประมาณ 0.6% และแสดงให้เห็นถึงประสิทธิภาพของแนวคิดใหม่ในสถาปัตยกรรม Transformer (การแสดงเนื้อหาและตำแหน่งคำแยกกัน decoder mask ที่ปรับปรุงแล้ว ฯลฯ)<sup>[\[4\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-syncedreview-deberta-4)</sup>

ความก้าวหน้าไม่ได้หยุดอยู่แค่นั้น: เมื่อโมเดลภาษาขยายขนาดและความซับซ้อนมากขึ้น ผลลัพธ์บน SuperGLUE ก็ยังคงปรับปรุงต่อไป<sup>[\[5\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-microsoft-scaling-5)</sup> ในช่วงปลายปี 2021 โมเดล **T-NLRv5** ของ Microsoft (ตระกูล Microsoft Turing NLR) ครองอันดับหนึ่งบนลีดเดอร์บอร์ด ซึ่งขยายช่องว่างเหนือระดับมนุษย์ออกไปอีก<sup>[\[5\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-microsoft-scaling-5)</sup> งาน GLUE ที่ยังไม่สามารถแก้ไขได้สำหรับเครื่องจักร (เช่น ความละเอียดอ่อนของ NLI) ถูก "ปิด" โดยโมเดลนี้ ซึ่งใกล้เคียงกับ **ความเท่าเทียมกันอย่างแท้จริงกับมนุษย์** แม้ในงานย่อยที่ยากที่สุด<sup>[\[5\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-microsoft-scaling-5)</sup>

ในช่วงปี 2022-2023 เกณฑ์ระดับมนุษย์บน SuperGLUE ถูกเอาชนะอย่างมั่นใจโดยโมเดลขนาดใหญ่อิสระหลายตัว<sup>[\[6\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-ainavigator-benchmarks-6)</sup> ตัวอย่างเช่น โมเดล **PaLM** จาก Google (540 พันล้านพารามิเตอร์) เมื่อ fine-tuning บนงาน SuperGLUE ทำคะแนนได้ประมาณ 90.4 คะแนน และโมเดล **GPT-4** (พัฒนาโดย OpenAI) ทำคะแนนได้สูงกว่าเล็กน้อย<sup>[\[6\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-ainavigator-benchmarks-6)</sup> ในช่วงกลางปี 2023 ตารางลีดเดอร์บอร์ดของ SuperGLUE มีโมเดลหลายตัวที่ทำคะแนนได้เกิน 90 (กล่าวคือ เกินระดับมนุษย์เฉลี่ย)<sup>[\[6\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-ainavigator-benchmarks-6)</sup> อาจกล่าวได้ว่า benchmark นี้ **ถูกแก้ไขได้จริงในทางปฏิบัติ** โดยระบบสมัยใหม่<sup>[\[6\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-ainavigator-benchmarks-6)</sup>: คะแนนของโมเดลที่ดีที่สุดสูงมากจนเกินความสามารถของคนส่วนใหญ่ที่ไม่มีความเชี่ยวชาญ<sup>[\[6\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-ainavigator-benchmarks-6)</sup> ความสำเร็จนี้แสดงให้เห็นถึงความก้าวหน้าอย่างมหาศาลใน NLP ในระยะเวลาอันสั้น แต่ในขณะเดียวกันก็ชี้ให้เห็นถึงความจำเป็นในการทดสอบใหม่ที่ยากยิ่งขึ้นสำหรับโมเดลล่าสุด<sup>[\[6\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-ainavigator-benchmarks-6)</sup> benchmark ที่ตามมาก็เริ่มปรากฏขึ้นแล้ว (เช่น MMLU, BIG-Bench และอื่น ๆ) ที่มุ่งทดสอบโมเดลในด้านความเข้าใจที่กว้างขึ้นและความรู้รอบด้านที่เกินขอบเขตของงาน SuperGLUE<sup>[\[6\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-ainavigator-benchmarks-6)</sup>

## อิทธิพลและการวิจัยต่อเนื่อง

SuperGLUE จึงได้รับการยืนยันว่าเป็น **ขั้นตอนสำคัญในการพัฒนาวิธีการประเมิน** ในการประมวลผลภาษา<sup>[\[3\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-microsoft-deberta-3)</sup> ในแวดวงผู้ที่ชื่นชอบและวิชาการ ผลลัพธ์ของมันกลายเป็นเหมือน "กระดาษทดสอบ" สำหรับสถาปัตยกรรมใหม่ของ Large Language Models: การบรรลุหรือเกินระดับมนุษย์บน SuperGLUE ถูกมองว่าเป็นสัญญาณของโมเดลขั้นสูงที่มีความเข้าใจภาษาเชิงลึก<sup>[\[3\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-microsoft-deberta-3)</sup> สิ่งนี้สะท้อนออกมาในทางปฏิบัติด้วย — โมเดลภาษาสมัยใหม่หลายตัวที่บรรลุผลลัพธ์สูงบน SuperGLUE ได้กลายเป็นพื้นฐานสำหรับระบบถาม-ตอบ, agent สนทนา, ระบบสรุปข้อความ และอื่น ๆ<sup>[\[3\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-microsoft-deberta-3)</sup> SuperGLUE ยังคงถูกใช้โดยนักวิจัยสำหรับการ fine-tuning และการเปรียบเทียบอัลกอริทึม แม้ว่าตำแหน่งแนวหน้าจะค่อย ๆ เปลี่ยนไปสู่เกณฑ์การประเมิน AI ใหม่

## ลิงก์ภายนอก

- เว็บไซต์อย่างเป็นทางการของ SuperGLUE
- บทความดั้งเดิมของ SuperGLUE (NeurIPS)
- บทความของ Microsoft เกี่ยวกับการบรรลุระดับมนุษย์ของ DeBERTa
- หน้า dataset ของ SuperGLUE บน Papers With Code

## บรรณานุกรม

- Liang, P. et al. (2022). *Holistic Evaluation of Language Models (HELM)*. arXiv:2211.09110.
- Chang, Y. et al. (2023). *A Survey on Evaluation of Large Language Models*. arXiv:2307.03109.
- Ni, S. et al. (2025). *A Survey on Large Language Model Benchmarks*. arXiv:2508.15361.
- Biderman, S. et al. (2024). *The Language Model Evaluation Harness (lm-eval): Guidance and Lessons Learned*. arXiv:2405.14782.
- Kiela, D. et al. (2021). *Dynabench: Rethinking Benchmarking in NLP*. arXiv:2104.14337.
- Ma, Z. et al. (2021). *Dynaboard: An Evaluation‑As‑A‑Service Platform for Holistic Next‑Generation Benchmarking*. arXiv:2106.06052.
- Goel, K. et al. (2021). *Robustness Gym: Unifying the NLP Evaluation Landscape*. arXiv:2101.04840.
- Xu, C. et al. (2024). *Benchmark Data Contamination of Large Language Models: A Survey*. arXiv:2406.04244.
- Liu, S. et al. (2025). *A Comprehensive Survey on Safety Evaluation of LLMs*. arXiv:2506.11094.
- Chiang, W.-L. et al. (2024). *Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference*. arXiv:2403.04132.
- Boubdir, M. et al. (2023). *Elo Uncovered: Robustness and Best Practices in Language Model Evaluation*. arXiv:2311.17295.
- Huang, L. et al. (2023). *A Survey on Hallucination in Large Language Models*. arXiv:2311.05232.

## หมายเหตุ

<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-neurips-main-1)</sup> <sup>[\[2\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-reddit-t5-2)</sup> <sup>[\[3\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-microsoft-deberta-3)</sup> <sup>[\[4\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-syncedreview-deberta-4)</sup> <sup>[\[5\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-microsoft-scaling-5)</sup> <sup>[\[6\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_note-ainavigator-benchmarks-6)</sup> \</references\>

1.  <span id="cite_note-neurips-main-1">↑ <sup>[1.00](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-0)</sup> <sup>[1.01](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-1)</sup> <sup>[1.02](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-2)</sup> <sup>[1.03](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-3)</sup> <sup>[1.04](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-4)</sup> <sup>[1.05](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-5)</sup> <sup>[1.06](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-6)</sup> <sup>[1.07](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-7)</sup> <sup>[1.08](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-8)</sup> <sup>[1.09](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-9)</sup> <sup>[1.10](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-10)</sup> <sup>[1.11](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-11)</sup> <sup>[1.12](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-12)</sup> <sup>[1.13](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-13)</sup> <sup>[1.14](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-14)</sup> <sup>[1.15](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-15)</sup> <sup>[1.16](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-16)</sup> <sup>[1.17](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-17)</sup> <sup>[1.18](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-18)</sup> <sup>[1.19](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-19)</sup> <sup>[1.20](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-20)</sup> <sup>[1.21](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-21)</sup> <sup>[1.22](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-22)</sup> <sup>[1.23](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-23)</sup> <sup>[1.24](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-24)</sup> <sup>[1.25](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-25)</sup> <sup>[1.26](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-26)</sup> <sup>[1.27](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-27)</sup> <sup>[1.28](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-28)</sup> <sup>[1.29](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-29)</sup> <sup>[1.30](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-30)</sup> <sup>[1.31](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-31)</sup> <sup>[1.32](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-32)</sup> <sup>[1.33](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-33)</sup> <sup>[1.34](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-34)</sup> <sup>[1.35](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-35)</sup> <sup>[1.36](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-36)</sup> <sup>[1.37](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-37)</sup> <sup>[1.38](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-38)</sup> <sup>[1.39](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-39)</sup> <sup>[1.40](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-40)</sup> <sup>[1.41](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-41)</sup> <sup>[1.42](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-42)</sup> <sup>[1.43](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-43)</sup> <sup>[1.44](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-44)</sup> <sup>[1.45](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-45)</sup> <sup>[1.46](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-46)</sup> <sup>[1.47](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-47)</sup> <sup>[1.48](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-48)</sup> <sup>[1.49](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-49)</sup> <sup>[1.50](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-50)</sup> <sup>[1.51](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-51)</sup> <sup>[1.52](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-52)</sup> <sup>[1.53](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-53)</sup> <sup>[1.54](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-54)</sup> <sup>[1.55](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-55)</sup> <sup>[1.56](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-56)</sup> <sup>[1.57](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-57)</sup> <sup>[1.58](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-58)</sup> <sup>[1.59](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-neurips-main_1-59)</sup> Wang, Alex et al. (2019). «SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems». *NeurIPS*. <a href="http://papers.neurips.cc/paper/8589-superglue-a-stickier-benchmark-for-general-purpose-language-understanding-systems.pdf" class="external autonumber" rel="nofollow">[1]</a></span>
2.  <span id="cite_note-reddit-t5-2">↑ <sup>[2.0](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-reddit-t5_2-0)</sup> <sup>[2.1](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-reddit-t5_2-1)</sup> <sup>[2.2](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-reddit-t5_2-2)</sup> <sup>[2.3](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-reddit-t5_2-3)</sup> <sup>[2.4](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-reddit-t5_2-4)</sup> «Google T5 algorithm scores 88.9 on SuperGLUE languge benchmark, compared to 89.8 human baseline». *Reddit /r/linguistics*. <a href="https://www.reddit.com/r/linguistics/comments/dmtr38/google_t5_algorithm_scores_889_on_superglue/" class="external autonumber" rel="nofollow">[2]</a></span>
3.  <span id="cite_note-microsoft-deberta-3">↑ <sup>[3.0](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-microsoft-deberta_3-0)</sup> <sup>[3.1](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-microsoft-deberta_3-1)</sup> <sup>[3.2](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-microsoft-deberta_3-2)</sup> <sup>[3.3](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-microsoft-deberta_3-3)</sup> <sup>[3.4](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-microsoft-deberta_3-4)</sup> <sup>[3.5](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-microsoft-deberta_3-5)</sup> <sup>[3.6](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-microsoft-deberta_3-6)</sup> <sup>[3.7](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-microsoft-deberta_3-7)</sup> «Microsoft DeBERTa surpasses human performance on the SuperGLUE benchmark». *Microsoft Research Blog*. <a href="https://www.microsoft.com/en-us/research/blog/microsoft-deberta-surpasses-human-performance-on-the-superglue-benchmark/" class="external autonumber" rel="nofollow">[3]</a></span>
4.  <span id="cite_note-syncedreview-deberta-4">↑ <sup>[4.0](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-syncedreview-deberta_4-0)</sup> <sup>[4.1](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-syncedreview-deberta_4-1)</sup> «Microsoft DeBERTa Tops Human Performance on SuperGLUE NLU Benchmark». *Synced Review*. <a href="https://syncedreview.com/2021/01/06/microsoft-deberta-tops-human-performance-on-superglue-nlu-benchmark/" class="external autonumber" rel="nofollow">[4]</a></span>
5.  <span id="cite_note-microsoft-scaling-5">↑ <sup>[5.0](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-microsoft-scaling_5-0)</sup> <sup>[5.1](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-microsoft-scaling_5-1)</sup> <sup>[5.2](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-microsoft-scaling_5-2)</sup> <sup>[5.3](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-microsoft-scaling_5-3)</sup> «Efficiently and effectively scaling up language model pretraining for best language representation model on GLUE and SuperGLUE». *Microsoft Research Blog*. <a href="https://www.microsoft.com/en-us/research/blog/efficiently-and-effectively-scaling-up-language-model-pretraining-for-best-language-representation-model-on-glue-and-superglue/" class="external autonumber" rel="nofollow">[5]</a></span>
6.  <span id="cite_note-ainavigator-benchmarks-6">↑ <sup>[6.0](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-ainavigator-benchmarks_6-0)</sup> <sup>[6.1](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-ainavigator-benchmarks_6-1)</sup> <sup>[6.2](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-ainavigator-benchmarks_6-2)</sup> <sup>[6.3](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-ainavigator-benchmarks_6-3)</sup> <sup>[6.4](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-ainavigator-benchmarks_6-4)</sup> <sup>[6.5](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-ainavigator-benchmarks_6-5)</sup> <sup>[6.6](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-ainavigator-benchmarks_6-6)</sup> <sup>[6.7](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(TH)#cite_ref-ainavigator-benchmarks_6-7)</sup> «The Ultimate Guide to AI Benchmarks». *The AI Navigator*. <a href="https://www.theainavigator.com/blog/the-ultimate-guide-to-ai-benchmarks/" class="external autonumber" rel="nofollow">[6]</a></span>
