---
title: "SuperGLUE (benchmark) (KO)"
source: "https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)"
wiki: "systems-analysis.info/int"
article: "SuperGLUE_(benchmark)_(KO)"
language: "ko"
categories:
  - "Category:Korean"
  - "Category:Large language models"
  - "Category:LLM benchmarks"
  - "Category:Machine learning"
revision_id: 7010
wiki_created_at: 2026-09-07T00:14:32Z
wiki_modified_at: 2026-09-07T00:14:32Z
downloaded_at: 2026-09-07T23:16:54Z
---

# SuperGLUE (benchmark) (KO)

**SuperGLUE** — 자연어 처리 시스템, 특히 **대형 언어 모델**(LLM)을 평가하기 위한 종합적인 **benchmark**(테스트 과제 모음)입니다<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-neurips-main-1)</sup>. 2019년 뉴욕대학교의 Alex Wang을 중심으로 한 연구팀이 Facebook AI Research 및 기타 기관의 참여 하에 발표하였습니다<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-neurips-main-1)</sup>.

    SuperGLUE가 만들어진 배경에는 2019년 중반에 이르러 기존 benchmark인 GLUE가 당시 최신 모델들에게 '쉬운 과제'가 되어버렸다는 사실이 있습니다. GLUE에서 최고 모델들의 종합 점수가 88.4에 달해 인간 평균 수준(87.1)을 초과하게 되었고[1], 이로 인해 추가적인 발전의 여지가 줄어들었습니다[1]. 이에 대응하여 연구자들은 언어 이해 능력을 더욱 엄격하게 검증할 수 있는 더 어려운 대안으로 SuperGLUE를 개발하였습니다[1]. SuperGLUE의 목표는 영어의 일반적 언어 이해 분야에서의 발전을 측정하는 중립적이고 '학습하기 어려운' 지표를 제공하는 것입니다[1]. SuperGLUE에서 눈에 띄는 성능 향상을 이루려면 Machine Learning 방법론에서의 실질적인 혁신—예를 들어 소수 샘플에서의 효율적인 학습, 멀티태스크 학습, 자기지도 학습 등—이 필요할 것으로 기대되었습니다[1]. 즉, SuperGLUE에는 인간에게는 쉽지만 기계 지능에게는 어려운 과제들이 포함되어[1] 진정한 깊이 있는 언어 이해를 갖춘 모델 개발을 촉진하고자 하였습니다.

## GLUE와의 특징 및 차이점

SuperGLUE는 GLUE의 형식을 대체로 계승하여 여러 과제를 종합한 단일 **통합 품질 지표**, 공개 **리더보드**, 그리고 모델 분석을 위한 **도구 모음**을 제공합니다<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-neurips-main-1)</sup>. 그러나 SuperGLUE는 전작과 비교하여 여러 개선 사항과 새로운 요소를 도입하였습니다<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-neurips-main-1)</sup>:

- **더 어려운 과제들**: SuperGLUE에는 **가장 어려운 8가지 과제**가 선별되었습니다<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-neurips-main-1)</sup>. 그 중 두 가지는 GLUE에서 가장 어려웠던 과제를 계승하였으며, 나머지는 현대 NLP 모델들에 대한 난이도를 기준으로 새로운 후보 중에서 선발되었습니다<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-neurips-main-1)</sup>. 이처럼 benchmark는 모델들이 과거에 가장 낮은 성능을 보였던 이해의 측면에 집중합니다.
- **형식의 다양성**: GLUE의 모든 과제가 문장 또는 문장 쌍 분류로 귀결되었다면, SuperGLUE는 더 **넓은 범위의 형식**을 포함합니다<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-neurips-main-1)</sup>. 분류 외에도 **상호참조 해소**와 **질의응답** 과제가 추가되어, 모델이 일관된 텍스트를 이해하고 **논리적 추론**을 수행할 것을 요구합니다<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-neurips-main-1)</sup>.
- **모든 과제에서의 인간 평가**: SuperGLUE의 각 과제에 대해 비전문가 **인간 기준 성능**이 산출되었으며<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-neurips-main-1)</sup>, benchmark 출시 당시 BERT와 같은 강력한 모델조차 인간보다 크게 뒤처졌음이 확인되었습니다<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-neurips-main-1)</sup>. 종합 약 90%에 해당하는 **인간 기준점**의 존재는 모델 발전의 '여지'를 보장하고 목표 지향점으로 기능합니다<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-neurips-main-1)</sup>.
- **투명한 규칙과 도구**: 리더보드 결과 게시 규칙이 공정한 비교와 dataset 기여자 명시를 위해 개정되었습니다<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-neurips-main-1)</sup>. 또한 SuperGLUE 데이터에 대한 모델의 fine-tuning 및 멀티태스크 학습을 용이하게 하기 위한 새로운 오픈소스 도구 모음이 공개되었습니다<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-neurips-main-1)</sup>.

이러한 조치들을 종합하면, SuperGLUE는 좁은 방식의 반칙이나 기존 GLUE의 특정 형식에 대한 과적합으로는 높은 점수를 얻을 수 없도록 함으로써 모델의 **일반화된 언어 능력**을 보다 신뢰성 있게 검증하는 테스트가 됩니다<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-neurips-main-1)</sup>.

## SuperGLUE 과제 구성

SuperGLUE는 텍스트 이해의 다양한 측면을 포괄하는 **8가지 과제**로 구성됩니다.

- **BoolQ** (Boolean Questions): **질의응답(QA)** 유형의 과제로, 각 예시에는 짧은 텍스트(위키백과 발췌)와 '예' 또는 '아니오'로 답해야 하는 질문이 주어집니다<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-neurips-main-1)</sup>. 질문은 사용자(Google 검색 쿼리)가 작성한 것으로, 텍스트에서 명시적 또는 암묵적 사실을 추출할 것을 요구하며 평가 지표는 정확도(accuracy)입니다<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-neurips-main-1)</sup>.
- **CB** (CommitmentBank): 세 가지 클래스를 가진 **논리적 함의**(textual entailment) 과제입니다<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-neurips-main-1)</sup>. dataset은 복합문을 포함하는 짧은 텍스트로 구성되며, 텍스트 저자가 내포된 명제의 진실성에 얼마나 **확신(committed)**을 가지는지 판단해야 합니다<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-neurips-main-1)</sup>. 실질적으로는 주어진 문맥에서 주장이 도출되는지를 검증합니다. 소규모 샘플(약 250개 예시)과 클래스 불균형으로 인해 과제가 어려우며, 클래스별 평균 정밀도와 F1 점수로 평가됩니다<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-neurips-main-1)</sup>.
- **COPA** (Choice of Plausible Alternatives): **인과 추론** 과제입니다<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-neurips-main-1)</sup>. 모델에게 전제(한 문장)가 주어지고 두 가지 선택지 중 올바른 원인 또는 결과를 선택해야 합니다<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-neurips-main-1)</sup>. COPA의 모든 예시는 수작업으로 작성되었으며 인과관계를 파악하기 위해 **상식**이 필요합니다. 주제는 블로그와 전문 백과사전의 상황을 포함하며, 평가 지표는 정확도(올바른 선택 비율)입니다<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-neurips-main-1)</sup>. 예시: '아이가 질병에 대한 면역을 얻었다'는 문장과 '원인은 무엇인가?'라는 질문이 주어지면, 인간은 즉시 '그는 백신을 맞았다'가 정답임을 이해하지만 모델은 인과관계를 추론해야 합니다<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-neurips-main-1)</sup>.
- **MultiRC** (Multi-Sentence Reading Comprehension): 다지선다형 요소를 포함하는 **다문장 텍스트 이해** 과제입니다<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-neurips-main-1)</sup>. 모델은 텍스트 단락, 내용에 관한 질문, 가능한 답변 목록을 받으며 어떤 답변이 옳은지 판단해야 합니다(각 질문에 여러 정답이 있을 수 있음)<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-neurips-main-1)</sup>. 특징: 질문에 답하기 위해 일반적으로 텍스트의 여러 문장에서 정보를 결합해야 하므로 모델의 **사실 연결 능력**을 검증합니다<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-neurips-main-1)</sup>. 품질은 두 가지 지표로 측정됩니다: 답변에 대한 F1(부분적으로 정확한 집합을 고려)과 Exact Match(완전히 정확한 답변 집합을 제공한 질문의 비율)<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-neurips-main-1)</sup>.
- **ReCoRD** (Reading Comprehension with Commonsense Reasoning Dataset): **지식을 활용한 독해 이해** 과제입니다<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-neurips-main-1)</sup>. 변형된 Cloze 테스트 형태로, 뉴스 텍스트(CNN/Daily Mail 기사)와 개체명이 빠진 문장이 주어지면 모델은 텍스트에서 어떤 개체가 빈칸에 적합한지 선택해야 합니다<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-neurips-main-1)</sup>. 답변 후보는 기사에 언급된 모든 개체로 설정되며, 의미상 중복될 수 있습니다<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-neurips-main-1)</sup>. 성공적인 해결을 위해 문맥 이해와 상식이 필요합니다. 평가 지표는 최대 token-level F1과 Exact Match(예측 답변의 정확한 일치)입니다<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-neurips-main-1)</sup>.
- **RTE** (Recognizing Textual Entailment): **텍스트 함의**(entailment vs. not entailment)에 대한 이진 분류 과제입니다<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-neurips-main-1)</sup>. dataset은 여러 텍스트 함의 인식 대회(RTE 1-5 시리즈)의 예시를 통합한 것입니다<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-neurips-main-1)</sup>. 각 과제는 텍스트 단편 쌍(전제-가설)을 포함하며, 모델은 가설이 텍스트에서 도출되는지 판단해야 합니다. 많은 대규모 dataset과 달리 RTE는 상당히 소규모(약 2,500개 학습 예시)이지만, transfer 학습으로 큰 이득을 얻었습니다. BERT 유형의 모델 등장으로 정확도가 ~56%(무작위 추측 수준)에서 ~86%로 향상되었습니다<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-neurips-main-1)</sup>. 그럼에도 SuperGLUE 출시 당시 모델 정확도는 인간 수준보다 약 8 퍼센트 포인트 뒤처졌으며<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-neurips-main-1)</sup>, 이로 인해 RTE는 인간 수준까지 격차가 남아있는 과제 중 하나로 포함되었습니다.
- **WiC** (Word-in-Context): **문맥 속 단어 의미 중의성 해소**(WSD) 과제입니다<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-neurips-main-1)</sup>. 각각 동일한 다의어가 등장하는 두 개의 독립 문장이 주어지며, 해당 단어가 두 경우 모두 **같은 의미로** 사용되었는지 판단해야 합니다<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-neurips-main-1)</sup>. 데이터는 사전 리소스(WordNet, VerbNet, Wiktionary)에서 가져와 광범위한 단어와 의미를 포괄합니다<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-neurips-main-1)</sup>. 이진 분류로 형식화되어 정답률로 평가됩니다. WiC는 모델에게 미묘한 의미 차이의 이해, 즉 실질적으로 **어휘 의미론**을 검증할 것을 요구합니다.
- **WSC** (Winograd Schema Challenge): **상식을 활용한 상호참조 해소** 과제입니다<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-neurips-main-1)</sup>. 각 과제는 대명사를 포함하는 하나의 문장과 같은 문장의 두 개체(명사)로 구성됩니다<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-neurips-main-1)</sup>. **해당 대명사가 어떤 명사를 지칭하는지** 판단해야 합니다<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-neurips-main-1)</sup>. 고전적인 winograd 문장의 예: '트로피가 여행 가방에 들어가지 않았다, 왜냐하면 그것이 너무 작았기 때문이다'—인간은 '그것'이 여행 가방을 지칭한다는 것(너무 작은 것은 여행 가방이었다)을 이해합니다. 이러한 예시는 **일상적 지식과 문맥** 없이는 해결할 수 없습니다<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-neurips-main-1)</sup>. GLUE에도 이 과제의 단순화된 버전(WNLI)이 있었지만, 모델들은 오랫동안 무작위 수준조차 넘어서지 못했습니다<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-neurips-main-1)</sup>. 유사한 예시가 담긴 외부 데이터 추가 등의 특수한 기법을 통해서야 2019년까지 WSC에서 모델 성능이 ~90%까지 향상되었습니다<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-neurips-main-1)</sup>. 그러나 인간은 WSC 과제를 거의 실수 없이 해결합니다(~96-100% 정답률)<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-neurips-main-1)</sup>. SuperGLUE에는 이진 분류 형식의 원본 WSC 버전이 포함되어(각 '대명사-개체' 쌍에 대해 모델이 동일한 지시 대상인지 응답), 상식적 추론을 요구하는 가장 어려운 테스트 중 하나로 남아 있습니다<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-neurips-main-1)</sup>.

SuperGLUE의 모든 테스트는 개발자에게 답변이 공개되지 않는 **비공개 테스트 세트**를 가집니다<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-neurips-main-1)</sup>. 모델은 예측 결과를 서버에 제출하며, 서버에서는 과제별로 평균화된 종합 점수가 산출됩니다(여러 지표를 가진 과제의 경우 내부 지표를 먼저 평균화)<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-neurips-main-1)</sup>. 이러한 단일 **SuperGLUE score**는 전반적인 언어 지능 수준에서 모델을 비교하는 것을 단순화합니다.

## 모델의 결과와 발전

SuperGLUE 출시 당시 저자들은 강력한 기준 모델(강화된 BERT)의 결과를 참고치로 제시하였으며, 이는 모든 과제에서 **인간 수준보다 크게 낮았습니다**<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-neurips-main-1)</sup>. 평균적으로 당시 최고 모델은 통합 지표에서 인간보다 약 **20점 낮은** 점수를 기록하였습니다<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-neurips-main-1)</sup>. 개별 과제에서 격차는 특히 컸습니다. 예를 들어 WSC 과제에서 모델은 인간의 100%에 비해 ~65% 정확도에 겨우 도달하는 수준이었습니다(약 35점 차이)<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-neurips-main-1)</sup>. '상대적으로 쉬운' 과제(BoolQ, CB, RTE, WiC)에서조차 자동화 시스템은 인간 수준보다 ~10점 뒤처졌습니다<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-neurips-main-1)</sup>. 이러한 차이는 SuperGLUE가 실제로 현재 기술에 심각한 도전을 제기하며, 단순한 방법으로는 해결될 수 없음을 확인해주었습니다.

그럼에도 불구하고 SuperGLUE 등장 불과 몇 달 후부터 **빠른 발전**이 시작되었습니다<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-neurips-main-1)</sup>. 2019년 말 Google 연구팀은 110억 개의 매개변수를 가진 **T5** (Text-To-Text Transfer Transformer) 모델을 발표하였으며, 이 모델은 종합 점수 88.9를 달성하여 인간 수준인 ~89.8에 근접하였습니다<sup>[\[2\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-reddit-t5-2)</sup>. 사실상 T5는 SuperGLUE의 이전 기록을 단번에 4.3점 향상시키고 오류율을 거의 3분의 1 가량 줄였으며<sup>[\[2\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-reddit-t5-2)</sup>, 인간 점수까지의 격차를 겨우 **0.9점**으로 좁혔습니다<sup>[\[2\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-reddit-t5-2)</sup>. 개발자들은 SuperGLUE가 인간에게는 쉬운 과제들로 의도적으로 구성되어 있기 때문에, 모델이 ~89% 수준에 도달한 것이 중요한 이정표였다고 밝혔습니다<sup>[\[2\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-reddit-t5-2)</sup>.

**평균적인 인간 수준을 최초로 초과**한 모델은 Microsoft의 **DeBERTa** (Decoding-enhanced BERT with disentangled attention)였습니다<sup>[\[3\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-microsoft-deberta-3)</sup>. 2021년 1월 연구팀은 15억 매개변수 버전의 DeBERTa가 인간 기준 89.8보다 약간 높은 **89.9점**을 기록했다고 발표하였습니다<sup>[\[3\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-microsoft-deberta-3)</sup>. 이것은 단일 모델이 SuperGLUE 지표에서 인간을 초월한 **최초의 사례**였습니다<sup>[\[3\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-microsoft-deberta-3)</sup>. 여기에 더해 여러 DeBERTa 모델의 앙상블은 기록을 ~90.3점으로 끌어올렸습니다<sup>[\[3\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-microsoft-deberta-3)</sup>. DeBERTa 모델은 이전 선두 모델(Google T5)을 약 0.6% 앞섰으며, Transformer 아키텍처의 새로운 아이디어(콘텐츠와 단어 위치의 분리 표현, 개선된 마스크 디코더 등)의 효과를 입증하였습니다<sup>[\[4\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-syncedreview-deberta-4)</sup>.

발전은 거기서 멈추지 않았습니다. 언어 모델의 규모와 복잡성이 증가함에 따라 SuperGLUE 점수는 계속 향상되었습니다<sup>[\[5\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-microsoft-scaling-5)</sup>. 2021년 말에는 Microsoft의 **T-NLRv5** (Microsoft Turing NLR 계열) 모델이 리더보드 정상에 올라 인간 수준을 더욱 크게 초과하였습니다<sup>[\[5\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-microsoft-scaling-5)</sup>. GLUE에서 기계에 마지막까지 남아있던 미해결 과제들(예: NLI의 미묘한 점들)은 이 모델에 의해 '종결'되어 가장 어려운 하위 과제에서도 **인간과의 완전한 동등 수준**에 근접하였습니다<sup>[\[5\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-microsoft-scaling-5)</sup>.

2022-2023년에 이르러 SuperGLUE의 인간 수준 임계값은 여러 독립적인 대형 모델들에 의해 확실히 초월되었습니다<sup>[\[6\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-ainavigator-benchmarks-6)</sup>. 예를 들어 Google의 **PaLM** 모델(5,400억 매개변수)은 SuperGLUE 과제에 fine-tuning 후 약 90.4점에 도달하였으며, OpenAI가 개발한 **GPT-4** 모델은 그보다 약간 더 높은 결과를 보였습니다<sup>[\[6\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-ainavigator-benchmarks-6)</sup>. 2023년 중반까지 SuperGLUE 리더보드에는 90점(즉 평균 인간 수준 초과) 이상의 점수를 가진 모델이 여럿 등재되었습니다<sup>[\[6\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-ainavigator-benchmarks-6)</sup>. benchmark가 현대 시스템에 의해 **사실상 해결되었다**고 말할 수 있으며<sup>[\[6\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-ainavigator-benchmarks-6)</sup>, 최고 모델들의 점수는 대부분의 비전문가 인간의 능력을 초월할 만큼 높습니다<sup>[\[6\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-ainavigator-benchmarks-6)</sup>. 이러한 성과는 NLP 분야에서 짧은 기간에 이루어진 엄청난 발전을 보여주지만, 동시에 최신 모델들을 위한 더 새롭고 어려운 테스트의 필요성을 시사합니다<sup>[\[6\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-ainavigator-benchmarks-6)</sup>. 이미 SuperGLUE의 범위를 넘어서는 더 광범위한 이해와 지식을 모델에서 검증하기 위한 후속 benchmark들(예: MMLU, BIG-Bench 등)이 등장하고 있습니다<sup>[\[6\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-ainavigator-benchmarks-6)</sup>.

## 영향과 후속 연구

SuperGLUE는 이로써 언어 처리 분야에서 **평가 방법론 발전의 중요한 이정표**로 자리매김하였습니다<sup>[\[3\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-microsoft-deberta-3)</sup>. 열정적인 연구자들과 학술 커뮤니티에서 그 결과는 새로운 LLM 아키텍처의 일종의 '리트머스 시험지'가 되었습니다. SuperGLUE에서 인간 수준에 도달하거나 초과하는 것은 깊은 언어 이해를 갖춘 첨단 모델의 징표로 받아들여집니다<sup>[\[3\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-microsoft-deberta-3)</sup>. 이는 실무에도 반영되어, SuperGLUE에서 높은 점수를 달성한 많은 현대 언어 모델들이 응용 질의응답 시스템, 대화 에이전트, 텍스트 요약 시스템 등의 기반이 되었습니다<sup>[\[3\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-microsoft-deberta-3)</sup>. SuperGLUE는 알고리즘의 fine-tuning과 비교를 위해 연구자들이 계속 활용하고 있지만, 최전선은 이제 점차 인공지능 평가의 새로운 지평으로 이동하고 있습니다.

## 외부 링크

- SuperGLUE 공식 웹사이트
- SuperGLUE 원본 논문 (NeurIPS)
- DeBERTa의 인간 수준 달성에 관한 Microsoft 논문
- Papers With Code의 SuperGLUE dataset 페이지

## 참고 문헌

- Liang, P. et al. (2022). *Holistic Evaluation of Language Models (HELM)*. arXiv:2211.09110.
- Chang, Y. et al. (2023). *A Survey on Evaluation of Large Language Models*. arXiv:2307.03109.
- Ni, S. et al. (2025). *A Survey on Large Language Model Benchmarks*. arXiv:2508.15361.
- Biderman, S. et al. (2024). *The Language Model Evaluation Harness (lm-eval): Guidance and Lessons Learned*. arXiv:2405.14782.
- Kiela, D. et al. (2021). *Dynabench: Rethinking Benchmarking in NLP*. arXiv:2104.14337.
- Ma, Z. et al. (2021). *Dynaboard: An Evaluation‑As‑A‑Service Platform for Holistic Next‑Generation Benchmarking*. arXiv:2106.06052.
- Goel, K. et al. (2021). *Robustness Gym: Unifying the NLP Evaluation Landscape*. arXiv:2101.04840.
- Xu, C. et al. (2024). *Benchmark Data Contamination of Large Language Models: A Survey*. arXiv:2406.04244.
- Liu, S. et al. (2025). *A Comprehensive Survey on Safety Evaluation of LLMs*. arXiv:2506.11094.
- Chiang, W.-L. et al. (2024). *Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference*. arXiv:2403.04132.
- Boubdir, M. et al. (2023). *Elo Uncovered: Robustness and Best Practices in Language Model Evaluation*. arXiv:2311.17295.
- Huang, L. et al. (2023). *A Survey on Hallucination in Large Language Models*. arXiv:2311.05232.

## 각주

<sup>[\[1\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-neurips-main-1)</sup> <sup>[\[2\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-reddit-t5-2)</sup> <sup>[\[3\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-microsoft-deberta-3)</sup> <sup>[\[4\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-syncedreview-deberta-4)</sup> <sup>[\[5\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-microsoft-scaling-5)</sup> <sup>[\[6\]](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_note-ainavigator-benchmarks-6)</sup> \</references\>

1.  <span id="cite_note-neurips-main-1">↑ <sup>[1.00](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-0)</sup> <sup>[1.01](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-1)</sup> <sup>[1.02](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-2)</sup> <sup>[1.03](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-3)</sup> <sup>[1.04](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-4)</sup> <sup>[1.05](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-5)</sup> <sup>[1.06](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-6)</sup> <sup>[1.07](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-7)</sup> <sup>[1.08](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-8)</sup> <sup>[1.09](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-9)</sup> <sup>[1.10](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-10)</sup> <sup>[1.11](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-11)</sup> <sup>[1.12](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-12)</sup> <sup>[1.13](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-13)</sup> <sup>[1.14](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-14)</sup> <sup>[1.15](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-15)</sup> <sup>[1.16](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-16)</sup> <sup>[1.17](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-17)</sup> <sup>[1.18](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-18)</sup> <sup>[1.19](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-19)</sup> <sup>[1.20](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-20)</sup> <sup>[1.21](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-21)</sup> <sup>[1.22](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-22)</sup> <sup>[1.23](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-23)</sup> <sup>[1.24](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-24)</sup> <sup>[1.25](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-25)</sup> <sup>[1.26](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-26)</sup> <sup>[1.27](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-27)</sup> <sup>[1.28](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-28)</sup> <sup>[1.29](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-29)</sup> <sup>[1.30](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-30)</sup> <sup>[1.31](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-31)</sup> <sup>[1.32](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-32)</sup> <sup>[1.33](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-33)</sup> <sup>[1.34](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-34)</sup> <sup>[1.35](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-35)</sup> <sup>[1.36](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-36)</sup> <sup>[1.37](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-37)</sup> <sup>[1.38](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-38)</sup> <sup>[1.39](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-39)</sup> <sup>[1.40](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-40)</sup> <sup>[1.41](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-41)</sup> <sup>[1.42](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-42)</sup> <sup>[1.43](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-43)</sup> <sup>[1.44](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-44)</sup> <sup>[1.45](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-45)</sup> <sup>[1.46](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-46)</sup> <sup>[1.47](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-47)</sup> <sup>[1.48](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-48)</sup> <sup>[1.49](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-49)</sup> <sup>[1.50](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-50)</sup> <sup>[1.51](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-51)</sup> <sup>[1.52](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-52)</sup> <sup>[1.53](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-53)</sup> <sup>[1.54](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-54)</sup> <sup>[1.55](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-55)</sup> <sup>[1.56](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-56)</sup> <sup>[1.57](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-57)</sup> <sup>[1.58](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-58)</sup> <sup>[1.59](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-neurips-main_1-59)</sup> Wang, Alex et al. (2019). «SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems». *NeurIPS*. <a href="http://papers.neurips.cc/paper/8589-superglue-a-stickier-benchmark-for-general-purpose-language-understanding-systems.pdf" class="external autonumber" rel="nofollow">[1]</a></span>
2.  <span id="cite_note-reddit-t5-2">↑ <sup>[2.0](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-reddit-t5_2-0)</sup> <sup>[2.1](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-reddit-t5_2-1)</sup> <sup>[2.2](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-reddit-t5_2-2)</sup> <sup>[2.3](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-reddit-t5_2-3)</sup> <sup>[2.4](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-reddit-t5_2-4)</sup> «Google T5 algorithm scores 88.9 on SuperGLUE languge benchmark, compared to 89.8 human baseline». *Reddit /r/linguistics*. <a href="https://www.reddit.com/r/linguistics/comments/dmtr38/google_t5_algorithm_scores_889_on_superglue/" class="external autonumber" rel="nofollow">[2]</a></span>
3.  <span id="cite_note-microsoft-deberta-3">↑ <sup>[3.0](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-microsoft-deberta_3-0)</sup> <sup>[3.1](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-microsoft-deberta_3-1)</sup> <sup>[3.2](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-microsoft-deberta_3-2)</sup> <sup>[3.3](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-microsoft-deberta_3-3)</sup> <sup>[3.4](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-microsoft-deberta_3-4)</sup> <sup>[3.5](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-microsoft-deberta_3-5)</sup> <sup>[3.6](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-microsoft-deberta_3-6)</sup> <sup>[3.7](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-microsoft-deberta_3-7)</sup> «Microsoft DeBERTa surpasses human performance on the SuperGLUE benchmark». *Microsoft Research Blog*. <a href="https://www.microsoft.com/en-us/research/blog/microsoft-deberta-surpasses-human-performance-on-the-superglue-benchmark/" class="external autonumber" rel="nofollow">[3]</a></span>
4.  <span id="cite_note-syncedreview-deberta-4">↑ <sup>[4.0](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-syncedreview-deberta_4-0)</sup> <sup>[4.1](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-syncedreview-deberta_4-1)</sup> «Microsoft DeBERTa Tops Human Performance on SuperGLUE NLU Benchmark». *Synced Review*. <a href="https://syncedreview.com/2021/01/06/microsoft-deberta-tops-human-performance-on-superglue-nlu-benchmark/" class="external autonumber" rel="nofollow">[4]</a></span>
5.  <span id="cite_note-microsoft-scaling-5">↑ <sup>[5.0](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-microsoft-scaling_5-0)</sup> <sup>[5.1](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-microsoft-scaling_5-1)</sup> <sup>[5.2](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-microsoft-scaling_5-2)</sup> <sup>[5.3](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-microsoft-scaling_5-3)</sup> «Efficiently and effectively scaling up language model pretraining for best language representation model on GLUE and SuperGLUE». *Microsoft Research Blog*. <a href="https://www.microsoft.com/en-us/research/blog/efficiently-and-effectively-scaling-up-language-model-pretraining-for-best-language-representation-model-on-glue-and-superglue/" class="external autonumber" rel="nofollow">[5]</a></span>
6.  <span id="cite_note-ainavigator-benchmarks-6">↑ <sup>[6.0](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-ainavigator-benchmarks_6-0)</sup> <sup>[6.1](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-ainavigator-benchmarks_6-1)</sup> <sup>[6.2](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-ainavigator-benchmarks_6-2)</sup> <sup>[6.3](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-ainavigator-benchmarks_6-3)</sup> <sup>[6.4](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-ainavigator-benchmarks_6-4)</sup> <sup>[6.5](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-ainavigator-benchmarks_6-5)</sup> <sup>[6.6](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-ainavigator-benchmarks_6-6)</sup> <sup>[6.7](https://systems-analysis.info/int/SuperGLUE_(benchmark)_(KO)#cite_ref-ainavigator-benchmarks_6-7)</sup> «The Ultimate Guide to AI Benchmarks». *The AI Navigator*. <a href="https://www.theainavigator.com/blog/the-ultimate-guide-to-ai-benchmarks/" class="external autonumber" rel="nofollow">[6]</a></span>
