---
title: "SafetyBench (KO)"
source: "https://systems-analysis.info/int/SafetyBench_(KO)"
wiki: "systems-analysis.info/int"
article: "SafetyBench_(KO)"
language: "ko"
categories:
  - "Category:Korean"
  - "Category:Large language models"
  - "Category:LLM benchmarks"
  - "Category:Machine learning"
revision_id: 6559
wiki_created_at: 2026-09-07T00:06:38Z
wiki_modified_at: 2026-09-07T00:06:38Z
downloaded_at: 2026-09-07T23:14:34Z
---

# SafetyBench (KO)

**SafetyBench** — 대형 언어 모델의 **안전성을 종합적으로 평가하기 위한 최초의 포괄적 benchmark**입니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 이 benchmark는 칭화대학교 연구팀이 개발하여 2023년에 발표되었습니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>.

대형 언어 모델(이하 LLM)의 발전(예: ChatGPT의 등장)과 대규모 보급이 이루어지면서 이러한 시스템의 안전성 문제에 대한 관심이 높아졌습니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 연구 결과에 따르면 대화형 모델이 사용자의 개인 정보를 유출하거나 유해한 발언을 생성할 수 있는 것으로 나타났습니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 따라서 LLM의 안전성 평가는 실제 환경에서 신뢰할 수 있는 활용을 위해 매우 중요한 과제가 되었습니다. 그러나 최근까지 모델 안전성의 모든 주요 측면을 포괄하는 종합적인 benchmark(테스트 세트)가 존재하지 않았으며, 기존 dataset은 독성이나 사회적 편견과 같은 개별적인 측면만을 검증할 뿐 전체적인 그림을 제공하지 못했습니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 이처럼 포괄적인 평가 방법이 부재하여 취약점 발견과 더 안전한 언어 모델 개발 모두가 어려웠습니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. SafetyBench는 이러한 공백을 메우기 위해 만들어졌습니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>.

## SafetyBench의 개발 및 설명

SafetyBench는 AI가 생성하는 콘텐츠와 관련된 일반적인 문제 또는 위협의 **7가지 범주**를 포괄하는 **11,435개의 객관식 문항**(multiple-choice)으로 구성되어 있습니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 중요한 특징은 **이중 언어성**입니다. 각 문항은 **영어**와 **중국어** 두 가지 언어로 제공되어, 영어 모델과 중국어 모델을 동일한 자료로 평가할 수 있습니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 본질적으로 SafetyBench는 모델이 안전한 행동과 콘텐츠에 관한 문제를 얼마나 잘 이해하고 올바른 답변을 제공하는지 자동으로 정확하게 테스트할 수 있는 최초의 대규모 도구가 되었습니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. MMLU와 같은 잘 알려진 benchmark와 유사한 단일 정답 형식은 평가의 객관성과 효율성을 보장하며, 모델 응답에 대한 수작업 검토 의존도를 낮춥니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>.

SafetyBench 개발자들은 안전하지 않은 콘텐츠와 관련된 일반적인 시나리오에 대해 기존에 제안된 분류 체계를 기반으로 삼았습니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 특히 benchmark의 범주는 Sun 외 공저자(2023)의 연구에서 설명된 8가지 시나리오를 기반으로 도출되었으나, 중국어와 영어 맥락 간의 비교 불가 문제를 피하기 위해 하나의 범주(정치적으로 민감한 주제)가 제외되었습니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 따라서 최종 세트에는 두 언어에 공통된 7가지 안전 범주가 포함됩니다.

## SafetyBench의 안전 범주

SafetyBench의 각 테스트 문항은 잠재적으로 위험하거나 바람직하지 않은 다양한 측면을 포괄하는 7가지 범주 중 하나에 속합니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 아래에 각 범주와 간략한 설명을 제시합니다.

- **공격적 콘텐츠** (Offensiveness) – 위협, 모욕, 무례함, 욕설, 비꼼 및 기타 허용 불가한 어조의 표현<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 모델은 이러한 공격적 표현을 인식하고 독성 또는 공격적인 콘텐츠에 저항할 수 있어야 합니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>.
- **편견 및 차별** (Unfairness and Bias) – 인종, 성별, 종교 등의 기준에 따른 사회적 bias 및 불공정 표현<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 모델은 편견이나 차별을 나타내는 언어 구조를 식별하고 피할 수 있어야 합니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>.
- **신체 건강** (Physical Health) – 사람의 신체 건강에 영향을 미칠 수 있는 상황 및 발언<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 모델은 다양한 삶의 상황에서 건강을 유지하기 위한 올바르고 안전한 행동과 조언을 알고 있어야 합니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>.
- **정신 건강** (Mental Health) – 심리적 안녕, 감정 및 정신 건강과 관련된 문제<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 모델은 정신 건강을 유지하고 부정적인 감정적 영향을 예방하기 위한 올바른 방법을 제시할 수 있어야 합니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>.
- **불법 활동** (Illegal Activities) – 불법 행위를 전제로 하는 시나리오<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 모델은 합법적 행동과 불법적 행동을 구별하고, 법률 규범에 대한 기본 지식을 갖추고 있으며, 법 위반을 조장하지 않아야 합니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>.
- **윤리 및 도덕** (Ethics and Morality) – 법에 직접적으로 저촉되지 않더라도 비윤리적 또는 비도덕적 행동과 관련된 상황<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 모델은 높은 윤리적 기준을 보여주고 비윤리적 행동이나 발언을 비판할 수 있어야 합니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>.
- **개인정보 및 재산** (Privacy and Property) – 개인 정보, 재산권, 금융 위험 등에 관한 문제<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 모델은 개인정보 보호 및 재산권 원칙을 민감하게 이해하고, 개인 정보의 비의도적 노출이나 재산 피해를 방지할 수 있어야 합니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>.

각 범주는 수백에서 수천 개의 문항으로 구성되어 있어, 관련 규범 및 원칙에 대한 모델의 지식을 종합적으로 검증할 수 있습니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>.

## 데이터 수집 및 준비

이처럼 대규모의 테스트 세트를 구성하기 위해 SafetyBench 저자들은 **다양한 데이터 소스**를 활용했습니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 연구에서는 세 가지 주요 소스에서 문항이 수집되었음을 명시하고 있습니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>.

- **기존 dataset**: 일부 범주(특히 공격성, 편견, 신체 건강, 윤리)에 대해서는 공개적으로 이용 가능한 dataset이 활용되었습니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 저자들은 이러한 세트의 원본 텍스트를 가져와 객관식 문항 형식으로 변환했습니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 예를 들어 Offensiveness 범주에는 부분적으로 COLD 코퍼스(중국어 공격성 감지 dataset)가 활용되었으며<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>, 영어의 경우 Jigsaw Toxic Comment 대회 데이터 등이 사용되었습니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 마찬가지로 Unfairness and Bias에는 중국어 세트(COLD, CDial-Bias)와 영어 자료가 활용되었습니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 이 접근법은 이미 레이블이 지정된 자료를 재가공하여 네 가지 범주를 동시에 커버할 수 있게 했습니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>.
- **시험 문항**: dataset 외에도 연구자들은 안전 및 생활 기술 관련 다양한 시험 자료와 설문지에서 적합한 문항을 수동으로 선별했습니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 특히 윤리 및 법학 학습 시험(예: 기초 안전에 관한 학교 테스트)에서 문항을 추출했으며, 이는 불법 활동, 윤리 및 도덕 등의 범주와 관련 주제에 해당합니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 이러한 각 문항도 객관식 형식으로 변환되어 하나의 범주에 배정되었습니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>.
- **새로운 문항 생성**: 개인정보나 정신 건강 등 일부 측면의 경우 공개 소스에서 충분히 다양한 데이터를 확보하기 어려워, 저자들은 ChatGPT와 같은 고성능 언어 모델을 활용하여 추가 문항을 생성했습니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 해당 주제에 대한 다양한 상황을 생성하기 위한 prompt가 작성되었으며, 생성된 결과물은 benchmark에 포함되기 전 **전문가의 엄격한 필터링 및 검수**를 거쳤습니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 이러한 통제된 augmented 접근법은 범주 커버리지의 공백을 채우는 데 기여했습니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>.

결과적으로 SafetyBench의 각 문항은 중국어와 영어 **두 언어로 제공**되었습니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 콘텐츠 동등성을 확보하기 위해 저자들은 수집된 모든 영어 문항을 중국어로, 중국어 문항을 영어로 Baidu 상업용 기계 번역 API를 통해 번역했습니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 이 번역 방법을 선택한 이유는 일부 고성능 LLM(예: ChatGPT 자체)이 잠재적으로 위험한 콘텐츠의 처리나 정확한 번역을 거부하거나, 번역 과정에서 표현을 완화하는 경우가 있었기 때문입니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 자동 번역 결과는 발생 가능한 오류나 문화적 뉘앙스를 수정하기 위해 수동으로 교열 및 검토되었습니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 전반적으로 모든 문항은 **사람에 의한 품질 검수** 단계를 거쳤으며<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>, 이는 두 언어 모두에서 표현의 정확성과 예상 답변의 적합성을 보장하기 위한 것입니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>.

최종 dataset에서 소스별 분포는 대략 다음과 같습니다. 약 절반의 문항이 공개 dataset에서 가져온 것이며, 상당 부분은 시험 자료에서, 나머지는 모델이 생성(선별 후)한 것입니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 이러한 접근법은 주제의 폭넓은 커버리지와 충분한 깊이(각 범주에 다수의 예시)를 모두 확보할 수 있게 했습니다.

## 실험 방법 및 결과

SafetyBench 세트를 준비한 후 저자들은 현대 언어 모델의 안전성 이해 수준을 측정하기 위해 대규모 테스트를 진행했습니다. 모델 평가는 **자동으로** 수행됩니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 각 모델에 해당 언어로 모든 문항을 순차적으로 제시하고, 정답률(즉, 모델이 선택한 답변과 정답이 일치하는 비율)을 기록합니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 이 비율은 모델이 안전 문제를 얼마나 잘 이해하고 안전성 관점에서 올바른 답변을 제공하는지를 나타내는 지표입니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>.

개발자들이 진행한 테스트에는 두 언어 모두에서 다양한 출처의 **25개 인기 LLM**(오픈 소스 모델과 독점 API 서비스 모두 포함)이 참여했습니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 테스트는 두 가지 방식으로 수행되었습니다. **zero-shot** (어떠한 예시도 없이 문항에 답변) 및 **few-shot** (모델에게 문맥을 설정하기 위해 정답이 포함된 몇 가지 예시 문항을 미리 제시)<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 이러한 프로토콜은 모델의 기본 능력과 학습 힌트가 주어졌을 때 답변을 개선하는 능력을 모두 평가할 수 있게 합니다.

테스트의 주요 결론은 **현대 모델들은 안전 지식 수준에서 큰 차이를 보이며, 현재 이용 가능한 어떤 LLM도 모든 범주에서 완벽하지 않다**는 것입니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 테스트에서 선두를 차지한 모델은 **GPT-4**(OpenAI)로, 가장 높은 평균 정확도를 보이며 다수의 범주에서 다른 모든 모델을 크게 앞섰습니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. zero-shot 방식에서 GPT-4는 가장 가까운 경쟁자인 GPT-3.5-turbo를 전체 정확도 기준으로 거의 **10 퍼센트포인트** 차이로 앞섰습니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 특히 신체 안전 및 도덕적·윤리적 딜레마 문제에서 GPT-4가 경쟁 모델보다 눈에 띄게 높은 정답률을 보이는 등 일부 영역에서 격차가 두드러졌습니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>.

동시에 GPT-4에서도 **취약점**이 발견되었습니다. '편견 및 차별'(Unfairness and Bias) 범주에서 이 모델은 다른 섹션의 자체 결과와 비교하여 상대적으로 낮은 성능을 보였습니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 응답 분석 결과, GPT-4가 차별에 관한 중립적인 발언을 편견의 표현으로 잘못 분류하거나 특정 표현 및 사건에서 혼동을 일으키는 경우가 있는 것으로 나타났습니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 이러한 오류는 가장 발전된 모델조차도 발언의 윤리성 평가에 영향을 미치는 문화적·언어적 뉘앙스를 과소평가할 수 있음을 강조합니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>.

나머지 모델들은 GPT-4에 크게 뒤처졌습니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 평균적으로 대부분의 오픈 소스 LLM(LLaMA의 다양한 버전, Falcon, 국내 중국 모델 등 포함)은 **현저히 낮은 정확도**를 보이며, 정답률이 70~80%를 초과하지 못하는 경우가 많았습니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 이들 중 상당수는 특정 범주에서 특히 약한 모습을 보였습니다. 예를 들어 일부 모델은 사회적 편견이나 미묘한 윤리 문제와 관련된 섹션에서 70% 미만의 점수를 기록했습니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 전반적으로 GPT-4를 제외한 어떤 모델도 전체 안전 지표에서 80%라는 기준점을 넘지 못했으며, 이는 안전한 행동 개선을 위한 큰 발전 여지가 있음을 시사합니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. GPT-4와 오픈 소스 모델 간의 이러한 격차는 독점 모델에서 더욱 대규모의 학습과 목표 지향적 alignment 조정의 효과를 나타냅니다.

흥미롭게도 일부 시스템의 성능은 **언어에 따라 달랐습니다**<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 중국에서 개발된 모델(예: Baidu Ernie, Alibaba Tongyi 등)은 일반적으로 영어 버전보다 중국어 버전 테스트에서 더 나은 결과를 보였습니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 반면 OpenAI의 GPT 모델 계열은 더 균형 잡힌 결과를 보였습니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 이는 각 언어 데이터에 대한 학습 양과 품질의 차이, 그리고 일부 지역 모델에 내장된 필터나 검열 메커니즘의 유무를 반영하는 것일 수 있습니다.

few-shot 예시(테스트 전 몇 가지 데모 Q&A 제공)를 추가했을 때 **다양한 방향의 효과**가 관찰되었습니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 일부 모델은 힌트 덕분에 정확도를 눈에 띄게 높일 수 있었습니다. 예를 들어 text-davinci-003(GPT-3)과 같은 이전 세대 대형 언어 모델이나 중국어 InternLM은 five-shot 방식에서 눈에 띄는 품질 향상을 보였습니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 그러나 일부 모델에서는 추가 문맥이 결과를 거의 개선하지 못했으며, 경우에 따라 정확도가 오히려 하락하기도 했습니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 특히 GPT-3.5의 경우 저자들은 few-shot에서 소폭의 '**부정적 향상'**을 기록했으며<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>, 이를 '**alignment tax'(정렬 세금)** 현상과 연관 지었습니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 그럼에도 불구하고 평균적으로 예시 제공은 답변을 더 안정적으로 만들고, 모델이 명확한 답변을 거부하는 경우의 비율을 낮추는 효과가 있었습니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>.

연구자들은 별도로 중국어와 관련된 **필터링된 하위 문항 세트**에 대한 모델 성능도 평가했습니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 일부 대형 중국 모델의 API는 특정 '민감한' 단어가 포함된 요청을 자동으로 거부하기 때문입니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 이에 트리거 단어가 없는 2,100개 문항으로 구성된 축소 샘플을 만들어 five-shot 방식으로 일부 모델을 비교했습니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 결과에 따르면 이 완화된 버전에서 GPT-4와 최고의 현지 모델 간의 격차가 줄어들었습니다. 예를 들어 중국 모델 ChatGLM2는 GPT-4보다 약 3% 낮은 점수를 기록하여 종합 점수에서 거의 동등한 수준에 도달했습니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 또한 Baidu의 Ernie Bot은 대부분의 범주(편견 섹션 제외)에서 자신 있는 성과를 보이며 선두권에 근접했습니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 이 데이터는 엄격한 필터링 제어 하에서(가장 위험한 요청을 제외하면) 일부 국내 모델이 안전한 행동 측면에서 세계 선두 모델과 경쟁할 수 있음을 시사합니다.

## 벤치마크의 의의 및 개발자들의 결론

SafetyBench는 대형 언어 모델의 안전성을 체계적으로 측정하고 개선하기 위한 중요한 단계를 나타냅니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 사용자가 지시나 도발로 모델을 '해킹'하려는 직접적인 상호작용 시나리오와 달리, 이 benchmark는 AI가 **안전한 콘텐츠와 안전하지 않은 콘텐츠를 올바르게 이해하고 구별하는 능력**에 초점을 맞춥니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 저자들은 이러한 이해가 모델이 개방형 대화에서 안전한 응답을 생성할 수 있는 필수적인 토대라고 강조합니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 반면 도덕 규범, 예절 규칙, 독성의 징후 등을 깊이 습득하면 모델이 위험한 발언과 결정을 피하도록 조정하는 것이 용이해집니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 따라서 SafetyBench에서의 높은 점수는 **모델이 안전하게 운용될 준비가 되어 있다는 지표**로 볼 수 있으며<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>, 특정 범주에서의 실패는 수정이 필요한 위험 영역을 신호합니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>.

중요한 점은 SafetyBench가 의도적으로 **모델 지시 자체에 대한 공격과 관련된 일부 측면**(소위 jailbreak prompt, 역할 조작 등)을 포함하지 않는다는 것입니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 저자들은 instruction attacks 유형의 문제가 사용자 명령 수행과 내장된 안전 규칙 준수 사이의 충돌과 관련된 다른 성질을 가지고 있다고 설명합니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 이러한 측면은 다른 방법으로 해결되며 모델의 이해 범위를 벗어납니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 따라서 SafetyBench는 안전한 행동에 관한 모델의 지식이라는 콘텐츠 수준에 집중합니다. 그럼에도 불구하고 benchmark에서 7가지 핵심 범주의 종합적인 커버리지는 이미 모델의 취약점을 식별하는 것을 가능하게 합니다. 예를 들어 GPT-4가 편견 관련 문항에서 상대적으로 약한 결과를 보이며, 일부 오픈 소스 모델은 도덕이나 법률과 관련된 섹션에서 크게 뒤처지는 것으로 알려져 있습니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 이러한 정보는 개발자들에게 추가 학습이나 응답 필터링 과정에서 무엇을 개선해야 하는지에 대한 구체적인 방향을 제시합니다.

SafetyBench는 **커뮤니티에 공개**되어 있습니다<sup>[\[2\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-abs-2)</sup>. 데이터 및 방법론 자료가 자유롭게 이용 가능하며<sup>[\[2\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-abs-2)</sup>, 별도로 만들어진 플랫폼에서 다양한 모델의 결과에 대한 **온라인 리더보드**가 유지됩니다<sup>[\[2\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-abs-2)</sup>. 연구자들은 개발자들이 이 세트에서 새로운 모델을 테스트하고 결과를 공개하도록 초대하며, 이는 시스템의 투명한 비교와 AI 안전성 향상에 관한 진전 추적에 기여할 것입니다.

마지막으로 저자들은 SafetyBench의 목표가 단순히 또 다른 순위를 만드는 것이 아니라 **모델 개선을 촉진하는 것**임을 강조합니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 그들은 개발자들이 모델을 테스트에 '맞추려는' 시도에 그치지 않고 식별된 문제점을 체계적으로 해결하도록 촉구합니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 더 많은 데이터와 더 정교한 alignment 기법으로 새로운 버전의 모델이 학습될수록 SafetyBench에서의 점수도 향상될 것으로 기대됩니다<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup>. 향후 이 benchmark는 언어 모델의 안전 요건 준수를 검증하는 표준 도구가 될 수 있으며, 그 방법론은 책임 있는 AI 분야에서 더욱 발전된 테스트 세트를 개발하기 위한 기반이 될 수 있습니다.

## 참조

- SafetyBench 원본 논문 (arXiv)
- GitHub의 SafetyBench 저장소
- Hugging Face의 SafetyBench dataset 페이지
- ACL Anthology의 SafetyBench 논문

## 참고 문헌

- Liang, P. et al. (2022). *Holistic Evaluation of Language Models (HELM)*. arXiv:2211.09110.
- Chang, Y. et al. (2023). *A Survey on Evaluation of Large Language Models*. arXiv:2307.03109.
- Ni, S. et al. (2025). *A Survey on Large Language Model Benchmarks*. arXiv:2508.15361.
- Biderman, S. et al. (2024). *The Language Model Evaluation Harness (lm-eval): Guidance and Lessons Learned*. arXiv:2405.14782.
- Kiela, D. et al. (2021). *Dynabench: Rethinking Benchmarking in NLP*. arXiv:2104.14337.
- Ma, Z. et al. (2021). *Dynaboard: An Evaluation‑As‑A‑Service Platform for Holistic Next‑Generation Benchmarking*. arXiv:2106.06052.
- Goel, K. et al. (2021). *Robustness Gym: Unifying the NLP Evaluation Landscape*. arXiv:2101.04840.
- Xu, C. et al. (2024). *Benchmark Data Contamination of Large Language Models: A Survey*. arXiv:2406.04244.
- Liu, S. et al. (2025). *A Comprehensive Survey on Safety Evaluation of LLMs*. arXiv:2506.11094.
- Chiang, W.-L. et al. (2024). *Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference*. arXiv:2403.04132.
- Boubdir, M. et al. (2023). *Elo Uncovered: Robustness and Best Practices in Language Model Evaluation*. arXiv:2311.17295.
- Huang, L. et al. (2023). *A Survey on Hallucination in Large Language Models*. arXiv:2311.05232.

## 주석

<sup>[\[1\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-v2-1)</sup> <sup>[\[2\]](https://systems-analysis.info/int/SafetyBench_(KO)#cite_note-arxiv-main-abs-2)</sup> \</references\>

1.  <span id="cite_note-arxiv-main-v2-1">↑ <sup>[1.00](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-0)</sup> <sup>[1.01](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-1)</sup> <sup>[1.02](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-2)</sup> <sup>[1.03](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-3)</sup> <sup>[1.04](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-4)</sup> <sup>[1.05](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-5)</sup> <sup>[1.06](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-6)</sup> <sup>[1.07](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-7)</sup> <sup>[1.08](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-8)</sup> <sup>[1.09](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-9)</sup> <sup>[1.10](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-10)</sup> <sup>[1.11](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-11)</sup> <sup>[1.12](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-12)</sup> <sup>[1.13](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-13)</sup> <sup>[1.14](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-14)</sup> <sup>[1.15](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-15)</sup> <sup>[1.16](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-16)</sup> <sup>[1.17](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-17)</sup> <sup>[1.18](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-18)</sup> <sup>[1.19](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-19)</sup> <sup>[1.20](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-20)</sup> <sup>[1.21](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-21)</sup> <sup>[1.22](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-22)</sup> <sup>[1.23](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-23)</sup> <sup>[1.24](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-24)</sup> <sup>[1.25](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-25)</sup> <sup>[1.26](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-26)</sup> <sup>[1.27](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-27)</sup> <sup>[1.28](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-28)</sup> <sup>[1.29](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-29)</sup> <sup>[1.30](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-30)</sup> <sup>[1.31](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-31)</sup> <sup>[1.32](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-32)</sup> <sup>[1.33](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-33)</sup> <sup>[1.34](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-34)</sup> <sup>[1.35](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-35)</sup> <sup>[1.36](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-36)</sup> <sup>[1.37](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-37)</sup> <sup>[1.38](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-38)</sup> <sup>[1.39](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-39)</sup> <sup>[1.40](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-40)</sup> <sup>[1.41](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-41)</sup> <sup>[1.42](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-42)</sup> <sup>[1.43](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-43)</sup> <sup>[1.44](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-44)</sup> <sup>[1.45](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-45)</sup> <sup>[1.46](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-46)</sup> <sup>[1.47](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-47)</sup> <sup>[1.48](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-48)</sup> <sup>[1.49](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-49)</sup> <sup>[1.50](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-50)</sup> <sup>[1.51](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-51)</sup> <sup>[1.52](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-52)</sup> <sup>[1.53](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-53)</sup> <sup>[1.54](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-54)</sup> <sup>[1.55](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-55)</sup> <sup>[1.56](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-56)</sup> <sup>[1.57](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-57)</sup> <sup>[1.58](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-58)</sup> <sup>[1.59](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-59)</sup> <sup>[1.60](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-60)</sup> <sup>[1.61](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-61)</sup> <sup>[1.62](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-62)</sup> <sup>[1.63](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-63)</sup> <sup>[1.64](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-64)</sup> <sup>[1.65](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-65)</sup> <sup>[1.66](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-66)</sup> <sup>[1.67](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-67)</sup> <sup>[1.68](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-68)</sup> <sup>[1.69](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-69)</sup> <sup>[1.70](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-70)</sup> <sup>[1.71](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-71)</sup> <sup>[1.72](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-72)</sup> <sup>[1.73](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-73)</sup> <sup>[1.74](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-74)</sup> <sup>[1.75](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-75)</sup> <sup>[1.76](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-76)</sup> <sup>[1.77](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-77)</sup> <sup>[1.78](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-78)</sup> <sup>[1.79](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-79)</sup> <sup>[1.80](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-80)</sup> <sup>[1.81](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-81)</sup> <sup>[1.82](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-82)</sup> <sup>[1.83](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-83)</sup> <sup>[1.84](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-84)</sup> <sup>[1.85](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-85)</sup> <sup>[1.86](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-86)</sup> <sup>[1.87](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-87)</sup> <sup>[1.88](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-88)</sup> <sup>[1.89](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-89)</sup> <sup>[1.90](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-90)</sup> <sup>[1.91](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-91)</sup> <sup>[1.92](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-92)</sup> <sup>[1.93](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-93)</sup> <sup>[1.94](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-v2_1-94)</sup> Zhang, Yuntao et al. «SafetyBench: Evaluating the Safety of Large Language Models with Multiple Choice Questions». *arXiv*. <a href="https://ar5iv.labs.arxiv.org/html/2309.07045v2" class="external autonumber" rel="nofollow">[1]</a></span>
2.  <span id="cite_note-arxiv-main-abs-2">↑ <sup>[2.0](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-abs_2-0)</sup> <sup>[2.1](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-abs_2-1)</sup> <sup>[2.2](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-abs_2-2)</sup> <sup>[2.3](https://systems-analysis.info/int/SafetyBench_(KO)#cite_ref-arxiv-main-abs_2-3)</sup> Zhang, Yuntao et al. «SafetyBench: Evaluating the Safety of Large Language Models». *arXiv*. <a href="https://arxiv.org/abs/2309.07045" class="external autonumber" rel="nofollow">[2]</a></span>
