---
title: "Jailbreaks (LLM) (TH)"
source: "https://systems-analysis.info/int/Jailbreaks_(LLM)_(TH)"
wiki: "systems-analysis.info/int"
article: "Jailbreaks_(LLM)_(TH)"
language: "th"
categories:
  - "Category:Large language models"
  - "Category:Machine learning"
  - "Category:Thai"
revision_id: 3303
wiki_created_at: 2026-09-06T23:18:41Z
wiki_modified_at: 2026-09-06T23:18:41Z
downloaded_at: 2026-09-07T22:56:14Z
---

# Jailbreaks (LLM) (TH)

**เจลเบรก** (อังกฤษ: **Jailbreak** — แปลตรงตัวว่า \\การหลบหนีจากคุก\\) ในบริบทของโมเดลภาษาขนาดใหญ่ (LLM) คือประเภทของการโจมตีแบบปฏิปักษ์ที่มีเป้าหมายเพื่อเลี่ยงกลไกความปลอดภัยและข้อจำกัดที่ฝังอยู่ภายในโมเดล เพื่อให้ได้รับคำตอบที่ต้องห้ามหรืออาจเป็นอันตราย<sup>[\[1\]](https://systems-analysis.info/int/Jailbreaks_(LLM)_(TH)#cite_note-lillog_intro-1)</sup> Jailbreak หมายถึง \\การชักนำโมเดลให้สร้างคำตอบที่เป็นอันตรายซึ่งขัดต่อนโยบายการใช้งานและบรรทัดฐานทางสังคม ผ่านการพัฒนา prompt แบบปฏิปักษ์\\<sup>[\[2\]](https://systems-analysis.info/int/Jailbreaks_(LLM)_(TH)#cite_note-yi_2024_survey-2)</sup>

ช่องโหว่พื้นฐานที่ถูกใช้ประโยชน์ในการโจมตีแบบ jailbreak อยู่ที่ลักษณะเฉพาะทางสถาปัตยกรรมของ LLM: โมเดลไม่สามารถแยกแยะระหว่างคำสั่งและข้อมูลตามประเภทได้ เนื่องจากทั้ง system prompt และข้อมูลนำเข้าจากผู้ใช้มีรูปแบบเดียวกัน นั่นคือสตริงข้อความในภาษาธรรมชาติ<sup>[\[3\]](https://systems-analysis.info/int/Jailbreaks_(LLM)_(TH)#cite_note-prompting_guide_vuln-3)</sup>

## ประวัติความเป็นมาและพัฒนาการ

### ช่วงเริ่มต้น: Prompt Injection (2022)

การค้นพบช่องโหว่ต่อ prompt injection ที่มีการบันทึกเป็นครั้งแรกเกิดขึ้นในเดือนพฤษภาคม 2022 เมื่อนักวิจัยจากบริษัท **Preamble** ค้นพบความเปราะบางของ ChatGPT ต่อการโจมตีดังกล่าว ในเดือนกันยายน 2022 **Riley Goodside** ได้เผยแพร่การสาธิตสาธารณะครั้งแรกของช่องโหว่ใน GPT-3 บน Twitter โดยอิสระ พร้อมตัวอย่างที่เป็นที่รู้จักซึ่งสั่งให้โมเดลเพิกเฉยต่อคำสั่งก่อนหน้า<sup>[\[4\]](https://systems-analysis.info/int/Jailbreaks_(LLM)_(TH)#cite_note-ncc_group_history-4)</sup>

### ยุค DAN (2022–2023)

ในช่วงกลางปี 2022 ปรากฏ prompt แรกในรูปแบบ **"Do Anything Now"** (**DAN**) ซึ่งเป็นคำสั่งสำหรับการแสดงบทบาทสมมติ นวัตกรรมสำคัญคือการใช้การเล่นตามบทบาทเพื่อเลี่ยงข้อจำกัดด้านความปลอดภัยโดยการสร้าง \\บุคลิกทางเลือก\\ ที่ปราศจากกฎเกณฑ์<sup>[\[5\]](https://systems-analysis.info/int/Jailbreaks_(LLM)_(TH)#cite_note-dan_evolution_arxiv-5)</sup> วิวัฒนาการของ DAN นำไปสู่การปรากฏของสถานการณ์ที่ซับซ้อนพร้อมระบบ token (กลไกการลงโทษ/รางวัล) และกลไกการรักษาบุคลิกตัวละคร<sup>[\[6\]](https://systems-analysis.info/int/Jailbreaks_(LLM)_(TH)#cite_note-dan_github-6)</sup>

### การกระจายตัวของวิธีการ (2023–2024)

ตั้งแต่ปี 2023 เริ่มมีการวิจัยเชิงวิชาการอย่างครอบคลุมเกี่ยวกับการโจมตีแบบ jailbreak ในปี 2024 ปรากฏ **การโจมตีแบบมัลติโมดัล** ซึ่งรวมถึงการซ่อนคำสั่งที่เป็นอันตรายในรูปภาพ ไฟล์เสียง รวมถึง visual prompt injection ผ่าน ASCII-art<sup>[\[7\]](https://systems-analysis.info/int/Jailbreaks_(LLM)_(TH)#cite_note-hiddenlayer_multimodal-7)</sup>

### ช่วงปัจจุบัน (2024–2025)

เทคนิคการโจมตียังคงมีความซับซ้อนเพิ่มขึ้นเรื่อยๆ ในเดือนพฤศจิกายน 2024 มีการค้นพบเทคนิค **"Time Bandit"** ที่ใช้ประโยชน์จากความสับสนด้านเวลาใน ChatGPT-4o โดยการตั้งคำถามราวกับว่ามาจากช่วงเวลาทางประวัติศาสตร์ (ช่วงปี ค.ศ. 1800–1900)<sup>[\[8\]](https://systems-analysis.info/int/Jailbreaks_(LLM)_(TH)#cite_note-bleeping_computer_time_bandit-8)</sup>

## วิธีการทางเทคนิคและการจำแนกประเภท

การโจมตีสามารถจำแนกตามระดับการเข้าถึงโมเดล:

- **การโจมตีแบบกล่องดำ**: โดยไม่มีการเข้าถึงส่วนประกอบภายในของโมเดล (พารามิเตอร์, ค่า gradient)
- **การโจมตีแบบกล่องขาว**: ด้วยการเข้าถึงพารามิเตอร์และค่า gradient ของโมเดลอย่างเต็มรูปแบบ<sup>[\[2\]](https://systems-analysis.info/int/Jailbreaks_(LLM)_(TH)#cite_note-yi_2024_survey-2)</sup>

### อนุกรมวิธาน JailbreakRadar

การจำแนกประเภทแบบ *JailbreakRadar* (Chu et al., 2024) แบ่งหมวดหมู่การโจมตีหลักออกเป็นหกประเภท:

1.  **การโจมตีโดยตรง:** Prompt ที่เป็นอันตรายโดยตรง
2.  **การโจมตีทางอ้อม:** กลยุทธ์การจัดการแบบหลายขั้นตอน
3.  **การโจมตีเชิงบริบท:** การใช้ประวัติการสนทนา
4.  **การโจมตีเชิงบทบาท:** เทคนิคการแสดงเป็นตัวละคร (เช่น DAN)
5.  **การโจมตีเชิงการเข้ารหัส:** วิธีการ obfuscation เพื่อซ่อนคำสั่งที่เป็นอันตราย
6.  **การโจมตีเชิงแม่แบบ:** กรอบการทำงานแบบปฏิปักษ์ที่มีโครงสร้าง<sup>[\[9\]](https://systems-analysis.info/int/Jailbreaks_(LLM)_(TH)#cite_note-chu_2024_radar-9)</sup>

### กลไกทางเทคนิค

- **การสร้าง suffix แบบปฏิปักษ์ (GCG):** วิธีการที่เสนอโดย Zou et al. (2023) ซึ่งสร้าง suffix แบบปฏิปักษ์ (ลำดับของ token) โดยอัตโนมัติ ซึ่งเมื่อเพิ่มเข้าไปใน prompt จะกระตุ้นให้เกิดคำตอบที่เป็นอันตรายด้วยความน่าจะเป็นสูง วิธีการนี้ใช้การปรับแต่งแบบ gradient และแสดงให้เห็นถึงอัตราความสำเร็จสูง (สูงถึง 84% บน GPT-4) และความสามารถในการถ่ายโอนระหว่างโมเดล<sup>[\[10\]](https://systems-analysis.info/int/Jailbreaks_(LLM)_(TH)#cite_note-zou_2023_gcg-10)</sup>
- **การเจลเบรกแบบหลายรอบ:** การวิจัยของ Anthropic (2024) แสดงให้เห็นว่าประสิทธิภาพของการโจมตีเป็นไปตามกฎกำลัง: เมื่อจำนวนตัวอย่างที่เป็นอันตรายใน prompt เพิ่มขึ้น สัดส่วนของคำตอบที่ไม่พึงประสงค์ก็เพิ่มขึ้นตาม<sup>[\[11\]](https://systems-analysis.info/int/Jailbreaks_(LLM)_(TH)#cite_note-anthropic_many_shot-11)</sup>

## กลไกการป้องกัน

- **ตัวจำแนกประเภทแบบรัฐธรรมนูญ (Anthropic):** การกรองข้อมูลนำเข้า/ส่งออกโดยอิงหลักการรัฐธรรมนูญ วิธีการนี้ช่วยลดอัตราความสำเร็จของ jailbreak จาก 86% เหลือ 4.4% ในการประเมินที่ควบคุม<sup>[\[12\]](https://systems-analysis.info/int/Jailbreaks_(LLM)_(TH)#cite_note-mit_tech_review_const_ai-12)</sup>
- **Reinforcement Learning from Human Feedback (RLHF):** การฝึกสามขั้นตอน (OpenAI) ซึ่งรวมถึงการปรับแต่งแบบมีผู้ดูแล การฝึกโมเดลรางวัล และการปรับแต่งนโยบาย แสดงให้เห็นถึงการลดลงอย่างมีนัยสำคัญในการสร้างเนื้อหาที่เป็นพิษ
- **การฝึกแบบปฏิปักษ์:** การฝึกโมเดลบนตัวอย่างของการโจมตีแบบ jailbreak เพื่อเพิ่มความทนทาน ประสิทธิภาพของแนวทางนี้ในการลดอัตราความสำเร็จของการโจมตีประเมินไว้ที่ 60–80%<sup>[\[1\]](https://systems-analysis.info/int/Jailbreaks_(LLM)_(TH)#cite_note-lillog_intro-1)</sup>
- **การป้องกันแบบหลายชั้น:** กลยุทธ์ที่แนะนำซึ่งรวมถึงการตรวจสอบความถูกต้องของข้อมูลนำเข้า การป้องกันในระดับโมเดล การตรวจสอบข้อมูลส่งออก และการติดตามอย่างต่อเนื่องในเวลาจริง

การโจมตีแบบ jailbreak ต่อโมเดลภาษาขนาดใหญ่เป็นปัญหาความปลอดภัยของ AI เชิงพื้นฐาน ที่แสดงให้เห็นถึงความตึงเครียดอย่างต่อเนื่องระหว่างความสามารถและการปรับแนวทางของโมเดล ภูมิทัศน์ของการโจมตียิ่งซับซ้อนขึ้นเรื่อยๆ เปลี่ยนจาก prompt injection ง่ายๆ ไปสู่การโจมตีแบบมัลติโมดัลและแบบอัตโนมัติที่ซับซ้อน การวิจัยแสดงให้เห็นว่าไม่มีกลไกการป้องกันใดในปัจจุบันที่ทนทานต่อความพยายาม jailbreak ทุกรูปแบบอย่างสมบูรณ์ ความสำเร็จในด้านนี้ต้องอาศัยการลงทุนอย่างต่อเนื่องในการวิจัยด้านความปลอดภัย แนวปฏิบัติการเปิดเผยข้อมูลอย่างรับผิดชอบ และความพยายามร่วมกันของนักวิจัย ภาคอุตสาหกรรม และหน่วยงานกำกับดูแล

## อ้างอิง

- Prompting Guide: Jailbreaking
- คลังข้อมูลการนำการโจมตี GCG ไปใช้งานบน GitHub

## บรรณานุกรม

- Bai, Y. et al. (2022). *Constitutional AI: Harmlessness from AI Feedback*. arXiv:2212.08073.
- Zou, A. et al. (2023). *Universal and Transferable Adversarial Attacks on Aligned Language Models*. arXiv:2307.15043.
- Shen, X. et al. (2023). *"Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models*. arXiv:2308.03825.
- Chao, P. et al. (2024). *JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models*. arXiv:2404.01318.
- Liao, Z.; Sun, H. (2024). *AmpleGCG: Learning a Universal and Transferable Generative Model of Adversarial Suffixes for Jailbreaking Both Open and Closed LLMs*. OpenReview UfqzXg95I5.
- Yi, S. et al. (2024). *Jailbreak Attacks and Defenses Against Large Language Models: A Survey*. arXiv:2407.04295.
- Chu, J. et al. (2025). *JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs*. arXiv:2402.05668.
- Liu, A. et al. (2025). *PiCo: Jailbreaking Multimodal Large Language Models via Pictorial Code Contextualization*. arXiv:2504.01444.
- Ghosal, D. et al. (2025). *Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Filtering*. *CVPR 2025*. PDF.
- Yan, Q. et al. (2025). *Hidden in Plain Sight: Probing Implicit Reasoning in Multimodal Language Models*. arXiv:2506.00258.
- Liu, Y. et al. (2025). *RePD: Defending Jailbreak Attack through a Retrieval-Based Detector*. *Findings of NAACL 2025*. ACL Anthology.

## หมายเหตุ

1.  <span id="cite_note-lillog_intro-1">↑ <sup>[1.0](https://systems-analysis.info/int/Jailbreaks_(LLM)_(TH)#cite_ref-lillog_intro_1-0)</sup> <sup>[1.1](https://systems-analysis.info/int/Jailbreaks_(LLM)_(TH)#cite_ref-lillog_intro_1-1)</sup> «A brief history of jailbreaking». *Lil'Log*. <a href="https://lilianweng.github.io/posts/2023-03-15-prompt-engineering/" class="external autonumber" rel="nofollow">[1]</a></span>
2.  <span id="cite_note-yi_2024_survey-2">↑ <sup>[2.0](https://systems-analysis.info/int/Jailbreaks_(LLM)_(TH)#cite_ref-yi_2024_survey_2-0)</sup> <sup>[2.1](https://systems-analysis.info/int/Jailbreaks_(LLM)_(TH)#cite_ref-yi_2024_survey_2-1)</sup> Yi, J., et al. «Jailbreak Attacks and Defenses Against Large Language Models: A Comprehensive Survey». *arXiv:2405.09443*. <a href="https://arxiv.org/abs/2405.09443" class="external autonumber" rel="nofollow">[2]</a></span>
3.  <span id="cite_note-prompting_guide_vuln-3">[↑](https://systems-analysis.info/int/Jailbreaks_(LLM)_(TH)#cite_ref-prompting_guide_vuln_3-0) «Jailbreaking LLMs». *Prompting Guide*. <a href="https://www.promptingguide.ai/risks/jailbreaking" class="external autonumber" rel="nofollow">[3]</a></span>
4.  <span id="cite_note-ncc_group_history-4">[↑](https://systems-analysis.info/int/Jailbreaks_(LLM)_(TH)#cite_ref-ncc_group_history_4-0) «Exploring prompt injection attacks». *NCC Group*. <a href="https://research.nccgroup.com/2022/12/05/exploring-prompt-injection-attacks/" class="external autonumber" rel="nofollow">[4]</a></span>
5.  <span id="cite_note-dan_evolution_arxiv-5">[↑](https://systems-analysis.info/int/Jailbreaks_(LLM)_(TH)#cite_ref-dan_evolution_arxiv_5-0) «Do Anything Now: Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models». *arXiv:2308.03825*. <a href="https://arxiv.org/abs/2308.03825" class="external autonumber" rel="nofollow">[5]</a></span>
6.  <span id="cite_note-dan_github-6">[↑](https://systems-analysis.info/int/Jailbreaks_(LLM)_(TH)#cite_ref-dan_github_6-0) «0xk1h0/ChatGPT_DAN». *GitHub*. <a href="https://github.com/0xk1h0/ChatGPT_DAN" class="external autonumber" rel="nofollow">[6]</a></span>
7.  <span id="cite_note-hiddenlayer_multimodal-7">[↑](https://systems-analysis.info/int/Jailbreaks_(LLM)_(TH)#cite_ref-hiddenlayer_multimodal_7-0) «Hiding in Plain Sight: Multimodal Jailbreaking of Large Language Models». *HiddenLayer*. <a href="https://hiddenlayer.com/research/multimodal-jailbreaking-of-large-language-models/" class="external autonumber" rel="nofollow">[7]</a></span>
8.  <span id="cite_note-bleeping_computer_time_bandit-8">[↑](https://systems-analysis.info/int/Jailbreaks_(LLM)_(TH)#cite_ref-bleeping_computer_time_bandit_8-0) «ChatGPT "Time-travel" jailbreak lets you bypass its safety guards». *BleepingComputer*. <a href="https://www.bleepingcomputer.com/news/security/chatgpt-time-travel-jailbreak-lets-you-bypass-its-safety-guards/" class="external autonumber" rel="nofollow">[8]</a></span>
9.  <span id="cite_note-chu_2024_radar-9">[↑](https://systems-analysis.info/int/Jailbreaks_(LLM)_(TH)#cite_ref-chu_2024_radar_9-0) Chu, Z., et al. «JailbreakRadar: A Comprehensive Benchmark for Jailbreak Attack and Defense». *arXiv:2402.12642*. <a href="https://arxiv.org/abs/2402.12642" class="external autonumber" rel="nofollow">[9]</a></span>
10. <span id="cite_note-zou_2023_gcg-10">[↑](https://systems-analysis.info/int/Jailbreaks_(LLM)_(TH)#cite_ref-zou_2023_gcg_10-0) Zou, A., et al. «Universal and Transferable Adversarial Attacks on Aligned Language Models». *arXiv:2307.15043*. <a href="https://arxiv.org/abs/2307.15043" class="external autonumber" rel="nofollow">[10]</a></span>
11. <span id="cite_note-anthropic_many_shot-11">[↑](https://systems-analysis.info/int/Jailbreaks_(LLM)_(TH)#cite_ref-anthropic_many_shot_11-0) «Many-shot Jailbreaking». *Anthropic*. <a href="https://www.anthropic.com/research/many-shot-jailbreaking" class="external autonumber" rel="nofollow">[11]</a></span>
12. <span id="cite_note-mit_tech_review_const_ai-12">[↑](https://systems-analysis.info/int/Jailbreaks_(LLM)_(TH)#cite_ref-mit_tech_review_const_ai_12-0) «How we're using 'constitutional AI' to make our models safer». *MIT Technology Review*. <a href="https://www.technologyreview.com/2023/05/09/1072782/how-were-using-constitutional-ai-to-make-our-models-safer/" class="external autonumber" rel="nofollow">[12]</a></span>
