---
title: "ERNIE (Baidu) (TH)"
source: "https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)"
wiki: "systems-analysis.info/int"
article: "ERNIE_(Baidu)_(TH)"
language: "th"
categories:
  - "Category:Large language models"
  - "Category:Machine learning"
  - "Category:Thai"
revision_id: 1813
wiki_created_at: 2026-09-06T22:54:26Z
wiki_modified_at: 2026-09-06T22:54:26Z
downloaded_at: 2026-09-07T22:47:57Z
---

# ERNIE (Baidu) (TH)

**ERNIE (Baidu)** — ชุดโมเดลภาษาที่ผ่านการพรีเทรน (Large Language Model, LLM) และโมเดลพื้นฐานแบบมัลติโมดัล ที่พัฒนาโดยบริษัท Baidu ตั้งแต่ปี 2019 ภายใต้ชื่อรวม **Enhanced Representation through kNowledge IntEgration** (การเพิ่มประสิทธิภาพการแทนความหมายผ่านการผสานความรู้) ชุดโมเดลนี้มุ่งเน้นการยกระดับคุณภาพความเข้าใจเชิงความหมายและการสร้างข้อความ ด้วยการผสานความรู้ภายนอกจากกราฟความรู้ (Knowledge Graph, KG) ออนโทโลยี และสารานุกรมอย่างชัดเจนในกระบวนการพรีเทรน<sup>[\[1\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-ERNIE1-1)[\[2\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-ERNIE3-2)</sup>

โมเดลในชุดนี้มีวิวัฒนาการจากรูปแบบแรกที่อิงกับ BERT พร้อมการมาสก์เอนทิตีและวลี (ERNIE 1.0, 2.0) ผ่านสถาปัตยกรรมแบบรวมศูนย์ขนาดใหญ่ (ERNIE 3.0, ERNIE 3.0 Titan) ไปสู่ระบบมัลติโมดัลที่ใช้ Mixture-of-Experts (MoE) แบบโอเพนซอร์ส (ERNIE 4.5) และโมเดลออมนิโมดัลแบบเนทีฟ (ERNIE 5.0) ชุดโมเดลนี้รองรับงานด้านความเข้าใจข้อความ (Natural Language Understanding, NLU) การสร้างข้อความ (Natural Language Generation, NLG) และในเวอร์ชันหลัง รองรับงานมัลติโมดัล (ข้อความ รูปภาพ วิดีโอ เสียง) โดยเน้นภาษาจีนและรองรับหลายภาษา<sup>[\[2\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-ERNIE3-2)[\[3\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-ERNIE45-3)[\[4\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-ERNIE50-4)</sup>

## ประวัติและความเป็นมา

การพัฒนาชุดโมเดล ERNIE เริ่มต้นในปี 2019 ในแผนกวิจัยของ Baidu เพื่อตอบสนองต่อความสำเร็จของโมเดล BERT (Devlin et al., 2018) และความจำเป็นในการปรับโมเดลพรีเทรนให้เหมาะกับภาษาจีน ซึ่งการขาดช่องว่างระหว่างคำอย่างชัดเจนและสัดส่วนสูงของอักขระที่มีความหมายหลายนัยทำให้การจับหน่วยความหมาย (เอนทิตี วลี) ทำได้ยากขึ้น<sup>[\[1\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-ERNIE1-1)</sup>

### ERNIE 1.0 (2019)

โมเดลแรกของชุด (Sun et al., arXiv:1904.09223, เมษายน 2019) นำเสนอกลยุทธ์การมาสก์ความรู้แบบหลายระดับ (knowledge masking) สำหรับสถาปัตยกรรมแบบ BERT: แทนที่จะมาสก์โทเค็นแต่ละตัวแบบสุ่ม โมเดลจะมาสก์ทั้งเอนทิตีและวลีที่ระบุผ่านการจดจำเอนทิตีที่มีชื่อ (Named Entity Recognition, NER) และรูปแบบวลี<sup>[\[1\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-ERNIE1-1)</sup>

สถาปัตยกรรมของ ERNIE 1.0 อิงกับ Transformer encoder (12 ชั้น มิติที่ซ่อนอยู่ 768 หัวความใส่ใจ 12 หัว) นอกจากนี้ยังมีการนำงาน Dialogue Language Model (DLM) มาใช้สำหรับการเรียนรู้จากข้อมูลบทสนทนา โมเดลนี้ฝึกบนประโยคภาษาจีนประมาณ 173 ล้านประโยคจากคอร์ปัส Wikipedia, Baidu Baike, ข่าว และฟอรั่ม Baidu Tieba ณ เวลาที่เผยแพร่ ERNIE 1.0 ทำได้ผล state-of-the-art (SOTA) บนงาน NLP ภาษาจีน 5 งาน ได้แก่ XNLI (ความแม่นยำ 78.4% บนชุดทดสอบ เทียบกับ 77.2% ของ BERT), MSRA-NER (F1 93.8% เทียบกับ 92.6%)<sup>[\[1\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-ERNIE1-1)</sup>

### ERNIE 2.0 (2019)

ERNIE 2.0 (Sun et al., arXiv:1907.12412, กรกฎาคม 2019; ได้รับการตอบรับในงานประชุม AAAI 2020) นำเสนอกรอบการพรีเทรนแบบหลายงานต่อเนื่อง (continual multi-task pre-training) ซึ่งงานพรีเทรนใหม่ถูกเพิ่มเข้าสู่ Transformer ร่วมอย่างลำดับ (แบบ incremental) โดยไม่ต้องเทรนโมเดลใหม่ทั้งหมด<sup>[\[5\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-ERNIE2-5)</sup>

งานพรีเทรนของ ERNIE 2.0 แบ่งออกเป็นสามกลุ่ม:

- **Word-aware** — การมาสก์ความรู้ (knowledge masking) การทำนายตัวพิมพ์ใหญ่ (capitalization prediction) ความสัมพันธ์ระหว่างโทเค็นกับเอกสาร (token-document relation)
- **Structure-aware** — การเรียงลำดับประโยคใหม่ (sentence reordering) การกำหนดระยะห่างระหว่างประโยค (sentence distance)
- **Semantic-aware** — การทำนายความสัมพันธ์เชิงวาทกรรม (discourse relation prediction) ความเกี่ยวข้องสำหรับการค้นหาข้อมูล (IR relevance prediction)<sup>[\[5\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-ERNIE2-5)</sup>

โมเดลนี้ฝึกบนคอร์ปัสภาษาอังกฤษและภาษาจีน (สารานุกรม หนังสือ Reddit บันทึกการค้นหา) ผลลัพธ์: คะแนนเฉลี่ย GLUE สำหรับโมเดล base อยู่ที่ 80.6 (เทียบกับ 78.3 ของ BERT) สำหรับโมเดล large อยู่ที่ 83.6; ERNIE 2.0 ทำผลได้เหนือกว่า BERT และ XLNet ใน 16 งาน<sup>[\[5\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-ERNIE2-5)</sup>

### ERNIE 3.0 (2021)

กรอบ ERNIE 3.0 (Sun et al., arXiv:2107.02137, กรกฎาคม 2021) ผสานรวมพาราไดม์การพรีเทรนแบบ auto-encoding (AE) และ auto-regressive (AR) ในสถาปัตยกรรมเดียว ซึ่งแบ่งออกเป็นโมดูลการแทนความหมายสากล (Universal Representation Module) และโมดูลเฉพาะงาน (Task-specific Representation Modules) สำหรับ NLU และ NLG<sup>[\[2\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-ERNIE3-2)</sup>

โมเดลที่มี 10 พันล้านพารามิเตอร์นี้ฝึกบนข้อความภาษาจีน 4 TB (11 หมวดหมู่: Baidu Baike, Wikipedia, บันทึกการค้นหา ระบบถามตอบ ข้อความเฉพาะโดเมน) และกราฟความรู้ที่มีข้อเท็จจริงมากกว่า 50 ล้านรายการ ERNIE 3.0 ทำได้ SOTA บนงาน NLP ภาษาจีน 54 งาน และบน benchmark SuperGLUE ภาษาอังกฤษได้ 90.6% (ณ วันที่ 3 กรกฎาคม 2021 สูงกว่าค่าอ้างอิงของมนุษย์ที่ 89.8%)<sup>[\[2\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-ERNIE3-2)</sup>

### ERNIE 3.0 Titan (2021)

ในงานวิจัย «ERNIE 3.0 Titan: Exploring Larger-scale Knowledge Enhanced Pre-training for Language Understanding and Generation» (Wang et al., arXiv:2112.12731, ธันวาคม 2021) อธิบายการขยายกรอบ ERNIE 3.0 ไปสู่โมเดลแบบ dense ที่มี 260 พันล้านพารามิเตอร์ ซึ่งพัฒนาบนแพลตฟอร์ม PaddlePaddle<sup>[\[6\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-Titan-6)</sup>

Titan นำเสนอกลไกเพิ่มเติม ได้แก่ self-supervised adversarial loss สำหรับประเมินความน่าเชื่อถือของข้อความที่สร้างขึ้น, controllable language modeling loss สำหรับควบคุมสไตล์ หัวข้อ อารมณ์ และความยาวของการสร้างข้อความ รวมถึงการดิสทิลล์แบบออนไลน์ (On-the-Fly Distillation, OFD) เพื่อได้โมเดลนักเรียนขนาดเล็กระหว่างกระบวนการพรีเทรน โมเดลนี้ประเมินผลบนชุดข้อมูล 68 ชุด และตามข้อมูลของผู้เขียน ทำผลได้เหนือกว่าโมเดลก่อนหน้าทั้ง 68 ชุดข้อมูล<sup>[\[6\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-Titan-6)</sup>

### ERNIE Bot และเวอร์ชันเชิงพาณิชย์ (2023–2025)

ในเดือนมีนาคม 2023 Baidu เปิดตัวแชทบอท **ERNIE Bot** (Wenxin Yiyan, 文心一言) ที่อิงบนโมเดล ERNIE 3.x และ PLATO (โมเดลบทสนทนา) การเข้าถึงแบบสาธารณะเปิดให้บริการในเดือนสิงหาคม 2023 หลังจากได้รับการอนุมัติจากหน่วยงานกำกับดูแล การอัปเดตรวมถึง ERNIE 3.5 (มิถุนายน 2023) และ ERNIE 4.0 (ตุลาคม 2023)<sup>[\[7\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-Wiki_ErnieBot-7)[\[8\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-BaiduBlog-8)</sup>

ตามรายงาน SuperBench ของมหาวิทยาลัยชิงหัว ERNIE Bot 4.0 ได้อันดับหนึ่งในบรรดา LLM ของจีนตามตัวชี้วัดรวม แสดงผลลัพธ์ที่ดีในความเข้าใจภาษาจีนและการปฏิบัติตามคำสั่ง<sup>[\[9\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-Tsinghua-9)[\[10\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-Multiplatform-10)</sup>

ในปี 2022 ควบคู่กับสายโมเดลภาษาหลัก ได้มีการนำเสนอส่วนขยายมัลติโมดัล **ERNIE-ViLG 2.0** (arXiv:2210.15257) — โมเดลสร้างรูปภาพจากข้อความ (text-to-image) ซึ่งเป็นหนึ่งในก้าวแรกสู่ความสามารถมัลติโมดัล<sup>[\[11\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-ViLG-11)</sup>

ในปี 2024 ในงานประชุม WAVE SUMMIT Baidu นำเสนอ **ERNIE 4.0 Turbo** — เวอร์ชันที่ปรับให้เหมาะสมสำหรับแอปพลิเคชันความเร็วสูง ซึ่งสามารถสร้างข้อความที่มีความยาวมากกว่าพันคำในเวลาประมาณ 20 วินาที โดยยังคงคุณภาพไว้<sup>[\[12\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-Turbo-12)</sup>

### ERNIE 4.5 (2025)

ในเดือนมิถุนายน 2025 Baidu เผยแพร่รายงานทางเทคนิคของ ERNIE 4.5 ซึ่งนำเสนอชุดโมเดล 10 โมเดล ได้แก่ LLM ข้อความและโมเดลมัลติโมดัล (Vision-Language, VL) สถาปัตยกรรมอิงกับ Mixture-of-Experts (MoE) แบบเฮเทอโรจีเนียสที่แยกผู้เชี่ยวชาญตามโมดัลลิตี ตัวเลือกรวมถึงโมเดล MoE ที่มีพารามิเตอร์รวมสูงสุด 424 พันล้านและพารามิเตอร์ที่ใช้งาน 47 พันล้าน รวมถึงโมเดลแบบ dense ที่มี 0.3 พันล้านพารามิเตอร์ ชุดโมเดลนี้เผยแพร่ภายใต้ลิขสิทธิ์ Apache 2.0 บน Hugging Face และ GitHub (PaddlePaddle/ERNIE)<sup>[\[3\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-ERNIE45-3)</sup>

### ERNIE 5.0 (2026)

ในเดือนกุมภาพันธ์ 2026 มีการเผยแพร่รายงานทางเทคนิคของ ERNIE 5.0 (arXiv:2602.04705) — โมเดลออมนิโมดัลแบบ autoregressive เนทีฟที่มี ultra-sparse MoE รองรับการสร้างแบบร่วมของข้อความ รูปภาพ เสียง และวิดีโอ โมเดลนี้มีอันดับสูงบน leaderboard LMArena<sup>[\[4\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-ERNIE50-4)[\[13\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-LMArena-13)</sup>

## รากฐานทางทฤษฎีและหลักการทางสถาปัตยกรรม

### การมาสก์ความรู้ (Knowledge Masking)

หลักการสำคัญของโมเดลในยุคแรกของชุดนี้คือการผสานความรู้ที่มีโครงสร้างในกระบวนการพรีเทรนผ่านกลยุทธ์การมาสก์ที่ปรับปรุงแล้ว ต่างจาก Masked Language Modeling (MLM) มาตรฐานที่มาสก์โทเค็นแต่ละตัว ERNIE 1.0 จะมาสก์ทั้งหน่วยความหมาย ได้แก่ เอนทิตี (ระบุผ่าน NER) และวลี สำหรับเอนทิตี $E = \{ t_{1},\ldots,t_{k}\}$ ลำดับโทเค็นทั้งหมดจะถูกมาสก์ด้วยความน่าจะเป็น $p$ แนวทางนี้บังคับให้โมเดลพึ่งพาบริบทที่กว้างขึ้นและเรียนรู้ความสัมพันธ์ระหว่างเอนทิตี<sup>[\[1\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-ERNIE1-1)</sup>

ERNIE 2.0 พัฒนาแนวทางนี้ต่อด้วยการเพิ่มการเรียนรู้แบบหลายงานแบบ incremental: งานระดับคำศัพท์ โครงสร้าง และความหมายถูกเพิ่มตามลำดับ โดยความรู้ที่เรียนรู้ก่อนหน้าได้รับการรักษาไว้ด้วยกลไกการพรีเทรนอย่างต่อเนื่อง (continual pre-training)<sup>[\[5\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-ERNIE2-5)</sup>

### กรอบรวมของ ERNIE 3.0

กรอบ ERNIE 3.0 อิงกับการแบ่งสถาปัตยกรรมออกเป็นสองระดับ:<sup>[\[2\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-ERNIE3-2)[\[6\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-Titan-6)</sup>

- **Universal Representation Module** — Transformer-XL แบบหลายชั้นร่วม ฝึกบนงานพรีเทรนที่หลากหลายและดึงคุณลักษณะคำศัพท์และไวยากรณ์ร่วม ในเวอร์ชัน Titan โมดูลนี้มี 48 ชั้น การแทนความหมายที่ซ่อนอยู่ขนาด 12288 หัวความใส่ใจ 192 หัว และขนาด feed-forward ภายใน 196608 ($16 \times d_{model}$)
- **Task-specific Representation Modules** — โมดูลแยกสำหรับ NLU (ความใส่ใจสองทิศทาง) และ NLG (ความใส่ใจทิศทางเดียวพร้อมหน่วยความจำแบบ recurrent ของ Transformer-XL) แต่ละโมดูลมี 12 ชั้น ขนาดเวกเตอร์ที่ซ่อนอยู่ 768 และหัวความใส่ใจ 12 หัว

การผสานรวมพาราไดม์ auto-encoding และ auto-regressive ทำให้โมเดลเดียวให้ผลสูงทั้งบน benchmark NLU (งานแบบ BERT) และ benchmark NLG (งานแบบ GPT)<sup>[\[2\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-ERNIE3-2)</sup>

### Universal Knowledge-Text Prediction (UKTP)

งาน UKTP เป็นกลไกหลักในการผสานความรู้ใน ERNIE 3.0/Titan โมเดลได้รับคู่ (สามส่วนจากกราฟความรู้ ประโยคจากสารานุกรม) และต้องทำนายความสัมพันธ์ในสามส่วนโดยใช้ข้อความ หรือทำนายโทเค็นที่มาสก์ในข้อความโดยใช้ข้อมูลจากสามส่วน<sup>[\[6\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-Titan-6)</sup>

เมื่อทำนายความสัมพันธ์ $r$ ในสามส่วน $(h,r,t)$ ที่เชื่อมโยงกับข้อความ $x$ งานจะถูกกำหนดเป็นการจำแนกแบบหลายคลาส:

$$
P_{\theta}(r \mid x,h,t) = {softmax}(W \cdot f_{\theta}(x,h,t) + b)
$$

โดยที่ $f_{\theta}$ คือผลลัพธ์ของ Transformer ตามตำแหน่งของการกล่าวถึงเอนทิตีหัวและหาง $W,b$ คือพารามิเตอร์ของตัวจำแนก $\theta$ คือพารามิเตอร์ของโมเดล<sup>[\[6\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-Titan-6)</sup>

### กลไกการสร้างที่น่าเชื่อถือและควบคุมได้ (Titan)

ERNIE 3.0 Titan นำเสนอฟังก์ชันความสูญเสียเพิ่มเติมสองประเภท<sup>[\[6\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-Titan-6)</sup>

**Self-supervised adversarial loss** สร้างชุดข้อมูล $D_{a} = \{ D_{\text{original}},D_{\text{generated}}\}$ โดยที่ $D_{\text{original}}$ คือย่อหน้าจริง และ $D_{\text{generated}}$ คือข้อความที่สร้างโดยเวอร์ชันก่อนหน้าของ ERNIE จาก prefix ของย่อหน้าต้นฉบับ งานคือการจำแนกแบบไบนารี (ต้นฉบับ/ที่สร้าง) จากสถานะที่ซ่อนอยู่ของโทเค็นพิเศษ \[CLS\]:

$$
L_{a}(D_{a}) = - \sum\limits_{n = 1}^{|D_{a}|}\log P_{\theta}\left( y_{n} = I_{h_{\lbrack{CLS}\rbrack}^{(n)} \in D_{\text{original}}} \mid h_{\lbrack{CLS}\rbrack}^{(n)} \right)
$$

โดยที่ $h_{\lbrack{CLS}\rbrack}^{(n)}$ คือการแทนเวกเตอร์ \[CLS\] สำหรับตัวอย่างที่ $n$ และ $I_{\cdot}$ คือตัวบ่งชี้การเป็นส่วนหนึ่งของข้อความต้นฉบับ<sup>[\[6\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-Titan-6)</sup>

**Controllable language modeling loss** ตัวอย่างการฝึกแต่ละรายการมาพร้อมชุดแอตทริบิวต์ (prompt) ที่อธิบายแนว หัวข้อ คำสำคัญ อารมณ์ และความยาว ความสูญเสียรวมการสร้างแบบมีเงื่อนไขและไม่มีเงื่อนไข:

$$
L_{c}(D_{c}) = \left\{ \begin{matrix}
{- \sum\limits_{n = 1}^{|D_{c}|}\log P_{\theta}(x_{t}^{(n)} \mid x_{< t}^{(n)}),} & {\text{если~}p \leq 0,5,} \\
{- \sum\limits_{n = 1}^{|D_{c}|}\log P_{\theta}(x_{t}^{(n)} \mid x_{< t}^{(n)},\text{prompts}_{n}),} & {\text{если~}p > 0,5,}
\end{matrix} \right.
$$

โดยที่ $p$ คือตัวแปรสุ่มสำหรับการสลับโหมด $x_{t}^{(n)}$ คือโทเค็นที่ $t$ ของตัวอย่างที่ $n$ คำจำกัดความนี้ป้องกันการพึ่งพา prompt มากเกินไปของโมเดล<sup>[\[6\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-Titan-6)</sup>

### สถาปัตยกรรม MoE แบบเฮเทอโรจีเนียส (ERNIE 4.5)

ERNIE 4.5 นำเสนอสถาปัตยกรรม Mixture-of-Experts แบบมัลติโมดัลเฮเทอโรจีเนียสที่แยกผู้เชี่ยวชาญตามโมดัลลิตี: ผู้เชี่ยวชาญข้อความและผู้เชี่ยวชาญด้านภาพ (ผู้เชี่ยวชาญด้านภาพมีขนาดประมาณ 1/3 ของผู้เชี่ยวชาญข้อความ) การกำหนดเส้นทางโทเค็นทำงานด้วยการแยกโมดัลลิตี (modality-isolated routing) และใช้บทลงโทษ orthogonalization สำหรับเราเตอร์ (router orthogonalization loss สัมประสิทธิ์ระดับ $10^{- 3}$) เพื่อให้แน่ใจว่าผู้เชี่ยวชาญมีความเชี่ยวชาญเฉพาะ ใช้ 3D RoPE (Rotary Position Embeddings) สำหรับการเข้ารหัสตำแหน่งในมิติเวลา ความสูง และความกว้างในข้อมูลวิดีโอ กลไก FlashMask ใช้สำหรับการทำงานอย่างมีประสิทธิภาพกับบริบทยาว (สูงสุด 131,000 โทเค็น) พร้อมมาสก์ $O(N)$<sup>[\[3\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-ERNIE45-3)</sup>

การพรีเทรนของ ERNIE 4.5 ดำเนินการหลายขั้นตอน: เริ่มจากการฝึกบนข้อมูลข้อความเท่านั้น จากนั้นบนข้อมูลภาพ และในระยะสุดท้ายคือการฝึกร่วมบนข้อมูลมัลติโมดัล (text-only → vision-only → joint) สำหรับการประมวลผลรูปภาพใช้ adaptive ViT encoder (Vision Transformer) พร้อมกลไก pixel shuffle สำหรับการปรับการแทนความหมายของโมดัลลิตีต่างๆ รองรับการประมวลผลวิดีโอยาว (สูงสุด 32,000 โทเค็น) ด้วย adaptive frame sampling<sup>[\[3\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-ERNIE45-3)</sup>

### ออมนิโมดัลลิตีแบบเนทีฟ (ERNIE 5.0)

ERNIE 5.0 นำไปใช้การสร้างแบบ autoregressive เนทีฟของหลายโมดัลลิตี (ข้อความ รูปภาพ เสียง วิดีโอ) ในสถาปัตยกรรมเดียวด้วย ultra-sparse MoE ซึ่งในแต่ละขั้นตอนมีผู้เชี่ยวชาญน้อยกว่า 3% ของจำนวนทั้งหมดที่ถูกเปิดใช้งาน ใช้งาน Next-Group-of-Tokens Prediction เป็นงานพรีเทรน — การทำนายกลุ่มโทเค็นแทนที่จะเป็นโทเค็นเดียว ซึ่งเพิ่มประสิทธิภาพการฝึก สำหรับโมดัลลิตีเสียงใช้ hierarchical codec เทคโนโลยี elastic training ช่วยให้สามารถสร้างโมเดลย่อยที่มีความลึก ความกว้าง และระดับความเบาบางที่แตกต่างกันภายในรอบการฝึกเดียว<sup>[\[4\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-ERNIE50-4)</sup>

## งานพรีเทรนและข้อมูล

### งานพรีเทรนของ ERNIE 3.0 / Titan

ใน ERNIE 3.0 และ Titan ใช้กลุ่มงานพรีเทรนดังต่อไปนี้:<sup>[\[2\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-ERNIE3-2)[\[6\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-Titan-6)</sup>

**งานแบบ Word-aware:**

- Knowledge Masked Language Modeling — การมาสก์วลีและเอนทิตีเพื่อเรียนรู้การพึ่งพาในบริบทท้องถิ่นและบริบทกว้าง
- Document Language Modeling — การสร้างแบบ autoregressive ของเอกสารที่มีความยาวสูงสุด 512 โทเค็น โดยใช้หน่วยความจำแบบ recurrent ของ Transformer-XL

**งานแบบ Structure-aware:**

- Sentence Reordering — สำหรับย่อหน้าที่แบ่งแบบสุ่มเป็น $m$ ส่วนและสลับสับเปลี่ยน โมเดลจะแก้งานจำแนกแบบ $k$ คลาสพร้อม $k = \sum\limits_{n = 1}^{m}n!$ เพื่อกู้คืนลำดับเดิม
- Sentence Distance — การจำแนก 3 คลาส: ประโยคที่ติดกัน ประโยคที่ไม่ติดกันในเอกสารเดียวกัน ประโยคจากเอกสารต่างกัน

**งานแบบ Knowledge-aware:**

- Universal Knowledge-Text Prediction (UKTP) — การสร้างข้อความและสามส่วนของกราฟความรู้ร่วม
- Credible and Controllable Generations — การรวม adversarial loss และ controllable LM loss

### ข้อมูลพรีเทรน

สำหรับ ERNIE 3.0/Titan ใช้ ERNIE 3.0 Corpus — คอร์ปัสภาษาจีนขนาดประมาณ 4 TB ใน 11 หมวดหมู่ ได้แก่ เว็บเพจ บันทึกการค้นหา ข้อมูลถามตอบ ข้อความนิยาย กฎหมาย การเงิน และการแพทย์ นิยายและบทกวี เหนือคอร์ปัสมีกราฟความรู้ที่มีข้อเท็จจริงมากกว่า 50 ล้านรายการ<sup>[\[2\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-ERNIE3-2)[\[6\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-Titan-6)</sup>

นอกจากนี้สำหรับ Titan ยังสร้าง:

- **Adversarial dataset** — ย่อหน้าต้นฉบับ 2 ล้านย่อหน้าและย่อหน้า «เชิงลบ» ที่สอดคล้องกัน ซึ่งสร้างโดย ERNIE 3.0 จาก prefix ของ 1–3 ประโยคแรกของต้นฉบับ ความยาวสูงสุด 512 โทเค็น
- **Controllable dataset** — ข้อความที่มีแอตทริบิวต์แนว หัวข้อ (26 หัวข้อ) คำสำคัญ อารมณ์ (positive/negative/neutral) และความยาว แอตทริบิวต์แนวเข้ารหัสด้วย «soft prompt» ที่เรียนรู้ได้ (จำนวน prompt สำหรับแนวเลือกแบบสุ่มจาก 0 ถึง 64)<sup>[\[6\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-Titan-6)</sup>

เวอร์ชันหลัง (ERNIE 4.5, 5.0) ใช้คอร์ปัสมัลติโมดัลที่ขยายเพิ่ม (ข้อความ รูปภาพ วิดีโอ เสียง) รวมถึงเนื้อหาเว็บที่คัดสรร สิ่งพิมพ์ทางวิทยาศาสตร์ และข้อมูลสังเคราะห์<sup>[\[3\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-ERNIE45-3)[\[4\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-ERNIE50-4)</sup>

## การฝึกและการปรับแบบขนาน

การพรีเทรนของ ERNIE 3.0 Titan ดำเนินการโดยใช้ optimizer Adam (อัตราการเรียนรู้ $10^{- 4}$, $\beta_{1} = 0,9$, $\beta_{2} = 0,95$, L2 regularization 0.1, การตัด norm ของ gradient ที่ 1.0) ความยาวบริบทสูงสุด 512 และความยาวหน่วยความจำแบบ recurrent 128 สำหรับงานสร้าง ใช้รูปแบบการฝึกแบบ progressive (การลู่เข้าอย่างรวดเร็วใน 4,000 ขั้นตอนแรก) และการลดอัตราการเรียนรู้แบบเชิงเส้น<sup>[\[6\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-Titan-6)</sup>

### การขนาน 4D แบบไฮบริด

สำหรับการฝึกโมเดลแบบ dense ที่มี 260 พันล้านพารามิเตอร์บนคลัสเตอร์เฮเทอโรจีเนียส (GPU NVIDIA V100 และ NPU Ascend 910) ใน PaddlePaddle มีการนำการขนาน 4D แบบไฮบริดมาใช้:<sup>[\[6\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-Titan-6)</sup>

- **Data parallelism (DP)** — การจำลองโมเดลบนอุปกรณ์หลายเครื่องพร้อมการประมวลผล batch แยกกัน
- **Tensor model parallelism (MP)** — การแบ่งพารามิเตอร์และ activation ของ Transformer ภายในชั้นตามอุปกรณ์
- **Pipeline model parallelism (PP)** — การกระจายชั้นของโมเดลตาม pipeline
- **Group Sharded** — ตัวแปรที่ปรับปรุงแล้วของ ZeRO-like sharded-data-parallel เพื่อลดการซ้ำซ้อนของสถานะ optimizer

รายงานแสดงข้อมูลเกี่ยวกับการขยาย (weak scaling) ไปสู่หลายพัน Ascend 910 cards ด้วยประสิทธิภาพประมาณ 91.7% ในแง่ throughput<sup>[\[6\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-Titan-6)</sup>

### การฝึก ERNIE 4.5

การฝึก ERNIE 4.5 ดำเนินการบนคลัสเตอร์ที่มี GPU NVIDIA H800 จำนวน 2,016 หน่วย โดยบรรลุ Model FLOPs Utilization (MFU) ที่ 47% ใช้ความแม่นยำแบบผสม FP8 จุดตรวจสอบที่ทนต่อความผิดพลาด (Zero Cost Checkpoint กู้คืนในเวลาน้อยกว่า 8 นาที) และการฝึกแบบ progressive พร้อมการเพิ่มความยาวลำดับและขนาด batch<sup>[\[3\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-ERNIE45-3)</sup>

### การดิสทิลล์แบบออนไลน์

เพื่อลดข้อกำหนดด้านทรัพยากรการคำนวณในการติดตั้งใช้งาน ERNIE 3.0 Titan ใช้การดิสทิลล์แบบออนไลน์ ซึ่งโมเดลครูและโมเดลนักเรียนหลายตัวฝึกพร้อมกัน:<sup>[\[6\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-Titan-6)</sup>

- **On-the-Fly Distillation (OFD)** — logit และการแทนความหมายของครูถูกใช้สำหรับการฝึกนักเรียนในขั้นตอนการฝึกเดียวกัน
- **Teacher assistants** — โมเดลขนาดกลางที่ลดช่องว่างความสามารถระหว่างครูและนักเรียนขั้นสุดท้าย
- **Auxiliary Layer Distillation (ALD)** — ชั้นเพิ่มเติมในนักเรียนที่ถูกละทิ้งระหว่าง fine-tuning เพื่อให้การแทนความหมายภายในตรงกันดีขึ้น

## การฝึกหลังพรีเทรนและการปรับพฤติกรรม

เวอร์ชันเชิงพาณิชย์ของ ERNIE (เริ่มจาก ERNIE Bot) ผ่านขั้นตอนการฝึกหลังพรีเทรน:<sup>[\[7\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-Wiki_ErnieBot-7)[\[3\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-ERNIE45-3)</sup>

- **Supervised Fine-Tuning (SFT)** — การ fine-tune บนข้อมูลคำสั่งที่มีป้ายกำกับ (ใน ERNIE 4.5 มี 2.3 ล้านตัวอย่าง)
- **Reinforcement Learning from Human Feedback (RLHF)** — การปรับพฤติกรรมของโมเดลด้วยข้อมูลป้อนกลับจากมนุษย์ ใช้อัลกอริทึม PPO (Proximal Policy Optimization) และ DPO (Direct Preference Optimization)
- **Unified Preference Optimization (UPO)** — วิธีการปรับแบบรวมความชอบ ที่นำมาใช้ใน ERNIE 4.5
- **Reinforcement Learning with Verifiers (RLVR)** — การเรียนรู้แบบเสริมแรงโดยใช้ตัวตรวจสอบเพื่อตรวจสอบความถูกต้องของคำตอบ ใช้วิธี GRPO (Group Relative Policy Optimization)
- **โหมดการใช้เหตุผล** (thinking modes) — โมเดล 4.5 รองรับโหมดที่มีการไตร่ตรอง (reflection, planning) และโหมดที่ไม่มี

## ผลลัพธ์หลักและ benchmark

### ตารางสรุปวิวัฒนาการของโมเดล

| เวอร์ชัน          | ปี    | พารามิเตอร์                           | สถาปัตยกรรม                            | Benchmark หลัก                                        | แหล่งที่มา                                                                                      |
|-----------------|------|-------------------------------------|---------------------------------------|------------------------------------------------------|----------------------------------------------------------------------------------------------|
| ERNIE 1.0       | 2019 | ~110 ล้าน (base)                     | Transformer encoder (แบบ BERT)        | XNLI 78.4%; MSRA-NER F1 93.8%                        | <sup>[\[1\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-ERNIE1-1)</sup>  |
| ERNIE 2.0       | 2019 | ~110–340 ล้าน                        | Continual multi-task                  | GLUE 80.6 (base) / 83.6 (large); 16 งาน SOTA         | <sup>[\[5\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-ERNIE2-5)</sup>  |
| ERNIE 3.0       | 2021 | 10 พันล้าน                            | Unified Transformer-XL (AE+AR)        | SuperGLUE 90.6% (\> human 89.8%); 54 งาน SOTA ภาษาจีน | <sup>[\[2\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-ERNIE3-2)</sup>  |
| ERNIE 3.0 Titan | 2021 | 260 พันล้าน (dense)                   | Dense + controllable generation       | 68 ชุดข้อมูล SOTA; zero/few-shot                        | <sup>[\[6\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-Titan-6)</sup>   |
| ERNIE 4.5       | 2025 | 0.3–424 พันล้าน (MoE, 47 พันล้านที่ใช้งาน) | Heterogeneous multimodal MoE          | C-Eval 90.6%; MMLU 86.5%; MMMU 67.3–70%              | <sup>[\[3\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-ERNIE45-3)</sup> |
| ERNIE 5.0       | 2026 | ~2.4 ล้านล้าน (ultra-sparse MoE)      | Unified autoregressive multimodal MoE | LMArena Elo ~1460 (ท็อป 10 ระดับโลก)                   | <sup>[\[4\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-ERNIE50-4)</sup> |

### ผลลัพธ์ของ ERNIE 3.0 และ Titan

ERNIE 3.0 (10 พันล้านพารามิเตอร์) บน benchmark SuperGLUE ภาษาอังกฤษได้ผลเฉลี่ย 90.6 ซึ่งสูงกว่าค่าอ้างอิงของมนุษย์ (89.8) และสูงกว่า GPT-3, T5 และ DeBERTa ณ เวลาที่เผยแพร่ ERNIE 3.0 Titan (260 พันล้าน) ถูกประเมินบน 68 ชุดข้อมูล รวมถึงงานจำแนก NLI QA และการสร้าง; ผู้เขียนรายงานว่าทำผลได้เหนือกว่าโมเดลก่อนหน้าในทั้ง 68 ชุดข้อมูล ผลลัพธ์เหล่านี้เป็นข้อมูลจากแหล่งต้นฉบับ; การทำซ้ำโดยอิสระได้รับการอธิบายไว้อย่างจำกัด<sup>[\[2\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-ERNIE3-2)[\[6\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-Titan-6)</sup>

### ผลลัพธ์ของ ERNIE 4.5

ตามรายงานทางเทคนิคปี 2025 ERNIE 4.5 (300B-A47B หลัง post-training) แสดงผลลัพธ์ดังต่อไปนี้บน benchmark ข้อความและมัลติโมดัล:<sup>[\[3\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-ERNIE45-3)</sup>

| Benchmark | ERNIE-4.5-300B-A47B | เงื่อนไข                  |
|-----------|---------------------|-------------------------|
| C-Eval    | 90.6%               | 5-shot                  |
| CMMLU     | 90.2%               | —                       |
| MMLU      | 86.5%               | มาตรฐาน                 |
| IFEval    | 88.0%               | instruction following   |
| GSM8K     | 91.8%               | —                       |
| MMMU      | 67.3–70.0%          | non-thinking / thinking |
| MathVista | 78.8%               | thinking mode           |
| OCRBench  | 883                 | —                       |

ตามรายงาน ERNIE 4.5 ทำผลได้เหนือกว่า DeepSeek-V3 ใน 22 จาก 28 benchmark; ตัวแปรมัลติโมดัล (ERNIE-4.5-VL-424B-A47B) แสดงผลที่สามารถแข่งขันได้บนงานการใช้เหตุผลเชิงภาพ ในงานการใช้เหตุผลเชิงภาพบางส่วนและการตีความรูปภาพทางเทคนิค (การวิเคราะห์แผนผัง ไดอะแกรมวิศวกรรม) รายงานผลสูงกว่าโมเดล GPT และ Gemini บางรุ่น<sup>[\[3\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-ERNIE45-3)[\[14\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-AINmultimodal-14)</sup>

**การเปรียบเทียบ ERNIE 4.5 กับคู่แข่ง** (benchmark คัดสรร หลัง post-training ตามรายงานทางเทคนิคปี 2025):

| Benchmark | ERNIE-4.5-300B-A47B | DeepSeek-V3-671B-A37B | Qwen3-235B-A22B    | หมายเหตุ                 |
|-----------|---------------------|-----------------------|--------------------|-------------------------|
| C-Eval    | 90.6%               | —                     | —                  | 5-shot                  |
| MMLU      | 86.5%               | ใกล้เคียงกัน             | —                  | มาตรฐาน                 |
| IFEval    | 88.0%               | —                     | 83.2%              | instruction following   |
| MMMU      | 67.3–70.0%          | —                     | —                  | thinking / non-thinking |
| MathVista | 78.8%               | —                     | 77.6% (Qwen2.5-VL) | thinking mode           |

เงื่อนไขการทำซ้ำ: โปรโตคอลมาตรฐาน (5-shot / zero-shot); ชุดข้อมูลเปิด (C-Eval 2023+, MMLU, MMMU) เครื่องหมาย «—» หมายความว่าข้อมูลที่เทียบเคียงได้ในเงื่อนไขเดียวกันไม่ได้ระบุไว้ในรายงาน<sup>[\[3\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-ERNIE45-3)</sup>

### การประเมิน ERNIE Bot 4.0

รายงาน SuperBench ของมหาวิทยาลัยชิงหัวจัดอันดับ LLM เชิงพาณิชย์ตามชุดงาน (ความเข้าใจภาษา คณิตศาสตร์ การเขียนโปรแกรม หลายงาน) ERNIE Bot 4.0 ได้อันดับหนึ่งในบรรดาโมเดลจีน นำหน้า GLM-4 (Zhipu AI) ประมาณ 0.41 คะแนนตามตัวชี้วัดรวม; โมเดล GPT-4 และ Anthropic Claude-3 ยังคงอยู่ในอันดับสูงกว่าในการจัดอันดับระดับโลก ERNIE Bot 4.0 แสดงผลที่ดีในความเข้าใจข้อความภาษาจีนและการปฏิบัติตามคำสั่ง ขณะที่มีคะแนนต่ำกว่าในการเขียนโปรแกรมและงานภาษาอังกฤษบางงาน<sup>[\[9\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-Tsinghua-9)[\[10\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-Multiplatform-10)</sup>

## การประยุกต์ใช้งาน

### ผลิตภัณฑ์และบริการของ Baidu

โมเดลในชุด ERNIE ถูกผสานรวมเข้ากับระบบนิเวศผลิตภัณฑ์ของ Baidu:<sup>[\[7\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-Wiki_ErnieBot-7)[\[8\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-BaiduBlog-8)[\[3\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-ERNIE45-3)</sup>

- **ERNIE Bot (Wenxin Yiyan)** — แชทบอทที่รองรับการสร้างแบบมัลติโมดัล (ข้อความ รูปภาพ วิดีโอ เสียง) ตามข้อมูลของ Baidu จำนวนผู้ใช้งานรายเดือนที่ใช้งานอยู่เกิน 200 ล้านราย (2024–2025) ตั้งแต่ปี 2025 ERNIE Bot ให้บริการฟรี<sup>[\[7\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-Wiki_ErnieBot-7)</sup>
- **Baidu Search** — การสร้างคำตอบและฟีเจอร์ AI ในผลการค้นหา; ตามข้อมูลของ Baidu สูงสุด 70% ของผลลัพธ์บนสุดได้รับการเพิ่มคุณค่าด้วยองค์ประกอบ AI
- **Qianfan (แพลตฟอร์ม MaaS)** — API สำหรับลูกค้าองค์กร; ตามข้อมูลของ Baidu มีองค์กรมากกว่า 760,000 แห่งใช้แพลตฟอร์มนี้ จำนวนการเรียก API ต่อวันเกิน 1.5 พันล้านครั้ง (2024) ในจำนวนลูกค้ามี Lenovo, Trip.com และอื่นๆ<sup>[\[7\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-Wiki_ErnieBot-7)</sup>
- **Comate** — AI assistant สำหรับการเขียนโปรแกรม
- **Wenxin Yige** — การสร้างรูปภาพบนพื้นฐานโมเดล ERNIE-ViLG<sup>[\[11\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-ViLG-11)</sup>
- การผสานรวมกับพันธมิตรภายนอก: Samsung Galaxy (สำหรับตลาดจีน) อวตารดิจิทัล ปลั๊กอิน (การค้นหา การวิเคราะห์ไฟล์)<sup>[\[7\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-Wiki_ErnieBot-7)</sup>

### การประยุกต์ใช้ในอุตสาหกรรม

โมเดลถูกนำไปใช้ในด้านต่อไปนี้:<sup>[\[7\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-Wiki_ErnieBot-7)[\[3\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-ERNIE45-3)</sup>

- ระบบถามตอบในโดเมนเฉพาะ (กฎหมาย การแพทย์ การเงิน) โดยใช้การพรีเทรนที่เสริมด้วยความรู้
- การสร้างและแก้ไขข้อความภาษาจีน (ข่าว ข้อความการตลาด เนื้อหาสร้างสรรค์)
- งานมัลติโมดัล: การวิเคราะห์รูปภาพ แผนผังวิศวกรรม วิดีโอ และข้อมูลตาราง (ในเวอร์ชัน 4.5 และสูงกว่า)
- การศึกษา (การเรียนรู้เฉพาะบุคคล) หุ่นยนต์ (การวางแผนงาน) การขับขี่อัตโนมัติ (Apollo)

### โค้ดเปิด

เริ่มจาก ERNIE 4.5 ทั้ง 10 รูปแบบของชุดโมเดลเผยแพร่ภายใต้ลิขสิทธิ์ Apache 2.0 พร้อม ERNIEKit (รองรับ SFT, LoRA, DPO, inference บน PaddlePaddle/FastDeploy) โค้ดของเวอร์ชันก่อนหน้า (ERNIE 1.0–3.0) มีให้บน GitHub PaddlePaddle/ERNIE<sup>[\[3\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-ERNIE45-3)[\[15\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-GH-15)</sup>

## ข้อจำกัดและปัญหาที่ยังเปิดอยู่

### ข้อจำกัดด้านสถาปัตยกรรมและชุดข้อมูล

ERNIE 3.0 Titan ที่มี 260 พันล้านพารามิเตอร์มีข้อจำกัดความยาวบริบทที่ 512 โทเค็นสำหรับโมดูลแต่ละตัว ซึ่งจำกัดการสร้างแบบจำลองข้อความยาวโดยไม่มีกลไกเพิ่มเติม (chunking, หน่วยความจำภายนอก) การฝึกดำเนินการส่วนใหญ่บนคอร์ปัสภาษาจีน ซึ่งทำให้คุณภาพระหว่างภาษาจีนและภาษาอื่นไม่เท่ากัน<sup>[\[6\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-Titan-6)</sup>

การผสานกราฟความรู้ช่วยปรับปรุงการทำงานกับข้อเท็จจริงที่มีโครงสร้าง แต่ความแม่นยำและความสมบูรณ์ของความรู้ถูกจำกัดด้วยคุณภาพของกราฟเองและขั้นตอนการจับคู่ «สามส่วน–ข้อความ» (entity linking, relation alignment) ซึ่งในบางกรณีนำไปสู่การเชื่อมโยงที่ผิดพลาดหรือไม่สมบูรณ์<sup>[\[6\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-Titan-6)</sup>

### การสร้างข้อมูลเท็จและความน่าเชื่อถือ

Adversarial loss และ controllable LM loss ช่วยปรับปรุงความทนทานต่อการสร้างที่ไม่น่าเชื่อถือ แต่ไม่สามารถขจัดปัญหาการสร้างข้อมูลเท็จ (การสร้างข้อความที่ผิดทางข้อเท็จจริง) ได้อย่างสมบูรณ์ พารามิเตอร์ของ adversarial classifier ฝึกบนข้อมูลที่ข้อความ «ที่สร้าง» ถูกสร้างโดยเวอร์ชันก่อนหน้าของ ERNIE ซึ่งจำกัดขอบเขตของรูปแบบข้อผิดพลาดที่รู้จัก<sup>[\[6\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-Titan-6)</sup>

### อคติทางสังคม

งานวิจัยที่ตีพิมพ์ใน PeerJ Computer Science (2025) วิเคราะห์อคติทางสังคมใน LLM ภาษาจีน (ERNIE และ Qwen) โดยใช้กลุ่มสังคม 240 กลุ่มและคำอธิบายที่สร้างขึ้นมากกว่า 30,000 รายการ ผลลัพธ์แสดงว่า ERNIE สร้างเนื้อหาเชิงลบและมีแบบแผนน้อยกว่า Baidu Search หรือ Qwen (ประมาณ 1/10 ของคำที่มีความหมายเชิงลบใน ERNIE เทียบกับ 1/3 ใน Qwen) อย่างไรก็ตาม ยังคงมีแบบแผนบางส่วน รวมถึงคำอธิบายที่อาจสร้างความไม่พอใจ<sup>[\[16\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-Bias1-16)[\[17\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-Bias2-17)</sup>

### การทำซ้ำ

บางส่วนของโมเดลในชุด (ERNIE 3.0 Titan, ERNIE Bot 4.0 และ 5.0 เชิงพาณิชย์) ไม่มีให้ในรูปแบบโค้ดเปิดและน้ำหนักโมเดลอย่างสมบูรณ์ ซึ่งจำกัดการทำซ้ำ benchmark และความเป็นไปได้ในการประเมินอิสระ การเผยแพร่ ERNIE 4.5 ภายใต้ Apache 2.0 ช่วยปรับปรุงสถานการณ์ แต่สถาปัตยกรรมและการตั้งค่าการฝึกอาจแตกต่างจากที่อธิบายในสิ่งพิมพ์ยุคแรก<sup>[\[3\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-ERNIE45-3)[\[6\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-Titan-6)</sup>

### ความปลอดภัยทางการแพทย์

การศึกษาการใช้ ERNIE Bot ในสถานการณ์ทางการแพทย์ (Si et al., 2025) พบแนวโน้มการสั่งจ่ายมากเกินความจำเป็น: 91.9% ของการทดสอบที่ไม่จำเป็นและ 57.8% ของยาที่ไม่จำเป็นในสถานการณ์จำลอง รวมถึงความไม่สมดุลด้านอายุและเศรษฐกิจสังคมในคำแนะนำ<sup>[\[18\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-Medical-18)</sup>

## แง่มุมด้านจริยธรรมและกฎระเบียบ

โมเดลในชุด ERNIE ดำเนินการภายใต้กฎระเบียบ AI สร้างสรรค์ของจีน โดยเฉพาะ «มาตรการชั่วคราวสำหรับการจัดการบริการ AI สร้างสรรค์» (Interim Measures for Generative AI Services, 2023) ซึ่งกำหนดให้ต้องมีการประเมินความปลอดภัย การลงทะเบียนอัลกอริทึม และข้อจำกัดด้านเนื้อหา<sup>[\[7\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-Wiki_ErnieBot-7)[\[17\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-Bias2-17)</sup>

ERNIE Bot แสดงระดับการปฏิเสธสูงต่อคำขอที่เกี่ยวข้องกับหัวข้อทางการเมืองที่ละเอียดอ่อน ตามกฎหมายของประเทศ งานวิจัย (Pan et al., 2026) วิเคราะห์การเซ็นเซอร์ทางการเมืองใน LLM ที่มีต้นกำเนิดจากจีน รวมถึงโมเดลของ Baidu<sup>[\[19\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-Censorship-19)</sup>

สิ่งพิมพ์ของ Baidu ไม่ได้อธิบายรายละเอียดขั้นตอนการตรวจสอบพฤติกรรมโมเดล เกณฑ์การกรองเนื้อหา และความสมดุลระหว่างข้อกำหนดกฎระเบียบท้องถิ่นกับหลักการทั่วไปด้านจริยธรรม AI เสมอไป ซึ่งยังคงเป็นพื้นที่เปิดสำหรับการวิจัยอิสระ<sup>[\[17\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-Bias2-17)</sup>

## แนวโน้มและทิศทางการวิจัย

จากรายงานทางเทคนิคและประกาศของ Baidu มีทิศทางการพัฒนาชุดโมเดล ERNIE ดังต่อไปนี้:

- การขยายมัลติโมดัลลิตีเพิ่มเติม (ข้อความ–รูปภาพ–วิดีโอ–เสียง) รวมถึงการประมวลผลข้อมูลภาพทางเทคนิคที่หนาแน่น (แผนผังวิศวกรรม รูปภาพทางการแพทย์)<sup>[\[3\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-ERNIE45-3)[\[4\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-ERNIE50-4)</sup>
- การรวม dense และสถาปัตยกรรม MoE เพื่อสมดุลระหว่างคุณภาพและประสิทธิภาพการคำนวณสำหรับขนาดโมเดลรวมขนาดใหญ่มาก<sup>[\[3\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-ERNIE45-3)</sup>
- การพัฒนาระบบ agentic (GenFlow, Famou) และเครื่องมือสำหรับการสร้างแบบมีโครงสร้าง (JSON, API calls) เพื่อการผสานรวมใน workflow การผลิต<sup>[\[3\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-ERNIE45-3)</sup>
- การเพิ่มประสิทธิภาพการฝึก: elastic training, การปรับปรุงการขยาย MoE (MFU), การควอนไทเซชัน (W4A8, 2-bit สำหรับการติดตั้งบน GPU เดียว)<sup>[\[3\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-ERNIE45-3)</sup>
- การวิจัยเพื่อลดอคติใน LLM ภาษาจีนโดยคำนึงถึงแง่มุมเฉพาะทางวัฒนธรรม และการพัฒนาวิธีการ benchmark สำหรับภาษาจีนและงานมัลติโมดัล<sup>[\[16\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-Bias1-16)[\[17\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-Bias2-17)</sup>
- การปรับปรุงการใช้เหตุผลในบริบทยาว การเรียนรู้แบบเสริมแรงที่ตรวจสอบได้ (verifiable RL) และการปรับให้เข้ากับโดเมนที่มีข้อมูลจำกัดผ่านคอร์ปัสสังเคราะห์<sup>[\[3\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-ERNIE45-3)[\[4\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-ERNIE50-4)</sup>

## การเปรียบเทียบกับ LLM จีนอื่นๆ

ERNIE แข่งขันกับ LLM จีนขนาดใหญ่หลายรุ่น ได้แก่ Qwen (Alibaba), GLM / ChatGLM (Zhipu AI), DeepSeek (DeepSeek AI) และ Tongyi Qianwen จุดแตกต่างหลักของชุดโมเดล ERNIE คือการเน้นการผสานกราฟความรู้ในการพรีเทรน (knowledge enhancement) การผสานรวมอย่างใกล้ชิดกับระบบนิเวศ Baidu (การค้นหา คลาวด์ API) และการมุ่งเน้นที่แพลตฟอร์ม PaddlePaddle ตามผลการจัดอันดับอิสระ (SuperBench, LMArena) โมเดลในชุด ERNIE ได้อันดับสูงในบรรดา LLM จีน โดยเฉพาะในงานความเข้าใจและการปฏิบัติตามคำสั่งภาษาจีน<sup>[\[9\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-Tsinghua-9)[\[13\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-LMArena-13)[\[16\]](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_note-Bias1-16)</sup>

## ดูเพิ่มเติม

- BERT
- สถาปัตยกรรม Transformer
- RLHF

## บรรณานุกรม

- Sun, Y. et al. (2019). *ERNIE: Enhanced Representation through Knowledge Integration*. arXiv:1904.09223. <a href="https://arxiv.org/abs/1904.09223" class="external free" rel="nofollow">https://arxiv.org/abs/1904.09223</a>
- Sun, Y. et al. (2019). *ERNIE 2.0: A Continual Pre-training Framework for Language Understanding*. AAAI 2020. arXiv:1907.12412. <a href="https://arxiv.org/abs/1907.12412" class="external free" rel="nofollow">https://arxiv.org/abs/1907.12412</a>
- Sun, Y. et al. (2021). *ERNIE 3.0: Large-scale Knowledge Enhanced Pre-training for Language Understanding and Generation*. arXiv:2107.02137. <a href="https://arxiv.org/abs/2107.02137" class="external free" rel="nofollow">https://arxiv.org/abs/2107.02137</a>
- Wang, S. et al. (2021). *ERNIE 3.0 Titan: Exploring Larger-scale Knowledge Enhanced Pre-training for Language Understanding and Generation*. arXiv:2112.12731. <a href="https://arxiv.org/abs/2112.12731" class="external free" rel="nofollow">https://arxiv.org/abs/2112.12731</a>
- Feng, Z. et al. (2022). *ERNIE-ViLG 2.0: Improving Text-to-Image Diffusion Model with Knowledge-Enhanced Mixture-of-Denoising-Experts*. arXiv:2210.15257. <a href="https://arxiv.org/abs/2210.15257" class="external free" rel="nofollow">https://arxiv.org/abs/2210.15257</a>
- Baidu ERNIE Team (2025). *ERNIE 4.5 Technical Report*. <a href="https://ernie.baidu.com/blog/publication/ERNIE_Technical_Report.pdf" class="external free" rel="nofollow">https://ernie.baidu.com/blog/publication/ERNIE_Technical_Report.pdf</a>
- Wang, H. et al. (2026). *ERNIE 5.0 Technical Report*. arXiv:2602.04705. <a href="https://arxiv.org/abs/2602.04705" class="external free" rel="nofollow">https://arxiv.org/abs/2602.04705</a>
- Song, X. et al. (2024). *Comparing diversity, negativity, and stereotypes in Chinese-language AI technologies*. arXiv:2408.15696. <a href="https://arxiv.org/abs/2408.15696" class="external free" rel="nofollow">https://arxiv.org/abs/2408.15696</a>
- PeerJ Computer Science (2025). *Comparing diversity, negativity, and stereotypes in Chinese-language AI technologies*. <a href="https://peerj.com/articles/cs-2694/" class="external free" rel="nofollow">https://peerj.com/articles/cs-2694/</a>
- Pan et al. (2026). *Political censorship in large language models originating from China*. PNAS Nexus.
- Si, Y. et al. (2025). *Quality safety and disparity of an AI chatbot*. PMC.
- Vaswani, A. et al. (2017). *Attention Is All You Need*. NeurIPS. <a href="https://arxiv.org/abs/1706.03762" class="external free" rel="nofollow">https://arxiv.org/abs/1706.03762</a>
- Devlin, J. et al. (2019). *BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding*. NAACL. <a href="https://arxiv.org/abs/1810.04805" class="external free" rel="nofollow">https://arxiv.org/abs/1810.04805</a>
- DaoInsights (2024). *Baidu's ERNIE bot tops the Tsinghua University LLM report ranking*. <a href="https://daoinsights.com/news/far-ahead-baidus-ernie-bot-tops-the-tsinghua-university-llm-report-ranking/" class="external free" rel="nofollow">https://daoinsights.com/news/far-ahead-baidus-ernie-bot-tops-the-tsinghua-university-llm-report-ranking/</a>
- Multiplatform.ai (2024). *ERNIE Bot Leads Tsinghua University's LLM Report Rankings in China*. <a href="https://multiplatform.ai/ernie-bot-leads-tsinghua-universitys-llm-report-rankings-in-china/" class="external free" rel="nofollow">https://multiplatform.ai/ernie-bot-leads-tsinghua-universitys-llm-report-rankings-in-china/</a>
- ArtificialIntelligence-News (2025). *Baidu ERNIE multimodal AI beats GPT and Gemini in benchmarks*. <a href="https://www.artificialintelligence-news.com/news/baidu-ernie-multimodal-ai-gpt-and-gemini-benchmarks/" class="external free" rel="nofollow">https://www.artificialintelligence-news.com/news/baidu-ernie-multimodal-ai-gpt-and-gemini-benchmarks/</a>
- ERNIE Official Blog (2025). *ERNIE-5.0-Preview-1022 now ranks \#2 globally on the LMArena Text leaderboard*. <a href="https://ernie.baidu.com/blog/posts/ernie-5.0-preview-1022-release-on-lmarena/" class="external free" rel="nofollow">https://ernie.baidu.com/blog/posts/ernie-5.0-preview-1022-release-on-lmarena/</a>
- ERNIE Official Blog (2025). *Announcing the Open Source Release of the ERNIE 4.5 Model Family*. <a href="https://ernie.baidu.com/blog/posts/ernie4.5/" class="external free" rel="nofollow">https://ernie.baidu.com/blog/posts/ernie4.5/</a>

## หมายเหตุ

1.  <span id="cite_note-ERNIE1-1">↑ <sup>[1.0](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-ERNIE1_1-0)</sup> <sup>[1.1](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-ERNIE1_1-1)</sup> <sup>[1.2](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-ERNIE1_1-2)</sup> <sup>[1.3](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-ERNIE1_1-3)</sup> <sup>[1.4](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-ERNIE1_1-4)</sup> <sup>[1.5](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-ERNIE1_1-5)</sup> Sun, Y. et al. (2019). *ERNIE: Enhanced Representation through Knowledge Integration*. arXiv:1904.09223. <a href="https://arxiv.org/abs/1904.09223" class="external free" rel="nofollow">https://arxiv.org/abs/1904.09223</a></span>
2.  <span id="cite_note-ERNIE3-2">↑ <sup>[2.00](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-ERNIE3_2-0)</sup> <sup>[2.01](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-ERNIE3_2-1)</sup> <sup>[2.02](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-ERNIE3_2-2)</sup> <sup>[2.03](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-ERNIE3_2-3)</sup> <sup>[2.04](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-ERNIE3_2-4)</sup> <sup>[2.05](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-ERNIE3_2-5)</sup> <sup>[2.06](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-ERNIE3_2-6)</sup> <sup>[2.07](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-ERNIE3_2-7)</sup> <sup>[2.08](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-ERNIE3_2-8)</sup> <sup>[2.09](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-ERNIE3_2-9)</sup> Sun, Y. et al. (2021). *ERNIE 3.0: Large-scale Knowledge Enhanced Pre-training for Language Understanding and Generation*. arXiv:2107.02137. <a href="https://arxiv.org/abs/2107.02137" class="external free" rel="nofollow">https://arxiv.org/abs/2107.02137</a></span>
3.  <span id="cite_note-ERNIE45-3">↑ <sup>[3.00](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-ERNIE45_3-0)</sup> <sup>[3.01](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-ERNIE45_3-1)</sup> <sup>[3.02](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-ERNIE45_3-2)</sup> <sup>[3.03](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-ERNIE45_3-3)</sup> <sup>[3.04](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-ERNIE45_3-4)</sup> <sup>[3.05](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-ERNIE45_3-5)</sup> <sup>[3.06](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-ERNIE45_3-6)</sup> <sup>[3.07](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-ERNIE45_3-7)</sup> <sup>[3.08](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-ERNIE45_3-8)</sup> <sup>[3.09](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-ERNIE45_3-9)</sup> <sup>[3.10](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-ERNIE45_3-10)</sup> <sup>[3.11](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-ERNIE45_3-11)</sup> <sup>[3.12](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-ERNIE45_3-12)</sup> <sup>[3.13](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-ERNIE45_3-13)</sup> <sup>[3.14](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-ERNIE45_3-14)</sup> <sup>[3.15](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-ERNIE45_3-15)</sup> <sup>[3.16](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-ERNIE45_3-16)</sup> <sup>[3.17](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-ERNIE45_3-17)</sup> <sup>[3.18](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-ERNIE45_3-18)</sup> <sup>[3.19](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-ERNIE45_3-19)</sup> Baidu ERNIE Team (2025). *ERNIE 4.5 Technical Report*. <a href="https://ernie.baidu.com/blog/publication/ERNIE_Technical_Report.pdf" class="external free" rel="nofollow">https://ernie.baidu.com/blog/publication/ERNIE_Technical_Report.pdf</a></span>
4.  <span id="cite_note-ERNIE50-4">↑ <sup>[4.0](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-ERNIE50_4-0)</sup> <sup>[4.1](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-ERNIE50_4-1)</sup> <sup>[4.2](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-ERNIE50_4-2)</sup> <sup>[4.3](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-ERNIE50_4-3)</sup> <sup>[4.4](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-ERNIE50_4-4)</sup> <sup>[4.5](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-ERNIE50_4-5)</sup> <sup>[4.6](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-ERNIE50_4-6)</sup> Wang, H. et al. (2026). *ERNIE 5.0 Technical Report*. arXiv:2602.04705. <a href="https://arxiv.org/abs/2602.04705" class="external free" rel="nofollow">https://arxiv.org/abs/2602.04705</a></span>
5.  <span id="cite_note-ERNIE2-5">↑ <sup>[5.0](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-ERNIE2_5-0)</sup> <sup>[5.1](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-ERNIE2_5-1)</sup> <sup>[5.2](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-ERNIE2_5-2)</sup> <sup>[5.3](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-ERNIE2_5-3)</sup> <sup>[5.4](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-ERNIE2_5-4)</sup> Sun, Y. et al. (2019). *ERNIE 2.0: A Continual Pre-training Framework for Language Understanding*. arXiv:1907.12412. AAAI 2020. <a href="https://arxiv.org/abs/1907.12412" class="external free" rel="nofollow">https://arxiv.org/abs/1907.12412</a></span>
6.  <span id="cite_note-Titan-6">↑ <sup>[6.00](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-Titan_6-0)</sup> <sup>[6.01](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-Titan_6-1)</sup> <sup>[6.02](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-Titan_6-2)</sup> <sup>[6.03](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-Titan_6-3)</sup> <sup>[6.04](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-Titan_6-4)</sup> <sup>[6.05](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-Titan_6-5)</sup> <sup>[6.06](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-Titan_6-6)</sup> <sup>[6.07](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-Titan_6-7)</sup> <sup>[6.08](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-Titan_6-8)</sup> <sup>[6.09](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-Titan_6-9)</sup> <sup>[6.10](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-Titan_6-10)</sup> <sup>[6.11](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-Titan_6-11)</sup> <sup>[6.12](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-Titan_6-12)</sup> <sup>[6.13](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-Titan_6-13)</sup> <sup>[6.14](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-Titan_6-14)</sup> <sup>[6.15](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-Titan_6-15)</sup> <sup>[6.16](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-Titan_6-16)</sup> <sup>[6.17](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-Titan_6-17)</sup> <sup>[6.18](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-Titan_6-18)</sup> <sup>[6.19](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-Titan_6-19)</sup> <sup>[6.20](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-Titan_6-20)</sup> Wang, S. et al. (2021). *ERNIE 3.0 Titan: Exploring Larger-scale Knowledge Enhanced Pre-training for Language Understanding and Generation*. arXiv:2112.12731. <a href="https://arxiv.org/abs/2112.12731" class="external free" rel="nofollow">https://arxiv.org/abs/2112.12731</a></span>
7.  <span id="cite_note-Wiki_ErnieBot-7">↑ <sup>[7.0](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-Wiki_ErnieBot_7-0)</sup> <sup>[7.1](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-Wiki_ErnieBot_7-1)</sup> <sup>[7.2](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-Wiki_ErnieBot_7-2)</sup> <sup>[7.3](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-Wiki_ErnieBot_7-3)</sup> <sup>[7.4](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-Wiki_ErnieBot_7-4)</sup> <sup>[7.5](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-Wiki_ErnieBot_7-5)</sup> <sup>[7.6](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-Wiki_ErnieBot_7-6)</sup> <sup>[7.7](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-Wiki_ErnieBot_7-7)</sup> Wikipedia. *Ernie Bot*. <a href="https://en.wikipedia.org/wiki/Ernie_Bot" class="external free" rel="nofollow">https://en.wikipedia.org/wiki/Ernie_Bot</a></span>
8.  <span id="cite_note-BaiduBlog-8">↑ <sup>[8.0](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-BaiduBlog_8-0)</sup> <sup>[8.1](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-BaiduBlog_8-1)</sup> Baidu Research Blog (2023). *ERNIE Bot: Baidu's Knowledge-Enhanced Large Language Model*. <a href="https://research.baidu.com/Blog/index-view?id=183" class="external free" rel="nofollow">https://research.baidu.com/Blog/index-view?id=183</a></span>
9.  <span id="cite_note-Tsinghua-9">↑ <sup>[9.0](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-Tsinghua_9-0)</sup> <sup>[9.1](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-Tsinghua_9-1)</sup> <sup>[9.2](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-Tsinghua_9-2)</sup> DaoInsights (2024). *Baidu's ERNIE bot tops the Tsinghua University LLM report ranking*. <a href="https://daoinsights.com/news/far-ahead-baidus-ernie-bot-tops-the-tsinghua-university-llm-report-ranking/" class="external free" rel="nofollow">https://daoinsights.com/news/far-ahead-baidus-ernie-bot-tops-the-tsinghua-university-llm-report-ranking/</a></span>
10. <span id="cite_note-Multiplatform-10">↑ <sup>[10.0](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-Multiplatform_10-0)</sup> <sup>[10.1](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-Multiplatform_10-1)</sup> Multiplatform.ai (2024). *ERNIE Bot Leads Tsinghua University's LLM Report Rankings in China*. <a href="https://multiplatform.ai/ernie-bot-leads-tsinghua-universitys-llm-report-rankings-in-china/" class="external free" rel="nofollow">https://multiplatform.ai/ernie-bot-leads-tsinghua-universitys-llm-report-rankings-in-china/</a></span>
11. <span id="cite_note-ViLG-11">↑ <sup>[11.0](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-ViLG_11-0)</sup> <sup>[11.1](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-ViLG_11-1)</sup> Feng, Z. et al. (2022). *ERNIE-ViLG 2.0: Improving Text-to-Image Diffusion Model with Knowledge-Enhanced Mixture-of-Denoising-Experts*. arXiv:2210.15257. <a href="https://arxiv.org/abs/2210.15257" class="external free" rel="nofollow">https://arxiv.org/abs/2210.15257</a></span>
12. <span id="cite_note-Turbo-12">[↑](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-Turbo_12-0) AI Base News (2025). *Free Limited-Time Trial! Baidu's ERNIE 4.0 Turbo Launches on the ERNIE Bot Official Website*. <a href="https://news.aibase.com/news/9958" class="external free" rel="nofollow">https://news.aibase.com/news/9958</a></span>
13. <span id="cite_note-LMArena-13">↑ <sup>[13.0](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-LMArena_13-0)</sup> <sup>[13.1](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-LMArena_13-1)</sup> ERNIE Official Blog (2025). *ERNIE-5.0-Preview-1022 now ranks \#2 globally on the LMArena Text leaderboard*. <a href="https://ernie.baidu.com/blog/posts/ernie-5.0-preview-1022-release-on-lmarena/" class="external free" rel="nofollow">https://ernie.baidu.com/blog/posts/ernie-5.0-preview-1022-release-on-lmarena/</a></span>
14. <span id="cite_note-AINmultimodal-14">[↑](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-AINmultimodal_14-0) ArtificialIntelligence‑News (2025). *Baidu ERNIE multimodal AI beats GPT and Gemini in benchmarks*. <a href="https://www.artificialintelligence-news.com/news/baidu-ernie-multimodal-ai-gpt-and-gemini-benchmarks/" class="external free" rel="nofollow">https://www.artificialintelligence-news.com/news/baidu-ernie-multimodal-ai-gpt-and-gemini-benchmarks/</a></span>
15. <span id="cite_note-GH-15">[↑](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-GH_15-0) PaddlePaddle/ERNIE. GitHub. <a href="https://github.com/PaddlePaddle/ERNIE" class="external free" rel="nofollow">https://github.com/PaddlePaddle/ERNIE</a></span>
16. <span id="cite_note-Bias1-16">↑ <sup>[16.0](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-Bias1_16-0)</sup> <sup>[16.1](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-Bias1_16-1)</sup> <sup>[16.2](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-Bias1_16-2)</sup> Song, X. et al. (2024). *Comparing diversity, negativity, and stereotypes in Chinese-language AI technologies: a case study on Baidu, Ernie and Qwen*. arXiv:2408.15696. <a href="https://arxiv.org/abs/2408.15696" class="external free" rel="nofollow">https://arxiv.org/abs/2408.15696</a></span>
17. <span id="cite_note-Bias2-17">↑ <sup>[17.0](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-Bias2_17-0)</sup> <sup>[17.1](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-Bias2_17-1)</sup> <sup>[17.2](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-Bias2_17-2)</sup> <sup>[17.3](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-Bias2_17-3)</sup> PeerJ Computer Science (2025). *Comparing diversity, negativity, and stereotypes in Chinese-language AI technologies*. <a href="https://peerj.com/articles/cs-2694/" class="external free" rel="nofollow">https://peerj.com/articles/cs-2694/</a></span>
18. <span id="cite_note-Medical-18">[↑](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-Medical_18-0) Si, Y. et al. (2025). *Quality safety and disparity of an AI chatbot*. PMC.</span>
19. <span id="cite_note-Censorship-19">[↑](https://systems-analysis.info/int/ERNIE_(Baidu)_(TH)#cite_ref-Censorship_19-0) Pan et al. (2026). *Political censorship in large language models originating from China*. PNAS Nexus.</span>
