DeepSeek (TH)

From Systems analysis Wiki
Jump to navigation Jump to search

DeepSeek — บริษัทวิจัยด้านปัญญาประดิษฐ์สัญชาติจีน ที่พัฒนา large language model (LLM) และระบบ multimodal บริษัทได้รับการยอมรับอย่างกว้างขวางจากการเผยแพร่น้ำหนักโมเดลแบบเปิด และประสิทธิภาพเชิงต้นทุนที่สูง ซึ่งส่งผลให้เกิดการปรับราคาในตลาด AI ช่วงปลายปี 2024 ถึงต้นปี 2025[1]

ประวัติ

ผู้ก่อตั้ง DeepSeek คือนักธุรกิจและผู้ร่วมก่อตั้ง hedge fund High‑Flyer นาม เหลียง เวินเฟิง ในฤดูใบไม้ผลิปี 2023 High‑Flyer ได้แยกหน่วยงานวิจัย AI ออกมา และในเดือนพฤษภาคมของปีเดียวกันได้กลายเป็นบริษัท DeepSeek AI จนถึงปี 2025 จำนวนพนักงานเติบโตขึ้นเป็นประมาณ 160 คน[2] ตั้งแต่วันแรก บริษัทประกาศนโยบายความเปิดกว้าง ได้แก่ การเผยแพร่น้ำหนัก («open‑weight») ภายใต้ใบอนุญาตแบบอนุญาต และมุ่งเน้นการวิจัยพื้นฐานด้าน AGI

แตกต่างจากสตาร์ทอัพส่วนใหญ่ DeepSeek ได้รับเงินสนับสนุนจากงบประมาณ R&D ของ High‑Flyer ซึ่งตามคำกล่าวของผู้ก่อตั้ง ทำให้สามารถมุ่งเน้นเป้าหมายระยะยาวแทนการสร้างรายได้ในทันที[3]

บริษัทสร้างกระแสอย่างมากในวงการเทคโนโลยีและการเงินในเดือนมกราคม 2025 หลังจากเปิดตัวโมเดล DeepSeek-R1 การประกาศว่าการฝึกโมเดลที่เทียบได้กับ GPT-4 มีค่าใช้จ่ายน้อยกว่า 6 ล้านดอลลาร์ (เทียบกับการประมาณการ 100+ ล้านดอลลาร์สำหรับ GPT-4) ทำให้หุ้นของบริษัทเทคโนโลยีรายใหญ่ร่วงลง และบังคับให้อุตสาหกรรมทบทวนกรอบคิดที่ว่า «คำนวณมากขึ้น = โมเดลดีขึ้น»[4]

ลักษณะทางสถาปัตยกรรม

Mixture‑of‑Experts (DeepSeekMoE)
โมเดลเรือธงส่วนใหญ่ของ DeepSeek ใช้สถาปัตยกรรม Mixture of Experts (MoE) แตกต่างจากโมเดลแบบ «หนาแน่น» ที่เปิดใช้งานพารามิเตอร์ทั้งหมดเมื่อประมวลผลคำขอ โมเดล MoE จะเปิดใช้งานเพียงส่วนเล็กน้อยของซับเน็ตเฉพาะทาง («ผู้เชี่ยวชาญ») สำหรับแต่ละ token DeepSeek พัฒนาการนำ MoE ไปใช้งานของตนเองด้วยผู้เชี่ยวชาญ «ร่วม» การแบ่งส่วนแบบละเอียด และการปรับสมดุลโหลดโดยไม่ต้องใช้ค่าสูญเสียเสริม ซึ่งทำให้เปิดใช้งานเพียงบางส่วนจากพารามิเตอร์หลายร้อยพันล้านและลดต้นทุนการคำนวณได้อย่างมาก[5]
Multi‑Head Latent Attention (MLA)
วิธีการบีบอัด KV‑cache ลงในเวกเตอร์แฝง ประหยัดหน่วยความจำได้ถึง 93% และรองรับ context window ขนาดสูงสุด 128,000 token เทคโนโลยีนี้เป็นกุญแจสำคัญสำหรับการทำงานอย่างมีประสิทธิภาพกับข้อความยาว[6]
FP8 training และ Multi‑Token Prediction
โมเดลในตระกูล V3 ใช้ความแม่นยำแบบผสม FP8 (ตัวเลขทศนิยม 8 บิต) และการทำนายหลาย token พร้อมกัน ซึ่งช่วยเร่งกระบวนการ fine-tuning และ inference[7]

ตระกูลโมเดล

  • DeepSeek LLM — โมเดลพื้นฐาน 7 และ 67 พันล้านพารามิเตอร์ (2023) เผยแพร่แบบ bilingual (EN/ZH) ครั้งแรก ซึ่งเหนือกว่า LLaMA‑2 70B ในหลายงาน[8]
  • DeepSeek‑Coder (2023) — ชุดโมเดลสำหรับการเขียนโปรแกรม (1.3 – 33 พันล้าน) และพัฒนาต่อเป็น Coder‑V2 (16 พันล้าน / 236 พันล้าน MoE, context 128K, รองรับ 338 ภาษาการเขียนโปรแกรม)[9]
  • DeepSeek‑V2 (พฤษภาคม 2024) — MoE‑LLM 236 พันล้าน (21 พันล้านที่ใช้งาน) พร้อม MLA ฝึกด้วย 8.1 ล้านล้าน token[10]
  • DeepSeek‑V3 (ธันวาคม 2024) — 671 พันล้าน (37 พันล้านที่ใช้งาน) การฝึก ≈2.8 ล้าน GPU‑ชั่วโมงบน Nvidia H800 ค่าใช้จ่าย ≈5.5 ล้านดอลลาร์[11]
  • DeepSeek‑R1 (มกราคม 2025) — ชุดโมเดลสำหรับการอนุมานเชิงตรรกะ (reasoning) โดยเวอร์ชัน R1‑0528 เข้าใกล้ระดับ OpenAI o3 บน AIME 2025 และ LiveCodeBench[12]
  • DeepSeek‑VL / VL2 — โมเดล multimodal VL (สูงสุด 4.5 พันล้านที่ใช้งาน) พร้อมการประมวลผลภาพแบบโมเสกไดนามิก 1024×1024[13]
  • DeepSeek‑Math 7B — โมเดลเฉพาะทาง ความแม่นยำ 51.7% บน benchmark MATH ใกล้เคียงกับ GPT‑4[14]
  • DeepSeek‑Prover‑V2 — MoE 671 พันล้านสำหรับการพิสูจน์ทฤษฎีบทใน Lean 4 ได้ 63.5% บน miniF2F
  • โมเดล R1 แบบ distilled — เวอร์ชันเปิดตั้งแต่ 1.5 ถึง 70 พันล้านพารามิเตอร์บนฐาน Llama และ Qwen[15]

ไทม์ไลน์การเผยแพร่สำคัญ

วันที่ การเผยแพร่และลักษณะสำคัญ
2 พ.ย. 2023 DeepSeek‑Coder v1: โมเดล open‑weight ชุดแรกสำหรับโค้ด
29 พ.ย. 2023 DeepSeek LLM 7B/67B: โมเดล bilingual ฝึกด้วย 2 ล้านล้าน token
11 ม.ค. 2024 DeepSeek‑MoE 16B: การเปิดตัวสถาปัตยกรรม MoE ครั้งแรก
6 ก.พ. 2024 DeepSeek‑Math 7B: โมเดลเฉพาะทางด้านคณิตศาสตร์ (51.7% บน MATH)
6 พ.ค. 2024 DeepSeek‑V2 236B: การนำสถาปัตยกรรม MLA และ MoE มาใช้
17 มิ.ย. 2024 DeepSeek‑Coder‑V2: context 128K รองรับ 338 ภาษาการเขียนโปรแกรม
13 ธ.ค. 2024 DeepSeek‑VL2: โมเดล multimodal บนพื้นฐาน MoE
27 ธ.ค. 2024 DeepSeek‑V3 671B: โมเดลเรือธง ฝึกด้วยค่าใช้จ่ายน้อยกว่า 6 ล้านดอลลาร์
20 ม.ค. 2025 DeepSeek‑R1 / R1‑Zero: โมเดลสำหรับการอนุมาน ฝึกด้วย Reinforcement Learning
27 ม.ค. 2025 Janus‑Pro: โมเดลสำหรับการสร้างภาพ เหนือกว่า DALL‑E 3

ประสิทธิภาพและ benchmark

  • DeepSeek‑V3 เหนือกว่า Llama 3.1 และ Qwen 2.5 และเข้าใกล้ระดับ GPT‑4 ใน MMLU และ GPQA‑Diamond[16]
  • DeepSeek‑Coder‑V2 ได้ 72.9% บน Arena‑Hard — เทียบเท่า GPT‑4o และสูงกว่าโมเดลเปิดทั้งหมด ยกเว้น Claude‑3.5‑Sonnet[17]
  • DeepSeek‑Math 7B — 51.7% บน MATH ใกล้เคียงกับ Gemini‑Ultra แต่มีขนาดเล็กกว่า 10 เท่า[18]
  • R1‑Zero เพิ่มผลลัพธ์ AIME 2024 pass@1 จาก 15.6% เป็น 71% โดยใช้เพียงการฝึกด้วย Reinforcement Learning[19]

การอนุญาตสิทธิ์และ open‑source

โมเดลส่วนใหญ่เผยแพร่ภายใต้ใบอนุญาต MIT หรือ Apache 2.0 ซึ่งอนุญาตให้ใช้งานเชิงพาณิชย์ได้ บริษัทเผยแพร่น้ำหนักบน Hugging Face และ GitHub แต่ยังคงปิดเป็นความลับสำหรับ dataset ฉบับสมบูรณ์และ pipeline การฝึก («open weight แต่ไม่ใช่ full open source»)

ผลกระทบต่ออุตสาหกรรม

  • การเปิดตัว R1 ทำให้ราคาหุ้นของ NVIDIA, Microsoft และบริษัทอื่น ๆ ลดลงภายในวันเดียว จากข่าว «โมเดลระดับ GPT‑4 ในราคา 6 ล้านดอลลาร์»[20]
  • การสาธิตความสำเร็จในการฝึกบนชิป Nvidia H800 ภายใต้ข้อจำกัดการส่งออก กระตุ้นการอภิปรายเรื่องประสิทธิผลของมาตรการคว่ำบาตรของสหรัฐฯ และเร่งการพัฒนาตัวเร่งความเร็ว AI ของจีน (เช่น Huawei Ascend 910B)

การวิจารณ์และข้อจำกัด

  • ความปลอดภัย: ในการทดสอบ HarmBench โมเดล R1 ยอมตอบคำขอที่ไม่พึงประสงค์ 100% («jailbreak»)
  • การเซนเซอร์ทางการเมือง: เวอร์ชัน chat กรองหัวข้อ «ละเอียดอ่อน» สำหรับรัฐบาลจีน (เหตุการณ์จัตุรัสเทียนอันเหมินปี 1989 สถานะของไต้หวัน เป็นต้น)
  • การจัดเก็บข้อมูล: การจัดเก็บข้อมูลผู้ใช้บนเซิร์ฟเวอร์ในประเทศจีนจำกัดการใช้งาน API โดยบริษัทตะวันตกที่ต้องปฏิบัติตาม GDPR และระบอบกฎหมายที่คล้ายกัน[21]

เอกสารอ้างอิง

  • Dai, D. et al. (2024). DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture‑of‑Experts Language Models. arXiv:2401.06066.
  • Ding, Y. et al. (2024). LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens. arXiv:2402.13753.
  • Fedus, W.; Zoph, B.; Shazeer, N. (2021). Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity. arXiv:2101.03961.
  • He, L. et al. (2025). Scaling Instruction‑Tuned LLMs to Million‑Token Contexts via Hierarchical Synthetic Data Generation. arXiv:2504.12637.
  • Jegham, N. et al. (2025). Visual Reasoning Evaluation of Grok, Deepseek Janus, Gemini, Qwen, Mistral, and ChatGPT. arXiv:2502.16428.
  • Lepikhin, D. et al. (2020). GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding. arXiv:2006.16668.
  • Peng, B. et al. (2023). YaRN: Efficient Context Window Extension of Large Language Models. arXiv:2309.00071.
  • Shen, Y. et al. (2025). Long‑VITA: Scaling Large Multi‑modal Models to 1 Million Tokens with Leading Short‑Context Accuracy. arXiv:2502.05177.
  • Su, J. et al. (2021). RoFormer: Enhanced Transformer with Rotary Position Embedding. arXiv:2104.09864.
  • Zhong, M. et al. (2024). Understanding the RoPE Extensions of Long‑Context LLMs: An Attention Perspective. arXiv:2406.13282.

หมายเหตุ

  1. DeepSeek's low-cost AI spotlights billions spent by US tech // Reuters. 2025-01-27.
  2. Who is Liang Wenfeng, the founder of DeepSeek? // Reuters. 2025-01-28.
  3. Who is Liang Wenfeng, the founder of DeepSeek? // Reuters. 2025-01-28.
  4. DeepSeek's low-cost AI spotlights billions spent by US tech // Reuters. 2025-01-27.
  5. DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model // Hugging Face. 2024.
  6. DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model // Hugging Face. 2024.
  7. DeepSeek-V3: A Parameter-Efficient MoE Large Language Model with Better Performance // arXiv. 2024.
  8. DeepSeek LLM: Scaling Open-Source Language Models with Longtermism // arXiv. 2024.
  9. DeepSeek-Coder-V2: A More Powerful and Economical Coder // arXiv. 2024.
  10. DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model // Hugging Face. 2024.
  11. DeepSeek-V3: A Parameter-Efficient MoE Large Language Model with Better Performance // arXiv. 2024.
  12. DeepSeek-R1: A 671B Parameter MoE LLM with Unprecedented Reasoning Capabilities // arXiv. 2025.
  13. GitHub - deepseek-ai/DeepSeek-VL: Towards Real-World Vision-Language Understanding // GitHub.
  14. DeepSeek-Math: Pushing the Limits of Mathematical Reasoning in Open-Source Models // arXiv. 2024.
  15. DeepSeek-R1: A 671B Parameter MoE LLM with Unprecedented Reasoning Capabilities // arXiv. 2025.
  16. DeepSeek-V3: A Parameter-Efficient MoE Large Language Model with Better Performance // arXiv. 2024.
  17. DeepSeek-Coder-V2: A More Powerful and Economical Coder // arXiv. 2024.
  18. DeepSeek-Math: Pushing the Limits of Mathematical Reasoning in Open-Source Models // arXiv. 2024.
  19. DeepSeek-R1: A 671B Parameter MoE LLM with Unprecedented Reasoning Capabilities // arXiv. 2025.
  20. DeepSeek's low-cost AI spotlights billions spent by US tech // Reuters. 2025-01-27.
  21. DeepSeek's low-cost AI spotlights billions spent by US tech // Reuters. 2025-01-27.

ดูเพิ่มเติม

  • Large language model ของ OpenAI
  • Mixture-of-Experts