---
title: "DeepSeek (HE)"
source: "https://systems-analysis.info/int/DeepSeek_(HE)"
wiki: "systems-analysis.info/int"
article: "DeepSeek_(HE)"
language: "he"
categories:
  - "Category:Hebrew"
  - "Category:Large language models"
  - "Category:LLM families"
  - "Category:Machine learning"
revision_id: 1602
wiki_created_at: 2026-09-06T22:51:11Z
wiki_modified_at: 2026-09-06T22:51:11Z
downloaded_at: 2026-09-07T22:46:47Z
---

# DeepSeek (HE)

**DeepSeek** — חברת מחקר סינית בתחום הבינה המלאכותית, המפתחת מודלי שפה גדולים (LLM) ומערכות מולטי-מודליות. החברה זכתה לתהודה רחבה בזכות הפצת המשקולות של מודליה בקוד פתוח וביעילות הכלכלית הגבוהה שלהן, מה שגרר תיקון מחירים בשוק הבינה המלאכותית בסוף 2024 ותחילת 2025.<sup>[\[1\]](https://systems-analysis.info/int/DeepSeek_(HE)#cite_note-1)</sup>

## היסטוריה

מייסד DeepSeek הוא היזם ושותף-המייסד של קרן הגידור *High‑Flyer*, ליאנג וונפנג. באביב 2023 הפרידה High‑Flyer את חטיבת המחקר בתחום הבינה המלאכותית, שבמאי אותה שנה הפכה לחברת *DeepSeek AI*. עד 2025 גדל הצוות לכ-160 עובדים.<sup>[\[2\]](https://systems-analysis.info/int/DeepSeek_(HE)#cite_note-2)</sup> מראשית דרכה הצהירה החברה על מדיניות של פתיחות — פרסום משקולות («open‑weight») תחת רישיונות מתירניים ומיקוד במחקר בסיסי בתחום AGI.

בשונה מרוב הסטארט-אפים, DeepSeek ממומנת מתקציב המחקר והפיתוח של High‑Flyer, מה שמאפשר, לדברי המייסד, להתרכז במטרות ארוכות טווח במקום במונטיזציה מיידית.<sup>[\[3\]](https://systems-analysis.info/int/DeepSeek_(HE)#cite_note-3)</sup>

החברה עוררה תהודה משמעותית בקהילה הטכנולוגית והפיננסית בינואר 2025 לאחר שחרור המודל **DeepSeek-R1**. ההצהרה כי אימון מודל ברמה הדומה ל-GPT-4 עלה פחות מ-6 מיליון דולר (לעומת הערכות של 100+ מיליון דולר עבור GPT-4) גרמה לצניחת מניות ענקיות הטכנולוגיה ואילצה את התעשייה לשקול מחדש את הפרדיגמה של «יותר חישוב = מודל טוב יותר».<sup>[\[4\]](https://systems-analysis.info/int/DeepSeek_(HE)#cite_note-4)</sup>

## מאפיינים ארכיטקטוניים

Mixture‑of‑Experts (DeepSeekMoE)  
רוב מודלי הדגל של DeepSeek משתמשים בארכיטקטורת תערובת מומחים (MoE). בשונה ממודלים «צפופים», שבהם כל הפרמטרים מופעלים בעיבוד שאילתה, במודלי MoE רק חלק קטן מתת-הרשתות המתמחות («מומחים») מופעל עבור כל token. DeepSeek פיתחה מימוש MoE משלה עם מומחים «משותפים», פילוח גרעיני עדין ואיזון עומסים ללא הפסדי עזר, מה שמאפשר להפעיל רק חלק ממאות המיליארדים של פרמטרים ומפחית דרסטית את עלויות החישוב.<sup>[\[5\]](https://systems-analysis.info/int/DeepSeek_(HE)#cite_note-5)</sup>

Multi‑Head Latent Attention (MLA)  
שיטה לדחיסת KV‑cache לוקטור סמוי, החוסכת עד 93% מהזיכרון ומאפשרת שימוש בחלונות הקשר בגודל של עד 128,000 tokens. טכנולוגיה זו היא המפתח לעבודה יעילה עם טקסטים ארוכים.<sup>[\[6\]](https://systems-analysis.info/int/DeepSeek_(HE)#cite_note-6)</sup>

FP8 training ו-Multi‑Token Prediction  
במודלי משפחת V3 נעשה שימוש בדיוק מעורב FP8 (מספרי נקודה צפה של 8 ביטים) ובחיזוי מקביל של מספר tokens, מה שמאיץ את תהליכי האימון וה-inference.<sup>[\[7\]](https://systems-analysis.info/int/DeepSeek_(HE)#cite_note-7)</sup>

## משפחת המודלים

- **DeepSeek LLM** — מודלים בסיסיים של 7 ו-67 מיליארד פרמטרים (2023), הגרסה הדו-לשונית הראשונה (EN/ZH) שעלתה על *LLaMA‑2 70B* במספר משימות.<sup>[\[8\]](https://systems-analysis.info/int/DeepSeek_(HE)#cite_note-8)</sup>
- **DeepSeek‑Coder** (2023) — סדרת מודלים לתכנות (1.3 – 33 מיליארד) ופיתוחה *Coder‑V2* (16 מיליארד / 236 מיליארד MoE, הקשר 128K, 338 שפות תכנות).<sup>[\[9\]](https://systems-analysis.info/int/DeepSeek_(HE)#cite_note-9)</sup>
- **DeepSeek‑V2** (מאי 2024) — 236 מיליארד (21 מיליארד פעילים) MoE‑LLM עם MLA; אומן על 8.1 טריליון tokens.<sup>[\[10\]](https://systems-analysis.info/int/DeepSeek_(HE)#cite_note-10)</sup>
- **DeepSeek‑V3** (דצמבר 2024) — 671 מיליארד (37 מיליארד פעילים); אימון ≈2.8 מיליון שעות-GPU על Nvidia H800 בעלות ≈5.5 מיליון דולר.<sup>[\[11\]](https://systems-analysis.info/int/DeepSeek_(HE)#cite_note-11)</sup>
- **DeepSeek‑R1** (ינואר 2025) — סדרת מודלי reasoning; גרסת R1‑0528 התקרבה ל-*OpenAI o3* על AIME 2025 ו-LiveCodeBench.<sup>[\[12\]](https://systems-analysis.info/int/DeepSeek_(HE)#cite_note-12)</sup>
- **DeepSeek‑VL / VL2** — מודלי VL מולטי-מודליים (עד 4.5 מיליארד פעילים) עם עיבוד תמונות פסיפס דינמי בגודל 1024×1024.<sup>[\[13\]](https://systems-analysis.info/int/DeepSeek_(HE)#cite_note-13)</sup>
- **DeepSeek‑Math** 7B — מודל מתמחה, 51.7% דיוק על ה-benchmark MATH; קרוב ל-GPT‑4.<sup>[\[14\]](https://systems-analysis.info/int/DeepSeek_(HE)#cite_note-14)</sup>
- **DeepSeek‑Prover‑V2** — 671 מיליארד MoE להוכחת משפטים ב-Lean 4; 63.5% על miniF2F.
- **מודלי R1 מזוקקים** — גרסאות פתוחות מ-1.5 עד 70 מיליארד פרמטרים המבוססות על Llama ו-Qwen.<sup>[\[15\]](https://systems-analysis.info/int/DeepSeek_(HE)#cite_note-15)</sup>

## ציר זמן של גרסאות מפתח

| תאריך       | גרסה ומאפיינים מרכזיים                                            |
|-------------|-------------------------------------------------------------------|
| 2 נוב 2023  | **DeepSeek‑Coder v1:** מודלי open‑weight ראשונים לקוד.            |
| 29 נוב 2023 | **DeepSeek LLM 7B/67B:** מודל דו-לשוני, אומן על 2 טריליון tokens. |
| 11 ינו 2024 | **DeepSeek‑MoE 16B:** הופעת הבכורה של ארכיטקטורת MoE.             |
| 6 פבר 2024  | **DeepSeek‑Math 7B:** מודל מתמחה למתמטיקה (51.7% על MATH).        |
| 6 מאי 2024  | **DeepSeek‑V2 236B:** הטמעת ארכיטקטורות MLA ו-MoE.                |
| 17 יונ 2024 | **DeepSeek‑Coder‑V2:** הקשר 128K, תמיכה ב-338 שפות תכנות.         |
| 13 דצמ 2024 | **DeepSeek‑VL2:** מודל מולטי-מודלי מבוסס MoE.                     |
| 27 דצמ 2024 | **DeepSeek‑V3 671B:** מודל דגל, אומן בפחות מ-6 מיליון דולר.       |
| 20 ינו 2025 | **DeepSeek‑R1 / R1‑Zero:** מודלי reasoning, אומנו באמצעות RL.     |
| 27 ינו 2025 | **Janus‑Pro:** מודל ליצירת תמונות, עולה על DALL‑E 3.              |

## ביצועים ו-benchmarks

- *DeepSeek‑V3* עלתה על *Llama 3.1* ו-*Qwen 2.5* והתקרבה לרמת GPT‑4 על MMLU ו-GPQA‑Diamond.<sup>[\[16\]](https://systems-analysis.info/int/DeepSeek_(HE)#cite_note-16)</sup>
- *DeepSeek‑Coder‑V2* קיבלה 72.9% על Arena‑Hard — שוויון עם GPT‑4o ומעל כל המודלים הפתוחים מלבד Claude‑3.5‑Sonnet.<sup>[\[17\]](https://systems-analysis.info/int/DeepSeek_(HE)#cite_note-17)</sup>
- *DeepSeek‑Math 7B* — 51.7% על MATH, קרוב ל-Gemini‑Ultra עם גודל קטן פי 10.<sup>[\[18\]](https://systems-analysis.info/int/DeepSeek_(HE)#cite_note-18)</sup>
- *R1‑Zero* העלתה את תוצאת AIME 2024 pass@1 מ-15.6% ל-71% אך ורק באמצעות אימון RL.<sup>[\[19\]](https://systems-analysis.info/int/DeepSeek_(HE)#cite_note-19)</sup>

## רישוי ו-open‑source

רוב המודלים מופצים תחת רישיון MIT או Apache 2.0, המתיר שימוש מסחרי. החברה מפרסמת משקולות ב-Hugging Face ו-GitHub, אך שומרת סגורים את ה-dataset המלאים ואת pipelines האימון («open weight, but not full open source»).

## השפעה על התעשייה

- השקת R1 גרמה לירידה חד-יומית בשערי המניות של NVIDIA, Microsoft וחברות אחרות על רקע הידיעות על «מודל ברמת GPT‑4 ב-6 מיליון דולר».<sup>[\[20\]](https://systems-analysis.info/int/DeepSeek_(HE)#cite_note-20)</sup>
- הדגמת הצלחת האימון על שבבי Nvidia H800 תחת מגבלות יצוא עוררה דיון על יעילות הסנקציות האמריקאיות וזרזה את פיתוח מאיצי הבינה המלאכותית הסינים (למשל, Huawei Ascend 910B).

## ביקורת ומגבלות

- אבטחה: במבחן HarmBench עבר המודל R1 100% מהבקשות הבלתי רצויות («jailbreak»).
- צנזורה פוליטית: גרסאות הצ'אט מסננות נושאים «רגישים» עבור הממשלה הסינית (אירועי כיכר טיאנאנמן ב-1989, מעמד טייוואן וכדומה).
- אחסון נתונים: אחסון נתוני משתמשים על שרתים בסין מגביל את השימוש ב-API על ידי תאגידים מערביים הכפופים ל-GDPR ולמשטרי משפט דומים.<sup>[\[21\]](https://systems-analysis.info/int/DeepSeek_(HE)#cite_note-21)</sup>

## ספרות

- Dai, D. et al. (2024). *DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture‑of‑Experts Language Models*. arXiv:2401.06066.
- Ding, Y. et al. (2024). *LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens*. arXiv:2402.13753.
- Fedus, W.; Zoph, B.; Shazeer, N. (2021). *Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity*. arXiv:2101.03961.
- He, L. et al. (2025). *Scaling Instruction‑Tuned LLMs to Million‑Token Contexts via Hierarchical Synthetic Data Generation*. arXiv:2504.12637.
- Jegham, N. et al. (2025). *Visual Reasoning Evaluation of Grok, Deepseek Janus, Gemini, Qwen, Mistral, and ChatGPT*. arXiv:2502.16428.
- Lepikhin, D. et al. (2020). *GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding*. arXiv:2006.16668.
- Peng, B. et al. (2023). *YaRN: Efficient Context Window Extension of Large Language Models*. arXiv:2309.00071.
- Shen, Y. et al. (2025). *Long‑VITA: Scaling Large Multi‑modal Models to 1 Million Tokens with Leading Short‑Context Accuracy*. arXiv:2502.05177.
- Su, J. et al. (2021). *RoFormer: Enhanced Transformer with Rotary Position Embedding*. arXiv:2104.09864.
- Zhong, M. et al. (2024). *Understanding the RoPE Extensions of Long‑Context LLMs: An Attention Perspective*. arXiv:2406.13282.

## הערות

1.  <span id="cite_note-1">[↑](https://systems-analysis.info/int/DeepSeek_(HE)#cite_ref-1) DeepSeek's low-cost AI spotlights billions spent by US tech // Reuters. 2025-01-27.</span>
2.  <span id="cite_note-2">[↑](https://systems-analysis.info/int/DeepSeek_(HE)#cite_ref-2) Who is Liang Wenfeng, the founder of DeepSeek? // Reuters. 2025-01-28.</span>
3.  <span id="cite_note-3">[↑](https://systems-analysis.info/int/DeepSeek_(HE)#cite_ref-3) Who is Liang Wenfeng, the founder of DeepSeek? // Reuters. 2025-01-28.</span>
4.  <span id="cite_note-4">[↑](https://systems-analysis.info/int/DeepSeek_(HE)#cite_ref-4) DeepSeek's low-cost AI spotlights billions spent by US tech // Reuters. 2025-01-27.</span>
5.  <span id="cite_note-5">[↑](https://systems-analysis.info/int/DeepSeek_(HE)#cite_ref-5) DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model // Hugging Face. 2024.</span>
6.  <span id="cite_note-6">[↑](https://systems-analysis.info/int/DeepSeek_(HE)#cite_ref-6) DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model // Hugging Face. 2024.</span>
7.  <span id="cite_note-7">[↑](https://systems-analysis.info/int/DeepSeek_(HE)#cite_ref-7) DeepSeek-V3: A Parameter-Efficient MoE Large Language Model with Better Performance // arXiv. 2024.</span>
8.  <span id="cite_note-8">[↑](https://systems-analysis.info/int/DeepSeek_(HE)#cite_ref-8) DeepSeek LLM: Scaling Open-Source Language Models with Longtermism // arXiv. 2024.</span>
9.  <span id="cite_note-9">[↑](https://systems-analysis.info/int/DeepSeek_(HE)#cite_ref-9) DeepSeek-Coder-V2: A More Powerful and Economical Coder // arXiv. 2024.</span>
10. <span id="cite_note-10">[↑](https://systems-analysis.info/int/DeepSeek_(HE)#cite_ref-10) DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model // Hugging Face. 2024.</span>
11. <span id="cite_note-11">[↑](https://systems-analysis.info/int/DeepSeek_(HE)#cite_ref-11) DeepSeek-V3: A Parameter-Efficient MoE Large Language Model with Better Performance // arXiv. 2024.</span>
12. <span id="cite_note-12">[↑](https://systems-analysis.info/int/DeepSeek_(HE)#cite_ref-12) DeepSeek-R1: A 671B Parameter MoE LLM with Unprecedented Reasoning Capabilities // arXiv. 2025.</span>
13. <span id="cite_note-13">[↑](https://systems-analysis.info/int/DeepSeek_(HE)#cite_ref-13) GitHub - deepseek-ai/DeepSeek-VL: Towards Real-World Vision-Language Understanding // GitHub.</span>
14. <span id="cite_note-14">[↑](https://systems-analysis.info/int/DeepSeek_(HE)#cite_ref-14) DeepSeek-Math: Pushing the Limits of Mathematical Reasoning in Open-Source Models // arXiv. 2024.</span>
15. <span id="cite_note-15">[↑](https://systems-analysis.info/int/DeepSeek_(HE)#cite_ref-15) DeepSeek-R1: A 671B Parameter MoE LLM with Unprecedented Reasoning Capabilities // arXiv. 2025.</span>
16. <span id="cite_note-16">[↑](https://systems-analysis.info/int/DeepSeek_(HE)#cite_ref-16) DeepSeek-V3: A Parameter-Efficient MoE Large Language Model with Better Performance // arXiv. 2024.</span>
17. <span id="cite_note-17">[↑](https://systems-analysis.info/int/DeepSeek_(HE)#cite_ref-17) DeepSeek-Coder-V2: A More Powerful and Economical Coder // arXiv. 2024.</span>
18. <span id="cite_note-18">[↑](https://systems-analysis.info/int/DeepSeek_(HE)#cite_ref-18) DeepSeek-Math: Pushing the Limits of Mathematical Reasoning in Open-Source Models // arXiv. 2024.</span>
19. <span id="cite_note-19">[↑](https://systems-analysis.info/int/DeepSeek_(HE)#cite_ref-19) DeepSeek-R1: A 671B Parameter MoE LLM with Unprecedented Reasoning Capabilities // arXiv. 2025.</span>
20. <span id="cite_note-20">[↑](https://systems-analysis.info/int/DeepSeek_(HE)#cite_ref-20) DeepSeek's low-cost AI spotlights billions spent by US tech // Reuters. 2025-01-27.</span>
21. <span id="cite_note-21">[↑](https://systems-analysis.info/int/DeepSeek_(HE)#cite_ref-21) DeepSeek's low-cost AI spotlights billions spent by US tech // Reuters. 2025-01-27.</span>

## ראו גם

- מודלי שפה גדולים של OpenAI
- Mixture-of-Experts
