---
title: "Nemotron (NVIDIA) (TH)"
source: "https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)"
wiki: "systems-analysis.info/int"
article: "Nemotron_(NVIDIA)_(TH)"
language: "th"
categories:
  - "Category:Large language models"
  - "Category:Machine learning"
  - "Category:Thai"
revision_id: 4857
wiki_created_at: 2026-09-06T23:40:59Z
wiki_modified_at: 2026-09-06T23:40:59Z
downloaded_at: 2026-09-07T23:05:15Z
---

# Nemotron (NVIDIA) (TH)

**Nemotron** — ตระกูลโมเดลภาษาขนาดใหญ่ (Large Language Model, LLM) แบบเปิดน้ำหนัก พัฒนาโดยแผนก Applied Deep Learning Research (ADLR) ของบริษัท NVIDIA โมเดลเหล่านี้อยู่ในขอบเขตของ Machine Learning และการประมวลผลภาษาธรรมชาติ (Natural Language Processing, NLP) และออกแบบมาสำหรับงานหลากหลาย ได้แก่ การสร้างข้อมูลสังเคราะห์ การอนุมานเชิงตรรกะ (reasoning) แอปพลิเคชันเอเจนต์ (agentic AI) และการนำไปใช้งานบนฮาร์ดแวร์ระดับอุตสาหกรรมของ NVIDIA ตระกูลนี้รวมโมเดลที่มีสถาปัตยกรรมและขนาดต่างกัน ตั้งแต่ 4 พันล้านถึงประมาณ 5 แสนล้านพารามิเตอร์ ครอบคลุม dense decoder-only Transformer ซีรีส์ Nemotron-4 (2024) โมเดลไฮบริด Mamba-Transformer (Nemotron-H, 2025) และโมเดลไฮบริด Mamba-Transformer พร้อม Mixture-of-Experts (MoE) ในซีรีส์ Nemotron 3 (2025) ณ เดือนมีนาคม พ.ศ. 2569 ตระกูลนี้มีโมเดลมากกว่า 20 รุ่น ครอบคลุมงานสร้างข้อความ การให้เหตุผล การทำความเข้าใจเอกสารด้วยภาพ การดึงข้อมูล และการประเมินคุณภาพคำตอบ โมเดลส่วนใหญ่เผยแพร่แบบเปิดน้ำหนักภายใต้สิทธิ์การใช้งาน NVIDIA Open Model License และ Llama Community License<sup>[\[1\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-N3_whitepaper-1)[\[2\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-N3_nano_report-2)[\[3\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-N4_340B-3)</sup>

## ประวัติและบริบท

การพัฒนาตระกูล Nemotron ดำเนินไปในบริบทของกลยุทธ์ NVIDIA ในการสร้าง stack เครื่องมือที่ครบวงจรสำหรับ AI เชิงสร้างสรรค์ ตั้งแต่โครงสร้างพื้นฐานการฝึก (DGX, ซูเปอร์คอมพิวเตอร์บน GPU H100/B200) ไปจนถึงแพลตฟอร์มการฝึก (NeMo Framework) การปรับแนว (NeMo-Aligner) การนำไปใช้งาน (NIM — NVIDIA Inference Microservices) และความปลอดภัย (NeMo Guardrails)<sup>[\[4\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-NV_developer-4)[\[5\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-NeMo_docs-5)</sup>

### Nemotron-3 8B (2023)

โมเดลแรกที่เปิดให้สาธารณะใช้งานในซีรีส์นี้ — decoder Transformer ที่มีพารามิเตอร์ 8 พันล้านตัวและ context ขนาด 4096 token ฝึกบน **3.8 ล้านล้าน token** ใน 53 ภาษาธรรมชาติและ 37 ภาษาโปรแกรม การฝึกดำเนินการบน GPU A100 จำนวน 1,024 ตัวเป็นเวลา 19 วัน ไม่มีการเผยแพร่รายงานทางเทคนิคอย่างเป็นทางการบน arXiv โมเดลเผยแพร่ผ่าน NeMo Framework และ Hugging Face มีตัวแปร Chat (SFT และ SteerLM) ใช้สิทธิ์ NVIDIA AI Foundation Models Community License<sup>[\[6\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-N3_8B_card-6)[\[7\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-MS_nemotron3-7)</sup>

*หมายเหตุเกี่ยวกับชื่อ*: «Nemotron-3» ปี 2023 และ «Nemotron 3» เดือนธันวาคม 2025 เป็นโมเดลคนละรุ่น รุ่นแรกเป็นต้นแบบในช่วงต้น รุ่นหลังคือรุ่นปัจจุบันที่ใช้สถาปัตยกรรม MoE

### Nemotron-4 15B (กุมภาพันธ์ 2024)

การเผยแพร่ครั้งแรกที่มีเอกสารอย่างเป็นทางการของซีรีส์ Nemotron-4 — โมเดลที่มีพารามิเตอร์ 15 พันล้านตัว ฝึกบน 8 ล้านล้าน token เป็นเวลาประมาณ 13 วันตามปฏิทินบนโหนด DGX H100 จำนวน 384 โหนด (สูงสุด 3,072 GPU) ใช้ tensor parallelism 8 เท่าและ data parallelism ตั้งแต่ 96 ถึง 288 ค่า MFU (Model FLOPs Utilization) สูงสุดอยู่ที่ 34.3% ที่ batch size 384 สถาปัตยกรรมเป็น decoder-only Transformer พร้อม Grouped-Query Attention (GQA) และ Rotary Position Embeddings (RoPE) จากผลการทดสอบ benchmark ในขณะที่เผยแพร่ โมเดลนี้เหนือกว่าโมเดลโอเพนซอร์สทุกรุ่นที่มีขนาดใกล้เคียงในงานหลายภาษา รวมถึงโมเดลบางรุ่นที่ใหญ่กว่าถึงสี่เท่า<sup>[\[8\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-N4_15B-8)</sup>

### Nemotron-4 340B (มิถุนายน 2024)

ในเดือนมิถุนายน 2024 มีการเผยแพร่โมเดลที่ใหญ่ที่สุดของซีรีส์ Nemotron-4 — 340 พันล้านพารามิเตอร์ (331.6 พันล้านพารามิเตอร์ที่ไม่ใช่ embedding) ผ่านการ pre-train บน 9 ล้านล้าน token (8 ล้านล้าน pre-training + 1 ล้านล้าน continued training พร้อมการถ่วงน้ำหนักแหล่งข้อมูลคุณภาพสูง) การฝึกดำเนินการบนโหนด DGX H100 จำนวน 768 โหนด (สูงสุด 6,144 GPU) ด้วย MFU สูงสุด 42.4% ตั้งแต่เดือนธันวาคม 2023 ถึงพฤษภาคม 2024 ตระกูลนี้มีสามรูปแบบ ได้แก่ Base, Instruct และ Reward ขนาดโมเดลถูกเลือกให้สามารถใส่ในโหนด DGX H100 เดียว (GPU 8 ตัว) เมื่อนำไปใช้ในรูปแบบความแม่นยำ FP8 มากกว่า 98% ของข้อมูลที่ใช้ในการปรับแนว (alignment) ถูกสร้างขึ้นแบบสังเคราะห์โดยโมเดลในซีรีส์เองในกระบวนการวนซ้ำ และ pipeline การสร้างข้อมูลสังเคราะห์ได้รับการเปิดเผยเพื่อให้ทำซ้ำได้<sup>[\[3\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-N4_340B-3)</sup>

### Llama-3.1-Nemotron-70B (ตุลาคม 2024)

ในเดือนตุลาคม 2024 มีการเผยแพร่โมเดลที่ fine-tune จาก Meta Llama-3.1-70B-Instruct ได้แก่ Llama-3.1-Nemotron-70B-Instruct และ Llama-3.1-Nemotron-70B-Reward ณ วันที่ 1 ตุลาคม 2024 โมเดล Instruct ติดอันดับที่ 1 พร้อมกันในสาม benchmark อัตโนมัติ ได้แก่ Arena Hard — 85.0, AlpacaEval 2 LC — 57.6% และ MT-Bench (GPT-4-Turbo) — 8.98 โมเดล Reward บรรลุ 94.1% บน RewardBench<sup>[\[9\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-LN70B_Instruct-9)[\[10\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-LN70B_Reward-10)</sup>

### การเปลี่ยนผ่านสู่สถาปัตยกรรมไฮบริด (2025)

ในปี 2025 ซีรีส์ได้ขยายด้วยสถาปัตยกรรมไฮบริดที่ผสมเลเยอร์ Mamba-2 (State Space Model, SSM — โมเดลปริภูมิสถานะ) และ Transformer ในช่วงมีนาคม–พฤษภาคม 2025 มีการเผยแพร่ตระกูลโมเดลสำหรับการให้เหตุผล **Llama-Nemotron** (Nano 8B, Super 49B, Ultra 253B) พร้อมโหมดการให้เหตุผลแบบสลับได้<sup>[\[11\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-Llama_Nemotron-11)</sup> ในเดือนเมษายน 2025 มีการเผยแพร่ **Nemotron-H** (arXiv:2504.03624) — ตระกูลโมเดลไฮบริด Mamba-Transformer ขนาด 8B, 56B และ 47B พารามิเตอร์<sup>[\[12\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-Nemotron_H-12)</sup> ในเดือนสิงหาคม 2025 มีการนำเสนอ **Nemotron Nano 2** (9B-v2, arXiv:2508.14444)<sup>[\[13\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-Nano2-13)</sup>

### Nemotron 3 (ธันวาคม 2025)

เมื่อวันที่ 15 ธันวาคม 2025 มีการประกาศรุ่นที่สาม — ตระกูล Nemotron 3 ที่มีโมเดล Nano, Super และ Ultra รวมสถาปัตยกรรม Mamba-2, Transformer และ MoE พร้อม context window สูงสุด 1 ล้าน token Nano ได้รับการเผยแพร่พร้อมรายงานทางเทคนิค ส่วน Super และ Ultra มีกำหนดในช่วงครึ่งแรกของปี 2026<sup>[\[1\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-N3_whitepaper-1)[\[2\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-N3_nano_report-2)[\[14\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-N3_blog-14)[\[15\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-NVIDIA_news-15)</sup>

## พื้นฐานทางทฤษฎี

### สถาปัตยกรรม Transformer ใน Nemotron-4

โมเดลซีรีส์ Nemotron-4 สร้างบนสถาปัตยกรรม decoder Transformer มาตรฐานพร้อม causal attention ความน่าจะเป็นของลำดับ token $x_{1},x_{2},\ldots,x_{T}$ ถูก factorize ดังนี้:

$$
p(x_{1},\ldots,x_{T}) = \prod\limits_{t = 1}^{T}p(x_{t} \mid x_{1},\ldots,x_{t - 1};\theta),
$$

โดยที่ $\theta$ คือพารามิเตอร์ของโมเดล<sup>[\[8\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-N4_15B-8)</sup>

กลไก attention ถูกใช้งานในรูปแบบ Grouped-Query Attention (GQA)<sup>[\[16\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-GQA-16)</sup>:

$$
\operatorname{Attention}(Q,K,V) = \operatorname{softmax}\!\left( \frac{QK^{\top}}{\sqrt{d_{k}}} \right)V,
$$

โดยที่ $Q,K,V$ คือเมทริกซ์ query, key และ value; $d_{k}$ คือมิติของ attention head

สำหรับการ encode ตำแหน่ง ใช้ Rotary Position Embeddings (RoPE)<sup>[\[17\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-RoPE-17)</sup>:

$$
\operatorname{RoPE}(x_{m},m) = \begin{pmatrix}
{x_{m}^{(1)}\cos m\theta_{1} - x_{m}^{(2)}\sin m\theta_{1}} \\
{x_{m}^{(1)}\sin m\theta_{1} + x_{m}^{(2)}\cos m\theta_{1}} \\
 \vdots 
\end{pmatrix},
$$

โดยที่ $x_{m}$ คือเวกเตอร์ที่ตำแหน่ง $m$, $\theta_{i} = 10000^{- 2i/d}$, $d$ คือมิติของเวกเตอร์

ในเลเยอร์ MLP ใช้ activation แบบ Squared ReLU: $f(x) = (\max(0,x))^{2}$ ซึ่งให้ความเบาบางของ activation โมเดลไม่มี bias ไม่ใช้ dropout และเมทริกซ์ embedding ของ input และ output ไม่ถูกผูกร่วมกัน (untied embeddings) tokenizer เป็น SentencePiece BPE พร้อมการรักษาช่องว่าง การแบ่งตัวเลขทีละตัวอักษร และ byte-level fallback ขนาด vocabulary คือ 256,000<sup>[\[8\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-N4_15B-8)[\[3\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-N4_340B-3)</sup>

##### พารามิเตอร์สถาปัตยกรรมของ Nemotron-4

| พารามิเตอร์       | Nemotron-4 15B              | Nemotron-4 340B              |
|-----------------|-----------------------------|------------------------------|
| จำนวนพารามิเตอร์  | 15 พันล้าน (3.2 พันล้าน embed.) | 340 พันล้าน (9.4 พันล้าน embed.) |
| จำนวนเลเยอร์     | 32                          | 96                           |
| มิติซ่อน           | 6144                        | 18,432                       |
| Attention head  | 48                          | 96                           |
| KV head (GQA)   | 8                           | 8                            |
| ความยาว context | 4096                        | 4096                         |
| ขนาด vocabulary | 256,000                     | 256,000                      |

แหล่งที่มา: arXiv:2402.16819, arXiv:2406.11704<sup>[\[8\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-N4_15B-8)[\[3\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-N4_340B-3)</sup>

### สถาปัตยกรรมไฮบริด Mamba-Transformer (Nemotron-H)

ในโมเดล Nemotron-H บล็อก self-attention มาตรฐานบางส่วนถูกแทนที่ด้วยบล็อก Mamba-2 — State Space Models (SSM) ในการสร้างแบบ autoregressive แต่ละเลเยอร์ self-attention ต้องการ $O(n)$ การดำเนินการและหน่วยความจำต่อขั้นตอน (เนื่องจาก KV cache) ในขณะที่เลเยอร์ Mamba มีต้นทุนหน่วยความจำและการคำนวณคงที่สำหรับแต่ละ token ที่สร้าง<sup>[\[12\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-Nemotron_H-12)</sup>

Mamba-2 อธิบายด้วยความสัมพันธ์เชิงเรียกซ้ำ<sup>[\[18\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-Mamba2-18)</sup>:

$$
h_{t} = Ah_{t - 1} + Bx_{t},\quad y_{t} = Ch_{t},
$$

โดยที่ $h_{t} \in {\mathbb{R}}^{N}$ คือ hidden state, $A \in {\mathbb{R}}^{N \times N}$ คือเมทริกซ์ทแยงมุมของการเปลี่ยนสถานะ, $B \in {\mathbb{R}}^{N \times 1}$ และ $C \in {\mathbb{R}}^{1 \times N}$ คือเมทริกซ์ projection, $x_{t}$ คือ input token, $y_{t}$ คือสัญญาณ output ใน Mamba-2 เมทริกซ์ $A$, $B$, $C$ ขึ้นอยู่กับ input (input-dependent) ซึ่งแตกต่างจาก linear SSM แบบคลาสสิก

ใน Nemotron-H-56B ใช้ **54 เลเยอร์ Mamba-2 + 54 เลเยอร์ MLP + 10 เลเยอร์ self-attention** โดยเลเยอร์ attention ถูกจัดวางอย่างสม่ำเสมอท่ามกลางเลเยอร์ Mamba โมเดลนี้ฝึกบน 20 ล้านล้าน token ในรูปแบบ FP8 (per-tensor scaling) บน GPU H100 จำนวน 6,144 ตัว รุ่น Nemotron-H-8B มี 24 เลเยอร์ Mamba-2, 24 เลเยอร์ MLP และ 4 เลเยอร์ self-attention ฝึกบน 15 ล้านล้าน token<sup>[\[12\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-Nemotron_H-12)</sup>

### สถาปัตยกรรม MoE ของ Nemotron 3

โมเดล Nemotron 3 (ธันวาคม 2025) รวมองค์ประกอบสถาปัตยกรรมสามส่วน ได้แก่ Mamba-2, Transformer และ Mixture-of-Experts<sup>[\[1\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-N3_whitepaper-1)[\[2\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-N3_nano_report-2)</sup> Nemotron 3 Nano (30B-A3B) มี 52 เลเยอร์ ได้แก่ **23 เลเยอร์ Mamba-2 + MoE, 23 เลเยอร์ MoE และ 6 เลเยอร์ GQA attention** แต่ละเลเยอร์ MoE ประกอบด้วย expert แบบ routed จำนวน 128 ตัว + 1 expert แบบ shared โดยมี 6 expert ที่ถูกเปิดใช้งานสำหรับแต่ละ token การกำหนดเส้นทางทำโดย MLP router ที่เรียนรู้ได้พร้อม sigmoid gating การปรับสมดุลโหลดทำโดยไม่ใช้ auxiliary loss function (aux-loss-free) จำนวนพารามิเตอร์ทั้งหมดคือ 31.6 พันล้าน แต่มีเพียง 3.2 พันล้านที่ active สำหรับแต่ละ token<sup>[\[2\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-N3_nano_report-2)</sup>

การ normalize ใช้ RMSNorm<sup>[\[19\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-RMSNorm-19)</sup> ใน Mamba ไม่มี positional embedding (ข้อมูลตำแหน่งถูกแสดงโดยนัยใน SSM state) ใช้ untied embedding ไม่มี dropout และไม่มี bias ในเลเยอร์ linear<sup>[\[2\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-N3_nano_report-2)</sup>

### ส่วนขยายใน Super/Ultra: LatentMoE และ Multi-Token Prediction

โมเดล Super และ Ultra จากตระกูล Nemotron 3 ใช้นวัตกรรมสถาปัตยกรรมเพิ่มเติมสองรายการ:

- **Latent MoE (LatentMoE)** — การฉาย input representation ลงในพื้นที่ latent ที่มีมิติน้อยกว่าก่อนการกำหนดเส้นทาง ซึ่งช่วยให้เพิ่มจำนวน expert ได้โดยมีต้นทุนการคำนวณใกล้เคียงเดิม และเพิ่มความแม่นยำโดยไม่ลด throughput<sup>[\[1\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-N3_whitepaper-1)</sup>
- **Multi-Token Prediction (MTP)** — การทำนายหลาย token ถัดไปพร้อมกัน ใช้เพื่อปรับปรุงคุณภาพข้อความยาวและการ speculative decoding ในขณะ inference<sup>[\[1\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-N3_whitepaper-1)</sup>

การฝึก Super และ Ultra ดำเนินการในรูปแบบ NVFP4 — รูปแบบตัวเลข 4 บิตของ NVIDIA บนสถาปัตยกรรม Blackwell<sup>[\[1\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-N3_whitepaper-1)</sup>

### วิธีการบีบอัดโมเดล: Minitron, Puzzle, MiniPuzzle

สำหรับการสร้างโมเดลขนาดกะทัดรัด NVIDIA ใช้วิธีการหลายแบบ:

- **Minitron** — วิธีการ pruning แบบมีโครงสร้าง (structured pruning) และ knowledge distillation มีการศึกษา pruning สองประเภท ได้แก่ แบบความลึก (การลบเลเยอร์) และแบบความกว้าง (การลดมิติของ hidden layer, attention head และ MLP projection ร่วมกัน) การใช้ Minitron กับ Nemotron-4 15B ช่วยให้ได้โมเดล 8B และ 4B โดยลดต้นทุนการฝึกได้ถึง 40 เท่าเมื่อเทียบกับการฝึกตั้งแต่เริ่มต้น<sup>[\[20\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-Minitron-20)</sup>
- **Puzzle** — อัลกอริทึม Neural Architecture Search (NAS) ที่เลือกการกำหนดค่าที่เหมาะสมสำหรับแต่ละบล็อก (จำนวน KV head, ขนาด FFN, การมี/ไม่มี attention) พร้อม block-wise distillation นี่คือวิธีที่ Llama 3.1 405B ถูกบีบอัดเป็น 253B (Llama-Nemotron Ultra) และ Llama 3.3 70B เป็น 49B (Llama-Nemotron Super)<sup>[\[21\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-Puzzle-21)</sup>
- **MiniPuzzle** — ส่วนขยายของ Puzzle สำหรับสถาปัตยกรรมไฮบริด ที่บีบอัด Nemotron-H-56B เป็น 47B โดยใช้เพียง 63 พันล้าน token ของ distillation เมื่อมีความแม่นยำใกล้เคียงกัน โมเดลที่บีบอัดแล้วทำงานเร็วขึ้น 20%<sup>[\[12\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-Nemotron_H-12)</sup>

## วิธีการปรับแนว

### HelpSteer และ HelpSteer2

**HelpSteer** — ชุดข้อมูลจำนวน 37,120 ตัวอย่าง (ใช้สิทธิ์ CC-BY-4.0) สร้างร่วมกับ Scale AI โดยแต่ละคำตอบมีคำอธิบายตาม 5 คุณลักษณะ ได้แก่ ความเป็นประโยชน์ (helpfulness), ความถูกต้อง (correctness), ความสอดคล้อง (coherence), ความซับซ้อน (complexity) และความยืดยาว (verbosity) ตามมาตรา Likert 0–4<sup>[\[22\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-HelpSteer-22)</sup>

**HelpSteer2** — ชุดข้อมูลที่ปรับปรุงแล้วจำนวน 21,362 ตัวอย่าง (10,681 prompt × 2 คำตอบ) โดยมากกว่า 95% ของ prompt มาจาก ShareGPT (คำขอผู้ใช้จริง) ประมาณ 50% ของคำอธิบายถูกกรองออกว่ามีคุณภาพไม่เพียงพอ ส่วนขยาย **HelpSteer2-Preference** เพิ่มการเปรียบเทียบความชอบของผู้อธิบายเป็นคู่ ทำให้สามารถฝึกทั้งโมเดล reward แบบ regression และแบบ pairwise บนข้อมูลชุดเดียวกัน<sup>[\[23\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-HelpSteer2-23)[\[24\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-HelpSteer2_Pref-24)</sup>

### SteerLM

**SteerLM** — วิธีการ SFT (Supervised Fine-Tuning — การ fine-tune แบบมีผู้สอน) พร้อมการปรับตามคุณลักษณะ (helpfulness, correctness, coherence, complexity, verbosity) ซึ่งช่วยให้ผู้ใช้ควบคุมสไตล์ของคำตอบในขณะ inference กระบวนการนี้ประกอบด้วย: (1) การฝึกโมเดลทำนายคุณลักษณะ (APM) บน HelpSteer, (2) การระบุข้อมูลด้วย APM, (3) SFT พร้อมการปรับตามค่าคุณลักษณะเป้าหมาย, (4) bootstrapping<sup>[\[25\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-SteerLM-25)</sup>

### DPO และ RPO

**DPO (Direct Preference Optimization)** — วิธีการ optimize preference โดยตรงโดยไม่มีโมเดล reward ที่ชัดเจน ใช้ใน pipeline ของ Nemotron-4 340B Instruct และ Nemotron Nano 2<sup>[\[26\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-DPO-26)</sup>

**RPO (Reward-aware Preference Optimization)** — framework คณิตศาสตร์แบบรวมของ NVIDIA ที่สรุป DPO, IPO, SimPO, REINFORCE และ SteerLM 2.0 เป็นกรณีพิเศษ RPO ลดระยะห่างระหว่างการทำนายของโมเดล reward โดยนัยและโมเดล reward ที่ชัดเจนเป้าหมาย<sup>[\[27\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-RPO-27)</sup>:

$$
\mathcal{L}_{\text{RPO}}(\theta) = {\mathbb{E}}_{(x,y_{w},y_{l}) \sim \mathcal{D}}\left\lbrack d\!\left( r_{\phi}(x,y_{w}) - r_{\phi}(x,y_{l}),\;\beta\log\frac{\pi_{\theta}(y_{w}|x)}{\pi_{\text{ref}}(y_{w}|x)} - \beta\log\frac{\pi_{\theta}(y_{l}|x)}{\pi_{\text{ref}}(y_{l}|x)} \right) \right\rbrack,
$$

โดยที่ $r_{\phi}$ คือโมเดล reward ที่ชัดเจน, $\pi_{\theta}$ คือ policy ที่กำลังฝึก, $\pi_{\text{ref}}$ คือ reference policy, $d( \cdot , \cdot )$ คือฟังก์ชันระยะห่าง, $y_{w}$/$y_{l}$ คือคำตอบที่ถูกชอบ/ถูกปฏิเสธ RPO ใช้ในการฝึก Nemotron-4 340B Instruct และโมเดล Llama-Nemotron<sup>[\[27\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-RPO-27)[\[3\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-N4_340B-3)[\[11\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-Llama_Nemotron-11)</sup>

### Reinforcement Learning แบบหลายสภาพแวดล้อม (RLVR)

ใน Nemotron 3 ใช้ Multi-environment RLVR (Reinforcement Learning from Verifiable Rewards) — การฝึกพร้อมกันในหลายสภาพแวดล้อมภายใต้ NeMo Gym ได้แก่ คณิตศาสตร์แข่งขัน การเขียนโปรแกรม QA, tool-use, instruction-following และ long-context อัลกอริทึมที่ใช้คือ GRPO (Group Relative Policy Optimization) พร้อม masked importance sampling<sup>[\[2\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-N3_nano_report-2)[\[14\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-N3_blog-14)</sup>

Nemotron 3 รองรับ **การควบคุมงบประมาณการให้เหตุผล** (reasoning budget control) แบบละเอียด: ผู้ใช้สามารถกำหนดขีดจำกัด token สำหรับ thinking trace ผ่าน flag ใน chat template (สลับ «detailed thinking on/off» หรือจำกัดความยาว)<sup>[\[2\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-N3_nano_report-2)</sup>

## รายการโมเดล

### ตารางสรุป

| โมเดล                         | ปี         | พารามิเตอร์ทั้งหมด / ที่ active | Context   | สถาปัตยกรรม                     | คุณสมบัติหลัก                                                    |
|-------------------------------|-----------|---------------------------|-----------|--------------------------------|--------------------------------------------------------------|
| **Nemotron-3 8B**             | 2023      | 8B / 8B                   | 4K        | Dense Transformer              | 3.8T token; GPU A100 จำนวน 1,024 ตัว; 19 วัน                   |
| **Nemotron-4 15B**            | 2024      | 15B / 15B                 | 4K        | Dense Transformer              | 8T token; MFU 34.3%                                          |
| **Nemotron-4 340B**           | 2024      | 340B / 340B               | 4K        | Dense Transformer              | 9T token; Base/Instruct/Reward; ข้อมูลสังเคราะห์ alignment \>98% |
| **Llama-3.1-Nemotron-70B**    | 2024      | 70B / 70B                 | 128K      | Dense Transformer (Llama 3.1)  | RLHF (REINFORCE); RewardBench 94.1%                          |
| **Llama-Nemotron Nano 8B**    | 2025      | 8B / 8B                   | 128K      | Dense Transformer (Llama 3.1)  | Reasoning toggle                                             |
| **Llama-Nemotron Nano 4B**    | 2025      | 4B / 4B                   | 128K      | Dense Transformer (Minitron)   | NAS compression จาก Llama 3.1 8B                             |
| **Llama-Nemotron Super 49B**  | 2025      | 49B / 49B                 | 128K      | Dense Transformer (Puzzle NAS) | จาก Llama 3.3 70B                                            |
| **Llama-Nemotron Ultra 253B** | 2025      | 253B / 253B               | 128K      | Dense Transformer (Puzzle NAS) | จาก Llama 3.1 405B; GPQA 76.0%                               |
| **Nemotron-H 8B**             | 2025      | 8B / 8B                   | —         | ไฮบริด Mamba-Transformer        | 15T token; 24 Mamba-2 + 24 MLP + 4 Attn                      |
| **Nemotron-H 56B**            | 2025      | 56B / 56B                 | —         | ไฮบริด Mamba-Transformer        | 20T token FP8; 54 Mamba-2 + 54 MLP + 10 Attn                 |
| **Nemotron-H 47B**            | 2025      | 47B / 47B                 | ~1M (FP4) | ไฮบริด Mamba-Transformer        | MiniPuzzle จาก 56B; เร็วขึ้น 20%                                |
| **Nemotron Nano 2 (9B)**      | 2025      | 12B → 9B / 9B             | 128K      | ไฮบริด Mamba-Transformer        | 20T token; Minitron distillation                             |
| **Nemotron 3 Nano**           | 2025      | 31.6B / 3.2B              | 1M        | ไฮบริด Mamba-Transformer MoE    | 25T token; 128 expert, 6 active                              |
| **Nemotron 3 Super**          | 2025–2026 | ~100B / ~10B              | 1M        | ไฮบริด LatentMoE                | MTP; NVFP4                                                   |
| **Nemotron 3 Ultra**          | 2025–2026 | ~500B / ~50B              | 1M        | ไฮบริด LatentMoE                | MTP; NVFP4                                                   |

### Nemotron-4 15B

สถาปัตยกรรม: 32 เลเยอร์, hidden dimension 6144, attention head 48 ตัว, KV head 8 ตัว (GQA) จำนวนพารามิเตอร์ทั้งหมดคือ 15 พันล้าน (embedding 3.2 พันล้าน + non-embedding 12.5 พันล้าน) ผลลัพธ์: MMLU — 64.2% (5-shot), BBH — 58.7% (3-shot), GSM8K — 46.0% (8-shot, maj@1), HumanEval — 31.6% (0-shot, pass@1) บน benchmark หลายภาษา (XCOPA, TyDiQA-GoldP, MGSM) โมเดลนี้เหนือกว่าโมเดลที่ใหญ่กว่าถึง 4 เท่า<sup>[\[8\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-N4_15B-8)</sup>

### Nemotron-4 340B

สถาปัตยกรรม: 96 เลเยอร์, hidden dimension 18,432, attention head 96 ตัว, KV head 8 ตัว

**Nemotron-4 340B Base** แสดงผล: MMLU — 81.1% (5-shot), BBH — 85.4% (3-shot), HumanEval — 57.3% (0-shot), ARC-Challenge — 94.3% (25-shot)<sup>[\[3\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-N4_340B-3)</sup>

**Nemotron-4 340B Instruct** ผ่านการปรับแนวหลายขั้นตอน: Code SFT → General SFT → DPO → RPO (3 รอบ) ผลลัพธ์: MT-Bench — 8.22 (GPT-4-Turbo judge), MMLU — 78.7% (0-shot), GSM8K — 92.3% (0-shot), HumanEval — 73.2% (0-shot), Arena Hard — 54.2, AlpacaEval 2.0 LC — 41.5%<sup>[\[3\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-N4_340B-3)</sup>

**Nemotron-4 340B Reward** แทนที่เลเยอร์ softmax สุดท้ายด้วยการฉายเชิงเส้นของ hidden state ของเลเยอร์สุดท้ายไปยังเวกเตอร์ 5 คุณลักษณะของ HelpSteer2 รางวัลสุดท้ายคือผลรวมถ่วงน้ำหนักของคุณลักษณะทั้งห้า การฝึกดำเนินการ 2 epoch บนข้อมูล HelpSteer2 (batch size 128) บน RewardBench โมเดลนี้ได้ 92.0% เหนือกว่า GPT-4o (84.7%), Gemini 1.5 Pro (88.1%) และ Claude-3-Opus (80.7%) ณ เวลาที่เผยแพร่<sup>[\[3\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-N4_340B-3)</sup>

| โมเดล                    | Arena Hard | AlpacaEval 2.0 LC | MT-Bench |
|--------------------------|------------|-------------------|----------|
| Nemotron-4-340B-Instruct | 54.2       | 41.5%             | 8.22     |
| Llama-3-70B-Instruct     | 41.1       | 34.4%             | 8.16     |
| Mixtral-8x22B-Instruct   | 36.4       | 30.9%             | 7.63     |
| Qwen-2-72B-Instruct      | 48.1       | 38.8%             | 8.26     |

แหล่งที่มา: arXiv:2406.11704, ตาราง 5<sup>[\[3\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-N4_340B-3)</sup>

### Llama-3.1-Nemotron-70B

**Llama-3.1-Nemotron-70B-Reward** ฝึกด้วยวิธีไฮบริดที่รวม Bradley-Terry (การเปรียบเทียบความชอบแบบคู่) และ SteerLM regression (การประเมินหลายคุณลักษณะตามมาตรา Likert) การฝึกดำเนินการบนข้อมูล HelpSteer2 เท่านั้น (CC-BY-4.0) บน RewardBench โมเดลนี้ได้ 94.1% ครองอันดับ 1 ณ เวลาที่เผยแพร่<sup>[\[10\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-LN70B_Reward-10)</sup>

**Llama-3.1-Nemotron-70B-Instruct** ปรับแนวด้วย RLHF โดยใช้อัลกอริทึม REINFORCE (ไม่ใช่ PPO) และโมเดล reward Nemotron-70B-Reward โดยใช้โครงสร้างพื้นฐาน NeMo-Aligner พร้อมการเร่งด้วย TRT-LLM<sup>[\[9\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-LN70B_Instruct-9)</sup>

### Llama-Nemotron: Nano, Super, Ultra (2025)

ตระกูลโมเดลสำหรับการให้เหตุผลที่ได้มาจาก Llama 3.x ด้วยวิธี NAS, distillation และ post-training ขั้นสูง<sup>[\[11\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-Llama_Nemotron-11)</sup>

| โมเดล                     | พารามิเตอร์ | โมเดลฐาน                       | Context |
|---------------------------|-----------|--------------------------------|---------|
| Llama-Nemotron-Nano 8B    | 8 พันล้าน   | Llama-3.1-8B-Instruct          | 128K    |
| Llama-Nemotron-Nano 4B    | 4 พันล้าน   | Minitron 4B (จาก Llama 3.1 8B) | 128K    |
| Llama-Nemotron-Super 49B  | 49 พันล้าน  | Llama-3.3-70B → NAS (Puzzle)   | 128K    |
| Llama-Nemotron-Ultra 253B | 253 พันล้าน | Llama-3.1-405B → NAS (Puzzle)  | 128K    |

กระบวนการฝึกห้าขั้นตอน: (1) NAS + FFN Fusion, (2) distillation + continued pre-training, (3) SFT (รวม reasoning trace จาก DeepSeek-R1), (4) RL ขนาดใหญ่ (GRPO สำหรับ Ultra, REINFORCE/RLOO + Online RPO สำหรับ Nano), (5) การปรับแนวสำหรับบทสนทนา<sup>[\[11\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-Llama_Nemotron-11)</sup>

โมเดลเหล่านี้เป็นกลุ่มแรกในบรรดา LLM โอเพนซอร์สที่รองรับ **โหมดการให้เหตุผลแบบสลับได้** (reasoning toggle): เมื่อเปิดใช้ («detailed thinking on» ใน system prompt) โมเดลจะสร้าง chain of thought ในแท็ก `<think>...</think>` ก่อนตอบ เมื่อปิด จะตอบโดยตรง<sup>[\[11\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-Llama_Nemotron-11)</sup>

ผลลัพธ์ของ Llama-Nemotron-Ultra 253B (reasoning ON): GPQA Diamond — 76.0%, AIME 2024 — 80.8%, MATH 500 — ~97% เพื่อเปรียบเทียบ: DeepSeek-R1 (671B total / 37B active) — GPQA Diamond 71.5%, AIME 2024 ~79.8% Ultra สามารถใช้งานบนโหนด 8×H100 เดียว<sup>[\[11\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-Llama_Nemotron-11)</sup>

### Nemotron-H (เมษายน 2025)

ตระกูลโมเดลไฮบริด dense ที่ฝึกตั้งแต่เริ่มต้นและออกแบบสำหรับงานที่มีการสร้างข้อความยาว Nemotron-H-56B ฝึกบน 20 ล้านล้าน token ในรูปแบบ FP8 — เป็นหนึ่งในการทดลอง FP8 pre-training สาธารณะที่ใหญ่ที่สุด Nemotron-H-47B ได้จากการบีบอัด 56B ด้วยวิธี MiniPuzzle (63 พันล้าน token ของ distillation) และสามารถใส่ในหน่วยความจำของ RTX 5090 เดียว (FP4) ที่ context ~1M token<sup>[\[12\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-Nemotron_H-12)</sup>

| Benchmark             | Nemotron-H-56B | Nemotron-H-47B | Qwen-2.5-72B | Llama-3.1-70B |
|-----------------------|----------------|----------------|--------------|---------------|
| MMLU-Pro (5-shot CoT) | 60.5           | 61.8           | 58.8         | 51.3          |
| MMLU (5-shot)         | 84.2           | 83.6           | 86.1         | 78.8          |
| GSM8K (8-shot CoT)    | 93.7           | 93.3           | 90.9         | 83.9          |
| HumanEval (0-shot)    | 60.4           | 61.0           | 56.7         | 57.3          |

แหล่งที่มา: arXiv:2504.03624<sup>[\[12\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-Nemotron_H-12)</sup>

เมื่อสร้างด้วย input 65,536 token และ output 1,024 token บน H100 โมเดล Nemotron-H-47B ทำงานเร็วกว่า Qwen-2.5-72B และ Llama-3.1-70B ถึง 2.9 เท่า<sup>[\[12\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-Nemotron_H-12)</sup>

### Nemotron Nano 2 (สิงหาคม 2025)

โมเดลไฮบริด Mamba-Transformer สำหรับการให้เหตุผล **Nemotron-Nano-12B-v2-Base** — โมเดลต้นกำเนิดที่มี 62 เลเยอร์ (ส่วนใหญ่เป็น Mamba-2 + MLP + 4 เลเยอร์ attention) ฝึกบน 20 ล้านล้าน token ในรูปแบบ FP8 ตั้งแต่เริ่มต้น จากโมเดลนี้ได้ **Nemotron-Nano-9B-v2** ด้วยวิธี Minitron แบบขยาย สามารถทำงานที่ context 128K token บน GPU NVIDIA A10G เดียว (22 GB, BF16) Post-training: SFT (80 พันล้าน token รวม trace จาก DeepSeek-R1-0528) → GRPO → DPO → RLHF บน benchmark การให้เหตุผล โมเดลนี้มีความแม่นยำใกล้เคียง Qwen3-8B แต่สร้างข้อความเร็วกว่า 3–6 เท่าสำหรับลำดับ output ที่ยาว<sup>[\[13\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-Nano2-13)</sup>

### Nemotron 3 Nano 30B-A3B

เผยแพร่ในเดือนธันวาคม 2025 พารามิเตอร์: 31.6 พันล้านทั้งหมด, 3.2 พันล้าน active สถาปัตยกรรม: 52 เลเยอร์, hidden dimension 2688, expert dimension 1856, Mamba state dimension 128 MoE: expert แบบ routed จำนวน 128 ตัว + 1 shared, 6 active ต่อ token Context สูงสุด 1,048,576 token ฝึกบน 25 ล้านล้าน token (WSD schedule, batch size 3072, ข้อมูล cutoff เดือนมิถุนายน 2025) สองเฟส: Phase 1 — 23.5 ล้านล้าน (ข้อมูลหลากหลาย), Phase 2 — 1.5 ล้านล้าน (ข้อมูลคุณภาพสูง) + 121 พันล้าน token long-context CPT<sup>[\[2\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-N3_nano_report-2)</sup>

| Benchmark             | Nemotron 3 Nano 30B-A3B | Qwen3-30B-A3B-Thinking | GPT-OSS-20B |
|-----------------------|-------------------------|------------------------|-------------|
| MMLU-Pro              | 78.30                   | 80.9                   | 75.0        |
| AIME25 (no tools)     | 89.06                   | 85.0                   | 91.7        |
| AIME25 (with tools)   | 99.17                   | —                      | 98.7        |
| GPQA (no tools)       | 73.04                   | 73.4                   | 71.5        |
| LiveCodeBench v6      | 68.25                   | 66.0                   | 61.0        |
| SWE-Bench (OpenHands) | 38.76                   | 22.0\*                 | 34.0        |
| TauBench V2 Avg       | 49.04                   | 47.7                   | 47.5        |
| Arena-Hard-V2 Avg     | 67.65                   | 57.8                   | 48.55       |
| RULER-100 @ 1M        | 86.34                   | 77.5                   | —           |

\* — ค่าจากแหล่งข้อมูลรอง เงื่อนไขการทำซ้ำ — NeMo Evaluator SDK แหล่งที่มา: arXiv:2512.20848<sup>[\[2\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-N3_nano_report-2)[\[28\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-NeMo_eval-28)</sup>

ปริมาณงาน (input 8K / output 16K, single H200): สูงกว่า Qwen3-30B-A3B-Thinking 3.3 เท่าและสูงกว่า GPT-OSS-20B 2.2 เท่า<sup>[\[2\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-N3_nano_report-2)</sup>

### โมเดลและทิศทางเพิ่มเติม

- **OpenReasoning-Nemotron** (กรกฎาคม 2025) — ซีรีส์โมเดล 1.5B–32B พารามิเตอร์ที่อิงจาก Qwen 2.5 และ fine-tune บนข้อมูลจาก DeepSeek-R1-0528 รุ่น 32B พร้อม GenSelect@64 เหนือกว่า o3 (high) บน benchmark คณิตศาสตร์บางรายการ<sup>[\[29\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-OpenReasoning-29)</sup>
- **Nemotron Elastic** (arXiv:2511.16664) — framework ของโมเดลย่อยแบบซ้อน (6B, 9B, 12B) ภายในโมเดลต้นกำเนิดเดียว ลดต้นทุนการฝึก 360 เท่าเมื่อเทียบกับการฝึกตั้งแต่เริ่มต้น<sup>[\[30\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-Elastic-30)</sup>
- **Nemotron-CrossThink** — framework การเรียนรู้ด้วยการเสริมแรงแบบหลายโดเมนนอกเหนือจากงานคณิตศาสตร์ ให้ผลดีขึ้น +30.1% บน MATH-500 และ +12.8% บน MMLU-PRO<sup>[\[31\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-CrossThink-31)</sup>
- **Nemotron-UltraLong** — ขยาย context เป็น 4 ล้าน token<sup>[\[32\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-UltraLong-32)</sup>
- **Jet-Nemotron** — NAS หลัง post-training สำหรับโมเดลไฮบริดขนาดกะทัดรัด 2B/4B<sup>[\[33\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-JetNemotron-33)</sup>

### โมเดลเฉพาะทาง

ในระบบนิเวศ Nemotron ยังมีโมเดลสำหรับงานเฉพาะทาง:

- **Nemotron Nano V2 VL** — โมเดล vision-language สำหรับ OCR, การประมวลผลตาราง และวิดีโอ<sup>[\[34\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-NanoV2_VL-34)</sup>
- **Nemotron Parse 1.1** — โมเดลสำหรับ OCR และการดึงโครงสร้างเอกสาร<sup>[\[35\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-Parse-35)</sup>
- **Nemotron ColEmbed V2** — โมเดล embedding สำหรับการค้นหาเอกสารด้วยภาพ (late interaction)<sup>[\[36\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-ColEmbed-36)</sup>
- **Nemotron Speech** — โมเดลสำหรับการรู้จำเสียงพูดอัตโนมัติ (ASR) การสังเคราะห์เสียงพูด (TTS) และการแปลภาษาด้วยเครื่อง (NMT)<sup>[\[1\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-N3_whitepaper-1)</sup>
- **Llama 3.1 Nemotron Safety Guard 8B V3** — โมเดลจำแนกเนื้อหาที่ไม่ปลอดภัยใน 23 หมวดหมู่ใน 9 ภาษา ด้วยความแม่นยำ 84.2%<sup>[\[37\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-Safety-37)</sup>

## ข้อมูลสำหรับ Pre-training

### องค์ประกอบข้อมูลของ Nemotron-4

คลังข้อมูล pre-training ของ Nemotron-4 15B และ 340B ประกอบด้วยสามหมวดหลัก ได้แก่ ข้อความภาษาอังกฤษ (70%) ข้อความหลายภาษาใน 53 ภาษา (15%) และซอร์สโค้ดใน 43 ภาษาโปรแกรม (15%) ใช้การลบข้อมูลซ้ำระดับเอกสาร (exact และ near-duplicate) รวมถึงการกรองด้วยโมเดลภาษาและ heuristic โมเดล 15B ฝึกบน 8 ล้านล้าน token โมเดล 340B บน 9 ล้านล้าน (8 ล้านล้าน pre-training + 1 ล้านล้าน continued training พร้อมการถ่วงน้ำหนักแหล่งคุณภาพสูง)<sup>[\[8\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-N4_15B-8)[\[3\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-N4_340B-3)</sup>

### Nemotron-CC

ชุดข้อมูล **Nemotron-CC** สร้างจาก Common Crawl โดยใช้ ensemble classifier คุณภาพและการถอดความข้อความสังเคราะห์ ปริมาณรวม — **6.3 ล้านล้าน token** (4.4 ล้านล้านที่ลบซ้ำแล้ว + 1.9 ล้านล้านการถอดความสังเคราะห์) Pipeline ถูกรวมไว้ในเครื่องมือ NeMo Curator งานวิจัยนี้ได้รับการยอมรับเป็น Long Paper ในการประชุม ACL 2025<sup>[\[38\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-NemotronCC-38)</sup>

### Nemotron-CC-Math

ชุดย่อยที่เน้นคณิตศาสตร์ขนาด **133 พันล้าน token** ดึงจาก Common Crawl ด้วย pipeline Lynx + LLM พร้อมการรักษาโครงสร้างสูตรคณิตศาสตร์และโค้ดใน LaTeX<sup>[\[39\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-NemotronCC_Math-39)</sup>

### ข้อมูล Nemotron 3 Nano

การ pre-train ของ Nemotron 3 Nano ดำเนินการบน 25 ล้านล้าน token ชุดข้อมูลประกอบด้วย: Nemotron-CC-v2.1, Nemotron-CC-Code-v1, Nemotron-CC-Math-v1 และ dataset STEM เฉพาะทาง สำหรับการฝึก long-context ใช้ CPT เพิ่มเติม 121 พันล้าน token<sup>[\[2\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-N3_nano_report-2)</sup>

### Nemotron Pretraining Dataset v1

ชุดข้อมูลที่สร้างสังเคราะห์ ครอบคลุม STEM ข้อความวิชาการ งานให้เหตุผล และโดเมนหลายภาษา มีโทเค็นภาษามากกว่า 10 ล้านล้านและตัวอย่าง SFT จำนวน 18 ล้านรายการ ชุดข้อมูลทั้งหมดเผยแพร่ภายใต้สิทธิ์การใช้งานแบบเปิดที่อนุญาต<sup>[\[40\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-NemotronNano2_page-40)</sup>

### Pipeline การสร้างข้อมูลสังเคราะห์ (SDG)

Pipeline ที่อธิบายในรายงานเกี่ยวกับ Nemotron-4 340B ประกอบด้วยขั้นตอนต่อไปนี้<sup>[\[3\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-N4_340B-3)</sup>:

- **การสร้าง prompt**: คำขอแบบ one-turn และ two-turn สังเคราะห์ ที่สร้างด้วย Mixtral-8x7B-Instruct-v0.1 (สิทธิ์ Apache 2.0) ตามแม่แบบ Open QA, Writing, Closed QA, Math & Coding
- **การสร้างคำตอบ**: คำตอบหลายรูปแบบจากโมเดลรุ่นกลางเพื่อให้มีความหลากหลาย
- **การประเมินคุณภาพ**: สามแนวทาง — Ground-Truth-as-a-Judge (สำหรับงานที่มีคำตอบตรวจสอบได้), LLM-as-Judge (การเปรียบเทียบคู่คำตอบด้วยโมเดลภาษา), Reward-Model-as-Judge (ใช้ Nemotron-4-340B-Reward)
- **การปรับแนวแบบวนซ้ำ** จากโมเดลอ่อนแอสู่โมเดลแข็งแกร่ง (weak-to-strong)

จากตัวอย่างที่มีคำอธิบายโดยมนุษย์ประมาณ 20,000 รายการ (10,000 สำหรับ SFT, 10,000 HelpSteer2 สำหรับโมเดล reward) ถูกสังเคราะห์เป็นข้อมูลการปรับแนวทั้งหมดที่เหลือ<sup>[\[3\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-N4_340B-3)</sup>

## การบีบอัดเชิงตัวเลขและการ optimize การ inference

โมเดล Nemotron 3 Nano รองรับ Post-Training Quantization (PTQ) เป็น FP8 โดยใช้ ModelOpt และ Megatron-LM ใช้ selective quantization — เลเยอร์ attention และ Mamba บางส่วนถูกเก็บใน BF16 เพื่อรักษาคุณภาพ ส่วนที่เหลือถูกบีบอัดเป็น FP8 ตามรายงานทางเทคนิค วิธีนี้ให้ความแม่นยำ ~99% retention<sup>[\[2\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-N3_nano_report-2)</sup> Nano ยังรองรับการ inference ในรูปแบบ NVFP4<sup>[\[1\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-N3_whitepaper-1)</sup>

## ตารางสรุป benchmark ตามรุ่น

| โมเดล                           | MMLU  | GSM8K | HumanEval | Arena Hard | MT-Bench | เงื่อนไข                  |
|---------------------------------|-------|-------|-----------|------------|----------|-------------------------|
| Nemotron-4 15B                  | 64.2% | 46.0% | 31.6%     | —          | —        | base; 5-/8-/0-shot      |
| Nemotron-4 340B Base            | 81.1% | —     | 57.3%     | —          | —        | 5-/—/0-shot             |
| Nemotron-4 340B Instruct        | 78.7% | 92.3% | 73.2%     | 54.2       | 8.22     | 0-shot                  |
| Llama-3.1-Nemotron-70B-Instruct | —     | —     | —         | 85.0       | 8.98     | ต.ค. 2024               |
| LN-Ultra 253B (reasoning ON)    | —     | —     | —         | —          | —        | GPQA 76.0%, AIME 80.8%  |
| Nemotron-H-47B                  | 83.6% | 93.3% | 61.0%     | —          | —        | 5-/8-CoT/0-shot         |
| Nemotron 3 Nano (30B-A3B)       | —     | —     | —         | 67.7 (v2)  | —        | AIME25 89.1%, LCB 68.3% |

ควรตีความการเปรียบเทียบด้วยความระมัดระวัง เนื่องจากเงื่อนไขการประเมิน รุ่นของ benchmark และวันที่ทดสอบแตกต่างกันระหว่างรุ่น ผลลัพธ์มาจากรายงานทางเทคนิคของ NVIDIA การประเมินอิสระอาจแตกต่างออกไป (ดูส่วน «ข้อจำกัด»)<sup>[\[8\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-N4_15B-8)[\[3\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-N4_340B-3)[\[9\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-LN70B_Instruct-9)[\[11\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-Llama_Nemotron-11)[\[12\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-Nemotron_H-12)[\[2\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-N3_nano_report-2)</sup>

## การนำไปใช้งาน

### การนำไปใช้ผ่าน NVIDIA NIM

NVIDIA NIM — ชุด microservice แบบ container ที่ optimize สำหรับการ inference พร้อม API ที่เข้ากันได้กับ OpenAI โมเดล Nemotron พร้อมใช้งานเป็น NIM microservice บนแพลตฟอร์ม build.nvidia.com NIM รองรับรูปแบบ FP8, BF16, NVFP4 และการนำไปใช้ผ่าน Helm chart บน Kubernetes โมเดลยังเข้ากันได้กับ framework อิสระ ได้แก่ vLLM, SGLang, Ollama, llama.cpp, TensorRT-LLM<sup>[\[4\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-NV_developer-4)[\[1\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-N3_whitepaper-1)</sup>

### การสร้างข้อมูลสังเคราะห์

Nemotron-4 340B ออกแบบมาเพื่อใช้เป็น generator ข้อมูลสังเคราะห์: โมเดล Instruct สร้างคำตอบ โมเดล Reward ประเมินและกรอง สิทธิ์ NVIDIA Open Model License อนุญาตอย่างชัดเจนให้ใช้ output ของโมเดลสำหรับการฝึกโมเดลอื่น ตามข้อมูลของ ServiceNow 15% ของข้อมูล pre-training ของโมเดล Apriel 1.6 มาจากชุดข้อมูล Nemotron<sup>[\[3\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-N4_340B-3)[\[15\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-NVIDIA_news-15)[\[41\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-NeMo_synth-41)</sup>

### ระบบเอเจนต์

โมเดลในตระกูล Nemotron 3 (Nano, Super, Ultra) ออกแบบสำหรับ workload แบบ multi-agent ได้แก่ การวางแผน, retrieval, การเรียกใช้เครื่องมือ (tool use) Nemotron 3 Nano แสดงผล 38.76% บน SWE-Bench (พร้อม OpenHands) และ 49.04% (เฉลี่ย) บน TauBench V2<sup>[\[2\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-N3_nano_report-2)</sup>

### การนำไปใช้ในองค์กร

พันธมิตรที่ประกาศไว้ ได้แก่ ServiceNow (โมเดลร่วม Apriel Nemotron 15B สำหรับระบบอัตโนมัติของ workflow), CrowdStrike (แพลตฟอร์ม Charlotte AI), Perplexity (การกำหนดเส้นทางคำขอ), Accenture, Cadence, Deloitte, Oracle, Palantir, Siemens, Synopsys, Zoom โมเดลพร้อมใช้งานบน AWS Amazon Bedrock และมีแผนรองรับ Google Cloud, CoreWeave, Nebius<sup>[\[15\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-NVIDIA_news-15)</sup>

## ข้อจำกัดและปัญหาที่ยังเปิดอยู่

### การประเมินอิสระและความแตกต่างจากข้อมูลของ NVIDIA

การประเมินอิสระพบความแตกต่างจากผลลัพธ์ที่ NVIDIA รายงาน ตามข้อมูลของ Artificial Analysis (กุมภาพันธ์ 2026) Llama-Nemotron-Ultra 253B (reasoning) ได้คะแนน Intelligence Index 15 ในขณะที่ค่ามัธยฐานอยู่ที่ 26 สำหรับโมเดลขนาดใกล้เคียงกัน บนแพลตฟอร์ม Chatbot Arena (LMArena) โมเดล Nemotron อยู่ในระดับกลางของการจัดอันดับ: Llama 3.3 Nemotron Super 49B — Elo 1327 ±12 (~อันดับที่ 147), Nemotron 3 Nano — 1317 ±6 (~อันดับที่ 164)<sup>[\[42\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-AA-42)[\[43\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-LMArena-43)</sup>

### ความยืดยาว

โมเดล Nemotron แสดงความยืดยาวสูง: ในการประเมินของ Artificial Analysis โมเดล Nemotron 3 Nano สร้าง 140 ล้าน token (ในขณะที่ค่ามัธยฐานของโมเดลอื่นอยู่ที่ 12 ล้าน) ซึ่งเพิ่มต้นทุน inference การวิเคราะห์ของ DEV Community พบว่า Llama-3.1-Nemotron-70B-Instruct แม้จะนำในเชิงทางการบน benchmark อัตโนมัติ แต่แสดงความแม่นยำไม่เพียงพอในงานเชิงปฏิบัติบางอย่าง และคะแนนสูงบน benchmark แบบ Arena อาจอธิบายได้จากการชอบคำตอบที่ยืดยาวและมั่นใจ แต่ไม่จำเป็นต้องแม่นยำ<sup>[\[42\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-AA-42)[\[44\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-Saplin-44)</sup>

### การสร้างภาพลวงตาและ bias

NVIDIA ระบุในการ์ดโมเดลว่าโมเดลอาจ «เสริมความลำเอียงและสร้างคำตอบที่เป็นพิษ โดยเฉพาะกับ prompt ที่เป็นพิษ» รวมถึง «สร้างข้อมูลที่ไม่ถูกต้อง ละเว้นรายละเอียดสำคัญ» ความเปิดกว้างของน้ำหนักและข้อมูลช่วยให้ตรวจสอบจากภายนอกได้<sup>[\[11\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-Llama_Nemotron-11)[\[13\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-Nano2-13)</sup>

### ความต้องการในการคำนวณ

โมเดล dense (Nemotron-4 340B) ต้องการคลัสเตอร์ 768 โหนด DGX H100 สำหรับการฝึกเต็มรูปแบบ ซึ่งจำกัดการทำซ้ำสำหรับกลุ่มวิจัยที่ไม่มีโครงสร้างพื้นฐานที่เทียบเท่า Context ยาว (1M) ในขณะ inference ต้องการ VRAM จำนวนมาก<sup>[\[3\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-N4_340B-3)[\[2\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-N3_nano_report-2)</sup>

### การเสื่อมลงเมื่อขยาย context

เมื่อขยาย context เป็นลำดับที่ยาวมาก พบการเสื่อมลงของผลลัพธ์บน benchmark ที่ใช้ context สั้น<sup>[\[32\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-UltraLong-32)</sup>

### การบีบอัดเชิงตัวเลขในสถาปัตยกรรมไฮบริด

การใช้ความละเอียดบิตต่ำมาก (FP8/NVFP4) ในสถาปัตยกรรม Mamba-Transformer นำไปสู่การลดลงเล็กน้อยของความแม่นยำ (~1% บน benchmark บางรายการ) ปัญหานี้แก้ไขด้วย selective quantization ซึ่งทำให้สถาปัตยกรรม compiler และ inference server ซับซ้อนขึ้น<sup>[\[2\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-N3_nano_report-2)[\[1\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-N3_whitepaper-1)</sup>

### ปัญหาที่ยังเปิดอยู่อื่นๆ

- การพึ่งพา benchmark ที่อาจมีการปนเปื้อนข้อมูลการฝึก
- การขาดโปรโตคอลมาตรฐานสำหรับการเปรียบเทียบสถาปัตยกรรมไฮบริด SSM-Transformer กับโมเดล Transformer ล้วน
- คำถามเกี่ยวกับความสามารถในการ scale ของแนวทาง MoE สำหรับงานที่ต้องการการให้เหตุผลเชิงลึกพร้อม expert ต่อ token น้อย
- ข้อมูลเกี่ยวกับ Super/Ultra ณ เดือนธันวาคม 2025 ยังเป็นเบื้องต้น benchmark ครบถ้วนคาดว่าจะมีหลังการเผยแพร่<sup>[\[1\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-N3_whitepaper-1)</sup>

## ด้านจริยธรรมและกฎระเบียบ

### ความปลอดภัยในขั้นตอนการพัฒนา

ในขั้นตอนการพัฒนาใช้: การประเมินตาม framework AEGIS (13 หมวดหมู่ความปลอดภัยเนื้อหา), เครื่องสแกนช่องโหว่ garak, การทดสอบด้วยมือ («red-teaming»)<sup>[\[37\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-Safety-37)</sup>

### ความปลอดภัยในขั้นตอน inference

**NeMo Guardrails** — ชุดเครื่องมือโอเพนซอร์ส (สิทธิ์ Apache 2.0) ที่เพิ่ม barrier แบบโปรแกรมได้ให้กับระบบ LLM microservice ดักจับ input และ output โดยใช้การตรวจสอบที่กำหนดค่าได้ ได้แก่ การควบคุมหัวข้อ การตรวจจับ PII การตรวจสอบ «การยึดโยง» ของ RAG การตรวจจับการโจมตี jailbreak (รวมถึงการป้องกันการ inject โค้ดบนพื้นฐาน YARA) การจำแนกความปลอดภัยแบบหลายภาษาและหลายโมดัล<sup>[\[45\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-Guardrails-45)</sup>

สำหรับ Nemotron 3 มีการเผยแพร่ชุด trace OpenTelemetry ประมาณ 11,000 รายการที่ระบุ label จากสถานการณ์จริงที่ใช้เครื่องมือ เพื่อใช้ประเมินความปลอดภัยของระบบเอเจนต์<sup>[\[1\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-N3_whitepaper-1)</sup>

### การอนุญาตสิทธิ์

| โมเดล                                  | สิทธิ์การใช้งาน                                   |
|----------------------------------------|-----------------------------------------------|
| Nemotron 3 (Nano, Super, Ultra)        | NVIDIA Nemotron Open Model License            |
| Nemotron-4 340B (Base/Instruct/Reward) | NVIDIA Open Model License Agreement           |
| Llama-Nemotron (Nano/Super/Ultra)      | Llama Community License (สืบทอดจาก Meta)       |
| Nemotron-3 8B (2023)                   | NVIDIA AI Foundation Models Community License |

**NVIDIA Nemotron Open Model License** อนุญาตการใช้งานเชิงพาณิชย์ การสร้างและเผยแพร่งานดัดแปลง ไม่ต้องการการอ้างอิง (เมื่อเก็บไฟล์ NOTICE ไว้) ไม่มีข้อจำกัดจำนวนผู้ใช้ และอนุญาตอย่างชัดเจนให้ใช้ output ของโมเดลสำหรับการฝึกโมเดลอื่น<sup>[\[46\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-License-46)</sup> โมเดลที่อิงจาก Llama สืบทอดข้อจำกัดของ Meta รวมถึงเกณฑ์ผู้ใช้งานที่ active รายเดือน 700 ล้านคน<sup>[\[11\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-Llama_Nemotron-11)</sup>

สำหรับ Nemotron 3 Nano NVIDIA เผยแพร่: น้ำหนักโมเดล, ข้อมูลการฝึกมากกว่า 10 ล้านล้าน token, สูตรการฝึก, codebase (NeMo, Megatron-LM, NeMo-Aligner), สภาพแวดล้อมการเรียนรู้ด้วยการเสริมแรงมากกว่า 10 รายการพร้อมงานมากกว่า 900,000 รายการ<sup>[\[1\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-N3_whitepaper-1)[\[2\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-N3_nano_report-2)</sup>

## แนวโน้มและทิศทางการวิจัย

การเผยแพร่ที่ใกล้จะมาถึงได้แก่ **Nemotron 3 Super** (~100 พันล้านพารามิเตอร์, ~10 พันล้าน active) และ **Nemotron 3 Ultra** (~500 พันล้าน, ~50 พันล้าน active) ที่คาดว่าจะมีในช่วงครึ่งแรกของปี 2026 ทั้งสองโมเดลจะใช้ LatentMoE, MTP และการฝึกในรูปแบบ NVFP4 บนสถาปัตยกรรม Blackwell<sup>[\[1\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-N3_whitepaper-1)</sup>

ทิศทางเชิงกลยุทธ์ของ NVIDIA — การวาง Nemotron ไว้ไม่ใช่เป็นโมเดลเดี่ยว แต่เป็นชั้นโครงสร้างพื้นฐานสำหรับระบบ multi-agent ที่โมเดลต่างๆ ในตระกูลใช้สำหรับงานต่างกันโดยกำหนดเส้นทางตามความซับซ้อน<sup>[\[1\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-N3_whitepaper-1)[\[15\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-NVIDIA_news-15)</sup>

ทิศทางการวิจัยหลัก:

- การพัฒนาสถาปัตยกรรมไฮบริด — การรวม SSM layer กับกลไก attention เพิ่มเติมเพื่อ context ยาวที่ scale ได้<sup>[\[12\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-Nemotron_H-12)</sup>
- วิธีการบีบอัดหลายระดับ — การพัฒนา Minitron, Puzzle, MiniPuzzle และ Nemotron Elastic<sup>[\[20\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-Minitron-20)[\[21\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-Puzzle-21)[\[30\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-Elastic-30)</sup>
- การเรียนรู้ด้วยการเสริมแรงแบบหลายโดเมน — Nemotron-CrossThink เพื่อขยาย RL นอกเหนือจากงานคณิตศาสตร์<sup>[\[31\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-CrossThink-31)</sup>
- การขยาย context — Nemotron-UltraLong ถึง 4 ล้าน token<sup>[\[32\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-UltraLong-32)</sup>
- ความหลายโมดัล — การขยายสู่ vision-language (Nemotron VL), เสียงพูด (Nemotron Speech), การค้นหาเอกสารด้วยภาพ (Nemotron ColEmbed V2)<sup>[\[34\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-NanoV2_VL-34)[\[35\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-Parse-35)[\[36\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-ColEmbed-36)</sup>
- โมเดลไฮบริดขนาดกะทัดรัด — Jet-Nemotron (NAS สำหรับโมเดล 2B/4B)<sup>[\[33\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-JetNemotron-33)</sup>
- สูตรเปิด — การเผยแพร่สูตรการฝึกและข้อมูลครบถ้วนเพื่อให้ชุมชนทำซ้ำและปรับแต่งได้<sup>[\[14\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_note-N3_blog-14)</sup>

## ดูเพิ่มเติม

- สถาปัตยกรรม Transformer
- Reinforcement Learning from Human Feedback (RLHF)
- Llama (ตระกูลโมเดลของ Meta)

## อ้างอิง

- NVIDIA Research — หน้า Nemotron 3: <a href="https://research.nvidia.com/labs/nemotron/Nemotron-3/" class="external free" rel="nofollow">https://research.nvidia.com/labs/nemotron/Nemotron-3/</a>
- NVIDIA Developer — Nemotron AI Models: <a href="https://developer.nvidia.com/nemotron" class="external free" rel="nofollow">https://developer.nvidia.com/nemotron</a>
- Hugging Face — เอกสาร Nemotron: <a href="https://huggingface.co/docs/transformers/main/en/model_doc/nemotron" class="external free" rel="nofollow">https://huggingface.co/docs/transformers/main/en/model_doc/nemotron</a>
- NeMo Framework Documentation: <a href="https://docs.nvidia.com/nemo/" class="external free" rel="nofollow">https://docs.nvidia.com/nemo/</a>

## บรรณานุกรม

- NVIDIA. *Nemotron-3-8B-Base-4k — Model Card*. Hugging Face, 2023. <a href="https://huggingface.co/nvidia/nemotron-3-8b-base-4k" class="external free" rel="nofollow">https://huggingface.co/nvidia/nemotron-3-8b-base-4k</a>
- Parmar, J. et al. (NVIDIA). *Nemotron-4 15B Technical Report*. arXiv:2402.16819, กุมภาพันธ์ 2024. <a href="https://arxiv.org/abs/2402.16819" class="external free" rel="nofollow">https://arxiv.org/abs/2402.16819</a>
- Adler, B. et al. (NVIDIA). *Nemotron-4 340B Technical Report*. arXiv:2406.11704, มิถุนายน 2024. <a href="https://arxiv.org/abs/2406.11704" class="external free" rel="nofollow">https://arxiv.org/abs/2406.11704</a>
- Wang, Z. et al. (NVIDIA). *HelpSteer2: Open-source Dataset for Training Top-Performing Reward Models*. arXiv:2406.08673, มิถุนายน 2024. <a href="https://arxiv.org/abs/2406.08673" class="external free" rel="nofollow">https://arxiv.org/abs/2406.08673</a>
- Muralidharan, S. et al. (NVIDIA). *Compact Language Models via Pruning and Knowledge Distillation*. arXiv:2407.14679, กรกฎาคม 2024. <a href="https://arxiv.org/abs/2407.14679" class="external free" rel="nofollow">https://arxiv.org/abs/2407.14679</a>
- Bercovich, A. et al. (NVIDIA). *Puzzle: Distillation-Based NAS for Inference-Optimized LLMs*. arXiv:2411.19146, พฤศจิกายน 2024. <a href="https://arxiv.org/abs/2411.19146" class="external free" rel="nofollow">https://arxiv.org/abs/2411.19146</a>
- Su, D. et al. (NVIDIA). *Nemotron-CC: Transforming Common Crawl into a Refined Long-Horizon Pretraining Dataset*. ACL 2025. arXiv:2412.02595. <a href="https://arxiv.org/abs/2412.02595" class="external free" rel="nofollow">https://arxiv.org/abs/2412.02595</a>
- Sun, S. et al. (NVIDIA). *Reward-aware Preference Optimization*. arXiv:2502.00203, มกราคม 2025. <a href="https://arxiv.org/abs/2502.00203" class="external free" rel="nofollow">https://arxiv.org/abs/2502.00203</a>
- Blakeman, A. et al. (NVIDIA). *Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models*. arXiv:2504.03624, เมษายน 2025. <a href="https://arxiv.org/abs/2504.03624" class="external free" rel="nofollow">https://arxiv.org/abs/2504.03624</a>
- Bercovich, A. et al. (NVIDIA). *Llama-Nemotron: Efficient Reasoning Models*. arXiv:2505.00949, พฤษภาคม 2025. <a href="https://arxiv.org/abs/2505.00949" class="external free" rel="nofollow">https://arxiv.org/abs/2505.00949</a>
- Basant, A. et al. (NVIDIA). *NVIDIA Nemotron Nano 2*. arXiv:2508.14444, สิงหาคม 2025. <a href="https://arxiv.org/abs/2508.14444" class="external free" rel="nofollow">https://arxiv.org/abs/2508.14444</a>
- Karimi Mahabadi, R. et al. (NVIDIA). *Nemotron-CC-Math*. arXiv:2508.15096, สิงหาคม 2025. <a href="https://arxiv.org/abs/2508.15096" class="external free" rel="nofollow">https://arxiv.org/abs/2508.15096</a>
- Taghibakhshi, A. et al. (NVIDIA). *Nemotron Elastic*. arXiv:2511.16664, พฤศจิกายน 2025. <a href="https://arxiv.org/abs/2511.16664" class="external free" rel="nofollow">https://arxiv.org/abs/2511.16664</a>
- Blakeman, A. et al. (NVIDIA). *NVIDIA Nemotron 3: Efficient and Open Intelligence*. arXiv:2512.20856, ธันวาคม 2025. <a href="https://arxiv.org/abs/2512.20856" class="external free" rel="nofollow">https://arxiv.org/abs/2512.20856</a>
- Blakeman, A. et al. (NVIDIA). *Nemotron 3 Nano*. arXiv:2512.20848, ธันวาคม 2025. <a href="https://arxiv.org/abs/2512.20848" class="external free" rel="nofollow">https://arxiv.org/abs/2512.20848</a>
- Dao, T.; Gu, A. *Transformers are SSMs*. ICML 2024. arXiv:2405.21060. <a href="https://arxiv.org/abs/2405.21060" class="external free" rel="nofollow">https://arxiv.org/abs/2405.21060</a>
- Ainslie, J. et al. *GQA*. arXiv:2305.13245, 2023. <a href="https://arxiv.org/abs/2305.13245" class="external free" rel="nofollow">https://arxiv.org/abs/2305.13245</a>
- Su, J. et al. *RoFormer (RoPE)*. Neurocomputing, 2024. arXiv:2104.09864. <a href="https://arxiv.org/abs/2104.09864" class="external free" rel="nofollow">https://arxiv.org/abs/2104.09864</a>
- Zhang, B.; Sennrich, R. *RMSNorm*. arXiv:1910.07467, 2019. <a href="https://arxiv.org/abs/1910.07467" class="external free" rel="nofollow">https://arxiv.org/abs/1910.07467</a>
- Rafailov, R. et al. *Direct Preference Optimization*. NeurIPS 2023. arXiv:2305.18290. <a href="https://arxiv.org/abs/2305.18290" class="external free" rel="nofollow">https://arxiv.org/abs/2305.18290</a>

## หมายเหตุ

1.  <span id="cite_note-N3_whitepaper-1">↑ <sup>[1.00](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-N3_whitepaper_1-0)</sup> <sup>[1.01](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-N3_whitepaper_1-1)</sup> <sup>[1.02](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-N3_whitepaper_1-2)</sup> <sup>[1.03](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-N3_whitepaper_1-3)</sup> <sup>[1.04](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-N3_whitepaper_1-4)</sup> <sup>[1.05](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-N3_whitepaper_1-5)</sup> <sup>[1.06](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-N3_whitepaper_1-6)</sup> <sup>[1.07](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-N3_whitepaper_1-7)</sup> <sup>[1.08](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-N3_whitepaper_1-8)</sup> <sup>[1.09](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-N3_whitepaper_1-9)</sup> <sup>[1.10](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-N3_whitepaper_1-10)</sup> <sup>[1.11](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-N3_whitepaper_1-11)</sup> <sup>[1.12](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-N3_whitepaper_1-12)</sup> <sup>[1.13](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-N3_whitepaper_1-13)</sup> <sup>[1.14](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-N3_whitepaper_1-14)</sup> Blakeman, A. et al. (NVIDIA). *NVIDIA Nemotron 3: Efficient and Open Intelligence*. arXiv:2512.20856, декабрь 2025. <a href="https://arxiv.org/abs/2512.20856" class="external free" rel="nofollow">https://arxiv.org/abs/2512.20856</a></span>
2.  <span id="cite_note-N3_nano_report-2">↑ <sup>[2.00](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-N3_nano_report_2-0)</sup> <sup>[2.01](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-N3_nano_report_2-1)</sup> <sup>[2.02](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-N3_nano_report_2-2)</sup> <sup>[2.03](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-N3_nano_report_2-3)</sup> <sup>[2.04](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-N3_nano_report_2-4)</sup> <sup>[2.05](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-N3_nano_report_2-5)</sup> <sup>[2.06](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-N3_nano_report_2-6)</sup> <sup>[2.07](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-N3_nano_report_2-7)</sup> <sup>[2.08](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-N3_nano_report_2-8)</sup> <sup>[2.09](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-N3_nano_report_2-9)</sup> <sup>[2.10](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-N3_nano_report_2-10)</sup> <sup>[2.11](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-N3_nano_report_2-11)</sup> <sup>[2.12](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-N3_nano_report_2-12)</sup> <sup>[2.13](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-N3_nano_report_2-13)</sup> <sup>[2.14](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-N3_nano_report_2-14)</sup> <sup>[2.15](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-N3_nano_report_2-15)</sup> <sup>[2.16](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-N3_nano_report_2-16)</sup> Blakeman, A. et al. (NVIDIA). *Nemotron 3 Nano: Open, Efficient Mixture‑of‑Experts Hybrid Mamba‑Transformer Model for Agentic Reasoning*. arXiv:2512.20848, декабрь 2025. <a href="https://arxiv.org/abs/2512.20848" class="external free" rel="nofollow">https://arxiv.org/abs/2512.20848</a></span>
3.  <span id="cite_note-N4_340B-3">↑ <sup>[3.00](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-N4_340B_3-0)</sup> <sup>[3.01](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-N4_340B_3-1)</sup> <sup>[3.02](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-N4_340B_3-2)</sup> <sup>[3.03](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-N4_340B_3-3)</sup> <sup>[3.04](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-N4_340B_3-4)</sup> <sup>[3.05](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-N4_340B_3-5)</sup> <sup>[3.06](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-N4_340B_3-6)</sup> <sup>[3.07](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-N4_340B_3-7)</sup> <sup>[3.08](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-N4_340B_3-8)</sup> <sup>[3.09](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-N4_340B_3-9)</sup> <sup>[3.10](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-N4_340B_3-10)</sup> <sup>[3.11](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-N4_340B_3-11)</sup> <sup>[3.12](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-N4_340B_3-12)</sup> <sup>[3.13](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-N4_340B_3-13)</sup> <sup>[3.14](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-N4_340B_3-14)</sup> Adler, B. et al. (NVIDIA). *Nemotron‑4 340B Technical Report*. arXiv:2406.11704, июнь 2024. <a href="https://arxiv.org/abs/2406.11704" class="external free" rel="nofollow">https://arxiv.org/abs/2406.11704</a></span>
4.  <span id="cite_note-NV_developer-4">↑ <sup>[4.0](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-NV_developer_4-0)</sup> <sup>[4.1](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-NV_developer_4-1)</sup> NVIDIA Developer. *NVIDIA Nemotron AI Models*. <a href="https://developer.nvidia.com/nemotron" class="external free" rel="nofollow">https://developer.nvidia.com/nemotron</a></span>
5.  <span id="cite_note-NeMo_docs-5">[↑](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-NeMo_docs_5-0) NVIDIA. *NeMo Framework Documentation*. <a href="https://docs.nvidia.com/nemo/" class="external free" rel="nofollow">https://docs.nvidia.com/nemo/</a></span>
6.  <span id="cite_note-N3_8B_card-6">[↑](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-N3_8B_card_6-0) NVIDIA. *Nemotron‑3‑8B‑Base‑4k — Model Card*. Hugging Face, 2023. <a href="https://huggingface.co/nvidia/nemotron-3-8b-base-4k" class="external free" rel="nofollow">https://huggingface.co/nvidia/nemotron-3-8b-base-4k</a></span>
7.  <span id="cite_note-MS_nemotron3-7">[↑](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-MS_nemotron3_7-0) Microsoft. *Introducing NVIDIA Nemotron‑3 8B LLMs on the Model Catalog*. Microsoft Tech Community, 14 ноября 2023. <a href="https://techcommunity.microsoft.com/blog/machinelearningblog/introducing-nvidia-nemotron-3-8b-llms-on-the-model-catalog/3983569" class="external free" rel="nofollow">https://techcommunity.microsoft.com/blog/machinelearningblog/introducing-nvidia-nemotron-3-8b-llms-on-the-model-catalog/3983569</a></span>
8.  <span id="cite_note-N4_15B-8">↑ <sup>[8.0](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-N4_15B_8-0)</sup> <sup>[8.1](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-N4_15B_8-1)</sup> <sup>[8.2](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-N4_15B_8-2)</sup> <sup>[8.3](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-N4_15B_8-3)</sup> <sup>[8.4](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-N4_15B_8-4)</sup> <sup>[8.5](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-N4_15B_8-5)</sup> <sup>[8.6](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-N4_15B_8-6)</sup> Parmar, J., Prabhumoye, S., Jennings, J. et al. (NVIDIA). *Nemotron‑4 15B Technical Report*. arXiv:2402.16819, февраль 2024. <a href="https://arxiv.org/abs/2402.16819" class="external free" rel="nofollow">https://arxiv.org/abs/2402.16819</a></span>
9.  <span id="cite_note-LN70B_Instruct-9">↑ <sup>[9.0](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-LN70B_Instruct_9-0)</sup> <sup>[9.1](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-LN70B_Instruct_9-1)</sup> <sup>[9.2](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-LN70B_Instruct_9-2)</sup> NVIDIA. *Llama‑3.1‑Nemotron‑70B‑Instruct — Model Card*. Hugging Face, октябрь 2024. <a href="https://huggingface.co/nvidia/Llama-3.1-Nemotron-70B-Instruct-HF" class="external free" rel="nofollow">https://huggingface.co/nvidia/Llama-3.1-Nemotron-70B-Instruct-HF</a></span>
10. <span id="cite_note-LN70B_Reward-10">↑ <sup>[10.0](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-LN70B_Reward_10-0)</sup> <sup>[10.1](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-LN70B_Reward_10-1)</sup> NVIDIA. *Llama‑3.1‑Nemotron‑70B‑Reward — Model Card*. Hugging Face, октябрь 2024. <a href="https://huggingface.co/nvidia/Llama-3.1-Nemotron-70B-Reward" class="external free" rel="nofollow">https://huggingface.co/nvidia/Llama-3.1-Nemotron-70B-Reward</a></span>
11. <span id="cite_note-Llama_Nemotron-11">↑ <sup>[11.0](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-Llama_Nemotron_11-0)</sup> <sup>[11.1](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-Llama_Nemotron_11-1)</sup> <sup>[11.2](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-Llama_Nemotron_11-2)</sup> <sup>[11.3](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-Llama_Nemotron_11-3)</sup> <sup>[11.4](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-Llama_Nemotron_11-4)</sup> <sup>[11.5](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-Llama_Nemotron_11-5)</sup> <sup>[11.6](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-Llama_Nemotron_11-6)</sup> <sup>[11.7](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-Llama_Nemotron_11-7)</sup> <sup>[11.8](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-Llama_Nemotron_11-8)</sup> Bercovich, A. et al. (NVIDIA). *Llama‑Nemotron: Efficient Reasoning Models*. arXiv:2505.00949, май 2025. <a href="https://arxiv.org/abs/2505.00949" class="external free" rel="nofollow">https://arxiv.org/abs/2505.00949</a></span>
12. <span id="cite_note-Nemotron_H-12">↑ <sup>[12.0](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-Nemotron_H_12-0)</sup> <sup>[12.1](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-Nemotron_H_12-1)</sup> <sup>[12.2](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-Nemotron_H_12-2)</sup> <sup>[12.3](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-Nemotron_H_12-3)</sup> <sup>[12.4](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-Nemotron_H_12-4)</sup> <sup>[12.5](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-Nemotron_H_12-5)</sup> <sup>[12.6](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-Nemotron_H_12-6)</sup> <sup>[12.7](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-Nemotron_H_12-7)</sup> <sup>[12.8](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-Nemotron_H_12-8)</sup> Blakeman, A. et al. (NVIDIA). *Nemotron‑H: A Family of Accurate and Efficient Hybrid Mamba‑Transformer Models*. arXiv:2504.03624, апрель 2025. <a href="https://arxiv.org/abs/2504.03624" class="external free" rel="nofollow">https://arxiv.org/abs/2504.03624</a></span>
13. <span id="cite_note-Nano2-13">↑ <sup>[13.0](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-Nano2_13-0)</sup> <sup>[13.1](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-Nano2_13-1)</sup> <sup>[13.2](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-Nano2_13-2)</sup> Basant, A. et al. (NVIDIA). *NVIDIA Nemotron Nano 2: An Accurate and Efficient Hybrid Mamba‑Transformer Reasoning Model*. arXiv:2508.14444, август 2025. <a href="https://arxiv.org/abs/2508.14444" class="external free" rel="nofollow">https://arxiv.org/abs/2508.14444</a></span>
14. <span id="cite_note-N3_blog-14">↑ <sup>[14.0](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-N3_blog_14-0)</sup> <sup>[14.1](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-N3_blog_14-1)</sup> <sup>[14.2](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-N3_blog_14-2)</sup> Blakeman, A. et al. (NVIDIA). *Inside NVIDIA Nemotron 3: Techniques, Tools, and Data That Make It Efficient and Accurate*. NVIDIA Technical Blog, 15 декабря 2025. <a href="https://developer.nvidia.com/blog/inside-nvidia-nemotron-3-techniques-tools-and-data-that-make-it-efficient-and-accurate/" class="external free" rel="nofollow">https://developer.nvidia.com/blog/inside-nvidia-nemotron-3-techniques-tools-and-data-that-make-it-efficient-and-accurate/</a></span>
15. <span id="cite_note-NVIDIA_news-15">↑ <sup>[15.0](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-NVIDIA_news_15-0)</sup> <sup>[15.1](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-NVIDIA_news_15-1)</sup> <sup>[15.2](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-NVIDIA_news_15-2)</sup> <sup>[15.3](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-NVIDIA_news_15-3)</sup> NVIDIA. *NVIDIA Debuts Nemotron 3 Family of Open Models*. NVIDIA Newsroom, 15 декабря 2025. <a href="https://nvidianews.nvidia.com/news/nvidia-debuts-nemotron-3-family-of-open-models" class="external free" rel="nofollow">https://nvidianews.nvidia.com/news/nvidia-debuts-nemotron-3-family-of-open-models</a></span>
16. <span id="cite_note-GQA-16">[↑](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-GQA_16-0) Ainslie, J. et al. *GQA: Training Generalized Multi‑Query Transformer Models from Multi‑Head Checkpoints*. arXiv:2305.13245, 2023. <a href="https://arxiv.org/abs/2305.13245" class="external free" rel="nofollow">https://arxiv.org/abs/2305.13245</a></span>
17. <span id="cite_note-RoPE-17">[↑](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-RoPE_17-0) Su, J. et al. *RoFormer: Enhanced Transformer with Rotary Position Embedding*. Neurocomputing, 2024. arXiv:2104.09864. <a href="https://arxiv.org/abs/2104.09864" class="external free" rel="nofollow">https://arxiv.org/abs/2104.09864</a></span>
18. <span id="cite_note-Mamba2-18">[↑](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-Mamba2_18-0) Dao, T.; Gu, A. *Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality*. ICML 2024. arXiv:2405.21060. <a href="https://arxiv.org/abs/2405.21060" class="external free" rel="nofollow">https://arxiv.org/abs/2405.21060</a></span>
19. <span id="cite_note-RMSNorm-19">[↑](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-RMSNorm_19-0) Zhang, B.; Sennrich, R. *Root Mean Square Layer Normalization*. arXiv:1910.07467, 2019. <a href="https://arxiv.org/abs/1910.07467" class="external free" rel="nofollow">https://arxiv.org/abs/1910.07467</a></span>
20. <span id="cite_note-Minitron-20">↑ <sup>[20.0](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-Minitron_20-0)</sup> <sup>[20.1](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-Minitron_20-1)</sup> Muralidharan, S. et al. (NVIDIA). *Compact Language Models via Pruning and Knowledge Distillation*. arXiv:2407.14679, июль 2024. <a href="https://arxiv.org/abs/2407.14679" class="external free" rel="nofollow">https://arxiv.org/abs/2407.14679</a></span>
21. <span id="cite_note-Puzzle-21">↑ <sup>[21.0](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-Puzzle_21-0)</sup> <sup>[21.1](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-Puzzle_21-1)</sup> Bercovich, A. et al. (NVIDIA). *Puzzle: Distillation‑Based NAS for Inference‑Optimized LLMs*. arXiv:2411.19146, ноябрь 2024. <a href="https://arxiv.org/abs/2411.19146" class="external free" rel="nofollow">https://arxiv.org/abs/2411.19146</a></span>
22. <span id="cite_note-HelpSteer-22">[↑](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-HelpSteer_22-0) Wang, Z. et al. (NVIDIA). *HelpSteer: Multi‑attribute Helpfulness Dataset for SteerLM*. arXiv:2311.09528, ноябрь 2023. <a href="https://arxiv.org/abs/2311.09528" class="external free" rel="nofollow">https://arxiv.org/abs/2311.09528</a></span>
23. <span id="cite_note-HelpSteer2-23">[↑](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-HelpSteer2_23-0) Wang, Z. et al. (NVIDIA). *HelpSteer2: Open‑source Dataset for Training Top‑Performing Reward Models*. arXiv:2406.08673, июнь 2024. <a href="https://arxiv.org/abs/2406.08673" class="external free" rel="nofollow">https://arxiv.org/abs/2406.08673</a></span>
24. <span id="cite_note-HelpSteer2_Pref-24">[↑](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-HelpSteer2_Pref_24-0) Wang, Z. et al. (NVIDIA). *HelpSteer2‑Preference: Complementing Ratings with Preferences*. arXiv:2410.01257, октябрь 2024. <a href="https://arxiv.org/abs/2410.01257" class="external free" rel="nofollow">https://arxiv.org/abs/2410.01257</a></span>
25. <span id="cite_note-SteerLM-25">[↑](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-SteerLM_25-0) Dong, Y. et al. (NVIDIA). *SteerLM: Attribute Conditioned SFT as an (User‑Steerable) Alternative to RLHF*. arXiv:2310.05344, октябрь 2023. <a href="https://arxiv.org/abs/2310.05344" class="external free" rel="nofollow">https://arxiv.org/abs/2310.05344</a></span>
26. <span id="cite_note-DPO-26">[↑](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-DPO_26-0) Rafailov, R. et al. *Direct Preference Optimization: Your Language Model is Secretly a Reward Model*. NeurIPS 2023. arXiv:2305.18290. <a href="https://arxiv.org/abs/2305.18290" class="external free" rel="nofollow">https://arxiv.org/abs/2305.18290</a></span>
27. <span id="cite_note-RPO-27">↑ <sup>[27.0](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-RPO_27-0)</sup> <sup>[27.1](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-RPO_27-1)</sup> Sun, S. et al. (NVIDIA). *Reward‑aware Preference Optimization: A Unified Mathematical Framework for Model Alignment*. arXiv:2502.00203, январь 2025. <a href="https://arxiv.org/abs/2502.00203" class="external free" rel="nofollow">https://arxiv.org/abs/2502.00203</a></span>
28. <span id="cite_note-NeMo_eval-28">[↑](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-NeMo_eval_28-0) NVIDIA. *NeMo Evaluator SDK — Evaluation Recipe for Nemotron 3 Nano*. Hugging Face Blog, декабрь 2025. <a href="https://huggingface.co/blog/nvidia/nemotron-3-nano-evaluation-recipe" class="external free" rel="nofollow">https://huggingface.co/blog/nvidia/nemotron-3-nano-evaluation-recipe</a></span>
29. <span id="cite_note-OpenReasoning-29">[↑](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-OpenReasoning_29-0) NVIDIA. *OpenReasoning‑Nemotron*. Hugging Face Collection, июль 2025. <a href="https://huggingface.co/collections/nvidia/openreasoning-nemotron-685824d42db3e24d8b8f5e39" class="external free" rel="nofollow">https://huggingface.co/collections/nvidia/openreasoning-nemotron-685824d42db3e24d8b8f5e39</a></span>
30. <span id="cite_note-Elastic-30">↑ <sup>[30.0](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-Elastic_30-0)</sup> <sup>[30.1](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-Elastic_30-1)</sup> Taghibakhshi, A. et al. (NVIDIA). *Nemotron Elastic: Towards Efficient Many‑in‑One Reasoning LLMs*. arXiv:2511.16664, ноябрь 2025. <a href="https://arxiv.org/abs/2511.16664" class="external free" rel="nofollow">https://arxiv.org/abs/2511.16664</a></span>
31. <span id="cite_note-CrossThink-31">↑ <sup>[31.0](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-CrossThink_31-0)</sup> <sup>[31.1](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-CrossThink_31-1)</sup> Akter, S. N. et al. (NVIDIA). *Nemotron‑CrossThink: Scaling Self‑Learning beyond Math Reasoning*. arXiv:2504.13941, апрель 2025. <a href="https://arxiv.org/abs/2504.13941" class="external free" rel="nofollow">https://arxiv.org/abs/2504.13941</a></span>
32. <span id="cite_note-UltraLong-32">↑ <sup>[32.0](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-UltraLong_32-0)</sup> <sup>[32.1](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-UltraLong_32-1)</sup> <sup>[32.2](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-UltraLong_32-2)</sup> Xu, C. et al. (NVIDIA). *From 128K to 4M: Efficient Training of Ultra‑Long Context Large Language Models*. arXiv:2504.06214, апрель 2025. <a href="https://arxiv.org/abs/2504.06214" class="external free" rel="nofollow">https://arxiv.org/abs/2504.06214</a></span>
33. <span id="cite_note-JetNemotron-33">↑ <sup>[33.0](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-JetNemotron_33-0)</sup> <sup>[33.1](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-JetNemotron_33-1)</sup> Gu, Y. et al. (NVIDIA). *Jet‑Nemotron: Efficient Language Model with Post Neural Architecture Search*. arXiv:2508.15884, август 2025. <a href="https://arxiv.org/abs/2508.15884" class="external free" rel="nofollow">https://arxiv.org/abs/2508.15884</a></span>
34. <span id="cite_note-NanoV2_VL-34">↑ <sup>[34.0](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-NanoV2_VL_34-0)</sup> <sup>[34.1](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-NanoV2_VL_34-1)</sup> Deshmukh, A. S. et al. (NVIDIA). *NVIDIA Nemotron Nano V2 VL*. arXiv:2511.03929, ноябрь 2025. <a href="https://arxiv.org/abs/2511.03929" class="external free" rel="nofollow">https://arxiv.org/abs/2511.03929</a></span>
35. <span id="cite_note-Parse-35">↑ <sup>[35.0](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-Parse_35-0)</sup> <sup>[35.1](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-Parse_35-1)</sup> Chumachenko, K. et al. (NVIDIA). *NVIDIA Nemotron Parse 1.1*. arXiv:2511.20478, ноябрь 2025. <a href="https://arxiv.org/abs/2511.20478" class="external free" rel="nofollow">https://arxiv.org/abs/2511.20478</a></span>
36. <span id="cite_note-ColEmbed-36">↑ <sup>[36.0](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-ColEmbed_36-0)</sup> <sup>[36.1](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-ColEmbed_36-1)</sup> de Souza P. Moreira, G. et al. (NVIDIA). *Nemotron ColEmbed V2: Top‑Performing Late Interaction Embedding Models for Visual Document Retrieval*. arXiv:2602.03992, февраль 2026. <a href="https://arxiv.org/abs/2602.03992" class="external free" rel="nofollow">https://arxiv.org/abs/2602.03992</a></span>
37. <span id="cite_note-Safety-37">↑ <sup>[37.0](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-Safety_37-0)</sup> <sup>[37.1](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-Safety_37-1)</sup> NVIDIA Developer Blog. *Safeguard Agentic AI Systems with the NVIDIA Safety Recipe*. 2025. <a href="https://developer.nvidia.com/blog/safeguard-agentic-ai-systems-with-the-nvidia-safety-recipe/" class="external free" rel="nofollow">https://developer.nvidia.com/blog/safeguard-agentic-ai-systems-with-the-nvidia-safety-recipe/</a></span>
38. <span id="cite_note-NemotronCC-38">[↑](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-NemotronCC_38-0) Su, D. et al. (NVIDIA). *Nemotron‑CC: Transforming Common Crawl into a Refined Long‑Horizon Pretraining Dataset*. ACL 2025 (Long Paper). arXiv:2412.02595. <a href="https://arxiv.org/abs/2412.02595" class="external free" rel="nofollow">https://arxiv.org/abs/2412.02595</a></span>
39. <span id="cite_note-NemotronCC_Math-39">[↑](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-NemotronCC_Math_39-0) Karimi Mahabadi, R. et al. (NVIDIA). *Nemotron‑CC‑Math: A 133 Billion‑Token‑Scale High Quality Math Pretraining Dataset*. arXiv:2508.15096, август 2025. <a href="https://arxiv.org/abs/2508.15096" class="external free" rel="nofollow">https://arxiv.org/abs/2508.15096</a></span>
40. <span id="cite_note-NemotronNano2_page-40">[↑](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-NemotronNano2_page_40-0) NVIDIA Research. *NVIDIA Nemotron Nano 2 and the Nemotron Pretraining Dataset v1*. <a href="https://research.nvidia.com/labs/adlr/NVIDIA-Nemotron-Nano-2/" class="external free" rel="nofollow">https://research.nvidia.com/labs/adlr/NVIDIA-Nemotron-Nano-2/</a></span>
41. <span id="cite_note-NeMo_synth-41">[↑](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-NeMo_synth_41-0) NVIDIA. *Synthetic Data Generation — NVIDIA NeMo Framework User Guide*. <a href="https://docs.nvidia.com/nemo-framework/user-guide/24.12/datacuration/syntheticdata.html" class="external free" rel="nofollow">https://docs.nvidia.com/nemo-framework/user-guide/24.12/datacuration/syntheticdata.html</a></span>
42. <span id="cite_note-AA-42">↑ <sup>[42.0](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-AA_42-0)</sup> <sup>[42.1](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-AA_42-1)</sup> Artificial Analysis. *NVIDIA Nemotron 3 Nano 30B‑A3B — Intelligence Index*. Февраль 2026. <a href="https://artificialanalysis.ai/models/nvidia-nemotron-3-nano-30b-a3b-reasoning" class="external free" rel="nofollow">https://artificialanalysis.ai/models/nvidia-nemotron-3-nano-30b-a3b-reasoning</a></span>
43. <span id="cite_note-LMArena-43">[↑](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-LMArena_43-0) LMArena (Chatbot Arena). Рейтинги моделей, по состоянию на февраль 2026. <a href="https://lmarena.ai/" class="external free" rel="nofollow">https://lmarena.ai/</a></span>
44. <span id="cite_note-Saplin-44">[↑](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-Saplin_44-0) Saplin, M. *Llama 3.1 Nemotron 70B: Quirks and Features*. DEV Community, 2024. <a href="https://dev.to/maximsaplin/llama-31-nemotron-70b-quirks-and-features-4nbg" class="external free" rel="nofollow">https://dev.to/maximsaplin/llama-31-nemotron-70b-quirks-and-features-4nbg</a></span>
45. <span id="cite_note-Guardrails-45">[↑](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-Guardrails_45-0) NVIDIA. *NeMo Guardrails*. GitHub, 2023–2025. <a href="https://github.com/NVIDIA/NeMo-Guardrails" class="external free" rel="nofollow">https://github.com/NVIDIA/NeMo-Guardrails</a> (arXiv:2310.10501).</span>
46. <span id="cite_note-License-46">[↑](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(TH)#cite_ref-License_46-0) NVIDIA. *NVIDIA Nemotron Open Model License*. <a href="https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-nemotron-open-model-license/" class="external free" rel="nofollow">https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-nemotron-open-model-license/</a></span>
