---
title: "Nemotron (NVIDIA) (NL)"
source: "https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)"
wiki: "systems-analysis.info/int"
article: "Nemotron_(NVIDIA)_(NL)"
language: "nl"
categories:
  - "Category:Dutch"
  - "Category:Large language models"
  - "Category:Machine learning"
revision_id: 4853
wiki_created_at: 2026-09-06T23:40:55Z
wiki_modified_at: 2026-09-06T23:40:55Z
downloaded_at: 2026-09-07T23:05:11Z
---

# Nemotron (NVIDIA) (NL)

**Nemotron** — een familie van grote taalmodellen (Large Language Model, LLM) met open gewichten, ontwikkeld door de afdeling Applied Deep Learning Research (ADLR) van NVIDIA. De modellen behoren tot het domein van machine learning en Natural Language Processing (NLP) en zijn bestemd voor een breed scala aan taken: synthetische datageneratie, redeneren (reasoning), agentische toepassingen (agentic AI) en inzet op industriële NVIDIA-hardware. De familie omvat modellen van verschillende architecturen en schalen: van 4 miljard tot ~500 miljard parameters, waaronder dense decoder‑only Transformer-modellen uit de Nemotron‑4-serie (2024), hybride Mamba‑Transformer-modellen (Nemotron‑H, 2025) en hybride Mamba‑Transformer-modellen met Mixture-of-Experts (MoE) in de Nemotron 3-serie (2025). Vanaf maart 2026 telt de familie meer dan 20 modellen, voor taken als tekstgeneratie, redeneren, visueel documentbegrip, informatieextractie en antwoordkwaliteitsevaluatie. Modellen worden voornamelijk uitgebracht met open gewichten onder de NVIDIA Open Model License en de Llama Community License.<sup>[\[1\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-N3_whitepaper-1)[\[2\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-N3_nano_report-2)[\[3\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-N4_340B-3)</sup>

## Geschiedenis en achtergrond

De ontwikkeling van de Nemotron-familie vindt plaats in het kader van NVIDIA's strategie om een volledige gereedschapsstapel voor generatieve AI te bouwen: van trainingsinfrastructuur (DGX, supercomputers op basis van GPU H100/B200) tot trainingsplatforms (NeMo Framework), uitlijning (NeMo‑Aligner), implementatie (NIM — NVIDIA Inference Microservices) en beveiliging (NeMo Guardrails).<sup>[\[4\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-NV_developer-4)[\[5\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-NeMo_docs-5)</sup>

### Nemotron‑3 8B (2023)

Het eerste publiek beschikbare model van de serie — een decoder-only Transformer met 8 miljard parameters en een context van 4096 tokens, getraind op **3,8 biljoen tokens** in 53 natuurlijke talen en 37 programmeertalen. De training werd uitgevoerd op 1024 GPU A100-eenheden gedurende 19 dagen. Er werd geen formeel technisch rapport op arXiv gepubliceerd. Het model werd verspreid via NeMo Framework en Hugging Face; er bestonden varianten voor Chat (SFT en SteerLM). Licentie — NVIDIA AI Foundation Models Community License.<sup>[\[6\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-N3_8B_card-6)[\[7\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-MS_nemotron3-7)</sup>

*Opmerking over naamgeving*: «Nemotron‑3» uit 2023 en «Nemotron 3» uit december 2025 zijn verschillende modellen. Het eerste was een vroeg prototype, het tweede is de actuele generatie met MoE-architectuur.

### Nemotron‑4 15B (februari 2024)

De eerste publiek gedocumenteerde release van de Nemotron‑4-serie — een model met 15 miljard parameters, getraind op 8 biljoen tokens in ~13 kalenderdagen op 384 DGX H100-knooppunten (tot 3072 GPU's). Er werd gebruik gemaakt van 8-voudige tensorparallellisme en dataparallellisme van 96 tot 288. De piek-MFU (Model FLOPs Utilization) bedroeg 34,3 % bij een batch size van 384. De architectuur is een decoder-only Transformer met Grouped‑Query Attention (GQA) en Rotary Position Embeddings (RoPE). Op basis van benchmarks ten tijde van publicatie overtrof het model alle vergelijkbare open modellen in meertalige taken, inclusief een aantal modellen die vier keer zo groot zijn.<sup>[\[8\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-N4_15B-8)</sup>

### Nemotron‑4 340B (juni 2024)

In juni 2024 werd het grootste model van de Nemotron‑4-serie uitgebracht — 340 miljard parameters (331,6 miljard niet-embedding-parameters), voorgetraind op 9 biljoen tokens (8 biljoen voortraining + 1 biljoen voortgezette training met herwogen hoogwaardige bronnen). De training werd uitgevoerd op 768 DGX H100-knooppunten (tot 6144 GPU's) met een piek-MFU van 42,4 %, van december 2023 tot mei 2024. De familie omvat drie varianten: Base, Instruct en Reward. De modelgrootte is zodanig gekozen dat het model past op één DGX H100-knooppunt (8 GPU's) bij implementatie in FP8-precisieformaat. Meer dan 98 % van de bij uitlijning (alignment) gebruikte data is synthetisch gegenereerd door de modellen van de serie zelf in een iteratief proces; de pipeline voor synthetische datageneratie is opengesteld voor reproductie.<sup>[\[3\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-N4_340B-3)</sup>

### Llama‑3.1‑Nemotron‑70B (oktober 2024)

In oktober 2024 werden modellen uitgebracht die zijn verfijnd op basis van Meta Llama‑3.1‑70B‑Instruct: Llama‑3.1‑Nemotron‑70B‑Instruct en Llama‑3.1‑Nemotron‑70B‑Reward. Per 1 oktober 2024 stond het Instruct-model gelijktijdig op de 1e plaats op drie automatische benchmarks: Arena Hard — 85,0, AlpacaEval 2 LC — 57,6 %, MT‑Bench (GPT‑4‑Turbo) — 8,98. Het Reward-model behaalde 94,1 % op RewardBench.<sup>[\[9\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-LN70B_Instruct-9)[\[10\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-LN70B_Reward-10)</sup>

### Overgang naar hybride architecturen (2025)

In 2025 werd de serie uitgebreid met hybride architecturen die Mamba‑2-lagen (State Space Model, SSM — toestandsruimtemodel) en Transformer combineren. Tussen maart en mei 2025 werd de familie redenerende modellen **Llama‑Nemotron** uitgebracht (Nano 8B, Super 49B, Ultra 253B) met een schakelbare redeneerstand.<sup>[\[11\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-Llama_Nemotron-11)</sup> In april 2025 werd **Nemotron‑H** gepubliceerd (arXiv:2504.03624) — een familie van hybride Mamba‑Transformer-modellen van 8B, 56B en 47B parameters.<sup>[\[12\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-Nemotron_H-12)</sup> In augustus 2025 werd **Nemotron Nano 2** gepresenteerd (9B‑v2, arXiv:2508.14444).<sup>[\[13\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-Nano2-13)</sup>

### Nemotron 3 (december 2025)

Op 15 december 2025 werd de derde generatie aangekondigd — de Nemotron 3-familie met de modellen Nano, Super en Ultra, die Mamba‑2-, Transformer- en MoE-architecturen combineert met een contextvenster van maximaal 1 miljoen tokens. Nano werd uitgebracht met een technisch rapport; Super en Ultra zijn gepland voor de eerste helft van 2026.<sup>[\[1\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-N3_whitepaper-1)[\[2\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-N3_nano_report-2)[\[14\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-N3_blog-14)[\[15\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-NVIDIA_news-15)</sup>

## Theoretische grondslagen

### Transformer-architectuur in Nemotron‑4

De modellen van de Nemotron‑4-serie zijn gebouwd op de standaard decoder-only Transformer-architectuur met causale aandacht. De kans op een tokenreeks $x_{1},x_{2},\ldots,x_{T}$ wordt gefactoriseerd als:

$$
p(x_{1},\ldots,x_{T}) = \prod\limits_{t = 1}^{T}p(x_{t} \mid x_{1},\ldots,x_{t - 1};\theta),
$$

waar $\theta$ — de parameters van het model zijn.<sup>[\[8\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-N4_15B-8)</sup>

Het aandachtsmechanisme is geïmplementeerd als Grouped‑Query Attention (GQA)<sup>[\[16\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-GQA-16)</sup>:

$$
\operatorname{Attention}(Q,K,V) = \operatorname{softmax}\!\left( \frac{QK^{\top}}{\sqrt{d_{k}}} \right)V,
$$

waar $Q,K,V$ — query-, sleutel- en waardematrices zijn; $d_{k}$ — de dimensie van de aandachtskop.

Voor positiecodering worden Rotary Position Embeddings (RoPE) toegepast<sup>[\[17\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-RoPE-17)</sup>:

$$
\operatorname{RoPE}(x_{m},m) = \begin{pmatrix}
{x_{m}^{(1)}\cos m\theta_{1} - x_{m}^{(2)}\sin m\theta_{1}} \\
{x_{m}^{(1)}\sin m\theta_{1} + x_{m}^{(2)}\cos m\theta_{1}} \\
 \vdots 
\end{pmatrix},
$$

waar $x_{m}$ — de vector op positie $m$, $\theta_{i} = 10000^{- 2i/d}$, $d$ — de dimensie van de vector.

In MLP-lagen wordt Squared ReLU-activatie gebruikt: $f(x) = (\max(0,x))^{2}$, wat zorgt voor schaarse activaties. De modellen bevatten geen biassen (bias), gebruiken geen dropout, en de invoer- en uitvoerembeddingmatrices zijn niet aan elkaar gekoppeld (untied embeddings). De tokenizer — SentencePiece BPE met behoud van spaties, cijfer-voor-cijfer opsplitsing en byte-level fallback, woordenschatgrootte — 256.000.<sup>[\[8\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-N4_15B-8)[\[3\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-N4_340B-3)</sup>

##### Architectuurparameters van Nemotron‑4

| Parameter           | Nemotron‑4 15B                  | Nemotron‑4 340B                  |
|---------------------|---------------------------------|----------------------------------|
| Aantal parameters   | 15 miljard (3,2 miljard embed.) | 340 miljard (9,4 miljard embed.) |
| Aantal lagen        | 32                              | 96                               |
| Verborgen dimensie  | 6144                            | 18.432                           |
| Aandachtskoppen     | 48                              | 96                               |
| KV‑koppen (GQA)     | 8                               | 8                                |
| Contextlengte       | 4096                            | 4096                             |
| Woordenschatgrootte | 256.000                         | 256.000                          |

Bronnen: arXiv:2402.16819, arXiv:2406.11704.<sup>[\[8\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-N4_15B-8)[\[3\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-N4_340B-3)</sup>

### Hybride Mamba‑Transformer-architectuur (Nemotron‑H)

In de Nemotron‑H-modellen zijn standaard zelf-aandachtsblokken gedeeltelijk vervangen door Mamba‑2-blokken — toestandsruimtemodellen (State Space Models, SSM). Bij autoregressieve generatie vereist elke zelf-aandachtslaag $O(n)$ bewerkingen en geheugen per stap (vanwege de KV-cache), terwijl Mamba-lagen constante geheugen- en rekenkosten bieden voor elk gegenereerd token.<sup>[\[12\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-Nemotron_H-12)</sup>

Mamba‑2 wordt beschreven door de recurrente betrekking<sup>[\[18\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-Mamba2-18)</sup>:

$$
h_{t} = Ah_{t - 1} + Bx_{t},\quad y_{t} = Ch_{t},
$$

waar $h_{t} \in {\mathbb{R}}^{N}$ — de verborgen toestand, $A \in {\mathbb{R}}^{N \times N}$ — de diagonale toestandsovergangsmatrix, $B \in {\mathbb{R}}^{N \times 1}$ en $C \in {\mathbb{R}}^{1 \times N}$ — projectiematrices, $x_{t}$ — het invoertoken, $y_{t}$ — het uitgangssignaal. In Mamba‑2 zijn de matrices $A$, $B$, $C$ invoerafhankelijk (input‑dependent), wat het model onderscheidt van klassieke lineaire SSM's.

In Nemotron‑H‑56B worden **54 Mamba‑2-lagen + 54 MLP-lagen + 10 zelf-aandachtslagen** gebruikt, waarbij de aandachtslagen gelijkmatig verdeeld zijn tussen de Mamba-lagen. Het model is getraind op 20 biljoen tokens in FP8-formaat (per‑tensor scaling) op 6144 GPU's H100. Versie Nemotron‑H‑8B bevat 24 Mamba‑2-lagen, 24 MLP-lagen en 4 zelf-aandachtslagen; de training werd uitgevoerd op 15 biljoen tokens.<sup>[\[12\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-Nemotron_H-12)</sup>

### MoE-architectuur van Nemotron 3

De Nemotron 3-modellen (december 2025) combineren drie architecturele componenten: Mamba‑2, Transformer en Mixture‑of‑Experts.<sup>[\[1\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-N3_whitepaper-1)[\[2\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-N3_nano_report-2)</sup> Nemotron 3 Nano (30B‑A3B) bevat 52 lagen: **23 Mamba‑2 + MoE-lagen, 23 MoE-lagen en 6 GQA-aandachtslagen**. Elke MoE-laag bevat 128 gerouteerde (routed) experts + 1 gedeelde (shared) expert, waarvan 6 experts geactiveerd worden per token. De routing wordt uitgevoerd door een trainbare MLP-router met sigmoid gating. Lastbalancering is geïmplementeerd zonder hulpverliesfunctie (aux‑loss‑free). Het totale aantal parameters bedraagt 31,6 miljard, maar er zijn slechts 3,2 miljard actief per token.<sup>[\[2\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-N3_nano_report-2)</sup>

Normalisatie — RMSNorm<sup>[\[19\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-RMSNorm-19)</sup>. Positie-embeddings in Mamba-gedeelten ontbreken (positie-informatie is impliciet in de SSM-toestand). Er worden niet-gekoppelde embeddings gebruikt, zonder dropout en zonder bias in lineaire lagen.<sup>[\[2\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-N3_nano_report-2)</sup>

### Uitbreidingen in Super/Ultra: LatentMoE en Multi‑Token Prediction

De Super- en Ultra-modellen uit de Nemotron 3-familie passen twee aanvullende architecturale innovaties toe:

- **Latent MoE (LatentMoE)** — projectie van de invoerrepresentatie naar een latente ruimte met kleinere dimensie vóór routing, waardoor het aantal experts kan worden vergroot bij vergelijkbare rekenkosten en de nauwkeurigheid wordt verhoogd zonder verlies aan doorvoer.<sup>[\[1\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-N3_whitepaper-1)</sup>
- **Multi‑Token Prediction (MTP)** — gelijktijdige voorspelling van meerdere volgende tokens, gebruikt om de kwaliteit van lange teksten te verbeteren en voor speculatief decoderen tijdens inferentie.<sup>[\[1\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-N3_whitepaper-1)</sup>

De training van Super en Ultra vindt plaats in NVFP4-formaat — NVIDIA's 4-bits getalformaat op de Blackwell-architectuur.<sup>[\[1\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-N3_whitepaper-1)</sup>

### Methoden voor modelcompressie: Minitron, Puzzle, MiniPuzzle

Voor het maken van compacte modellen past NVIDIA verschillende benaderingen toe:

- **Minitron** — een methode voor gestructureerd snoeien (pruning) en kennisdistillatie (knowledge distillation). Er worden twee soorten pruning onderzocht: op diepte (verwijdering van lagen) en op breedte (gezamenlijke verkleining van de dimensies van verborgen lagen, aandachtskoppen en MLP-projecties). Toepassing van Minitron op Nemotron‑4 15B levert modellen van 8B en 4B op, met een reductie van trainingskosten tot 40 keer ten opzichte van training vanaf nul.<sup>[\[20\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-Minitron-20)</sup>
- **Puzzle** — een Neural Architecture Search (NAS)-algoritme dat de optimale configuratie van elk blok bepaalt (aantal KV-koppen, FFN-grootte, aanwezigheid/afwezigheid van aandacht) met bloksgewijze distillatie. Op deze manier werd Llama 3.1 405B gecomprimeerd tot 253B (Llama‑Nemotron Ultra) en Llama 3.3 70B tot 49B (Llama‑Nemotron Super).<sup>[\[21\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-Puzzle-21)</sup>
- **MiniPuzzle** — een uitbreiding van Puzzle voor hybride architecturen, waarmee Nemotron‑H‑56B werd gecomprimeerd tot 47B met slechts 63 miljard distillatietokens. Bij vergelijkbare nauwkeurigheid is het gecomprimeerde model 20 % sneller bij inferentie.<sup>[\[12\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-Nemotron_H-12)</sup>

## Uitlijningsmethoden

### HelpSteer en HelpSteer2

**HelpSteer** — een dataset met 37.120 voorbeelden (licentie CC‑BY‑4.0), gecreëerd in samenwerking met Scale AI, waarbij elk antwoord is geannoteerd op vijf attributen: behulpzaamheid (helpfulness), correctheid (correctness), samenhang (coherence), complexiteit (complexity) en uitvoerigheid (verbosity) op een Likert-schaal van 0–4.<sup>[\[22\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-HelpSteer-22)</sup>

**HelpSteer2** — een bijgewerkte dataset met 21.362 voorbeelden (10.681 prompts × 2 antwoorden), waarbij meer dan 95 % van de prompts afkomstig zijn uit ShareGPT (echte gebruikersverzoeken). Ongeveer 50 % van de annotaties werd gefilterd als onvoldoende kwalitatief. De uitbreiding **HelpSteer2‑Preference** voegt paarsgewijze annotatorvoorkeuren toe, waardoor zowel regressie- als paarsgewijze beloningsmodellen op dezelfde data getraind kunnen worden.<sup>[\[23\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-HelpSteer2-23)[\[24\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-HelpSteer2_Pref-24)</sup>

### SteerLM

**SteerLM** — een SFT (Supervised Fine‑Tuning)-methode met conditionering op attributen (helpfulness, correctness, coherence, complexity, verbosity), waarmee de gebruiker de antwoordstijl tijdens inferentie kan sturen. De pipeline omvat: (1) training van een attribuutvoorspellingsmodel (APM) op HelpSteer, (2) annotatie van data met behulp van APM, (3) SFT met conditionering op doelattribuutwaarden, (4) bootstrapping.<sup>[\[25\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-SteerLM-25)</sup>

### DPO en RPO

**DPO (Direct Preference Optimization)** — een methode voor directe voorkeuroptimalisatie zonder expliciet beloningsmodel, gebruikt in de pipelines van Nemotron‑4 340B Instruct en Nemotron Nano 2.<sup>[\[26\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-DPO-26)</sup>

**RPO (Reward‑aware Preference Optimization)** — een geünificeerd wiskundig framework van NVIDIA, dat DPO, IPO, SimPO, REINFORCE en SteerLM 2.0 als bijzondere gevallen generaliseert. RPO minimaliseert de afstand tussen de voorspellingen van het impliciete beloningsmodel en het doelgerichte expliciete model<sup>[\[27\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-RPO-27)</sup>:

$$
\mathcal{L}_{\text{RPO}}(\theta) = {\mathbb{E}}_{(x,y_{w},y_{l}) \sim \mathcal{D}}\left\lbrack d\!\left( r_{\phi}(x,y_{w}) - r_{\phi}(x,y_{l}),\;\beta\log\frac{\pi_{\theta}(y_{w}|x)}{\pi_{\text{ref}}(y_{w}|x)} - \beta\log\frac{\pi_{\theta}(y_{l}|x)}{\pi_{\text{ref}}(y_{l}|x)} \right) \right\rbrack,
$$

waar $r_{\phi}$ — het expliciete beloningsmodel, $\pi_{\theta}$ — het te trainen beleid, $\pi_{\text{ref}}$ — het referentiebeleid, $d( \cdot , \cdot )$ — de afstandsfunctie, $y_{w}$/$y_{l}$ — het geprefereerde/verworpen antwoord. RPO wordt toegepast bij het trainen van Nemotron‑4 340B Instruct en de Llama‑Nemotron-modellen.<sup>[\[27\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-RPO-27)[\[3\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-N4_340B-3)[\[11\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-Llama_Nemotron-11)</sup>

### Multi-omgeving versterkend leren (RLVR)

In Nemotron 3 wordt Multi‑environment RLVR (Reinforcement Learning from Verifiable Rewards) toegepast — gelijktijdig trainen in meerdere omgevingen binnen NeMo Gym: competitieve wiskunde, programmeren, QA, tool‑use, instruction‑following, lange context. Het algoritme is GRPO (Group Relative Policy Optimization) met masked importance sampling.<sup>[\[2\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-N3_nano_report-2)[\[14\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-N3_blog-14)</sup>

Nemotron 3 ondersteunt granulaire **redeneerbudgetbeheersing** (reasoning budget control): de gebruiker kan een tokenlimiet voor de redeneerreeks instellen via een vlag in de chat template (schakelaar «detailed thinking on/off» of beperking van de lengte).<sup>[\[2\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-N3_nano_report-2)</sup>

## Modelreeks

### Overzichtstabel

| Model                         | Jaar      | Totaal / Actieve parameters | Context   | Architectuur                   | Belangrijkste kenmerken                                          |
|-------------------------------|-----------|-----------------------------|-----------|--------------------------------|------------------------------------------------------------------|
| **Nemotron‑3 8B**             | 2023      | 8B / 8B                     | 4K        | Dense Transformer              | 3,8B tokens; 1024 GPU A100; 19 dagen                             |
| **Nemotron‑4 15B**            | 2024      | 15B / 15B                   | 4K        | Dense Transformer              | 8B tokens; MFU 34,3 %                                            |
| **Nemotron‑4 340B**           | 2024      | 340B / 340B                 | 4K        | Dense Transformer              | 9B tokens; Base/Instruct/Reward; \>98 % synthetische uitlijndata |
| **Llama‑3.1‑Nemotron‑70B**    | 2024      | 70B / 70B                   | 128K      | Dense Transformer (Llama 3.1)  | RLHF (REINFORCE); RewardBench 94,1 %                             |
| **Llama‑Nemotron Nano 8B**    | 2025      | 8B / 8B                     | 128K      | Dense Transformer (Llama 3.1)  | Reasoning toggle                                                 |
| **Llama‑Nemotron Nano 4B**    | 2025      | 4B / 4B                     | 128K      | Dense Transformer (Minitron)   | NAS-compressie uit Llama 3.1 8B                                  |
| **Llama‑Nemotron Super 49B**  | 2025      | 49B / 49B                   | 128K      | Dense Transformer (Puzzle NAS) | Uit Llama 3.3 70B                                                |
| **Llama‑Nemotron Ultra 253B** | 2025      | 253B / 253B                 | 128K      | Dense Transformer (Puzzle NAS) | Uit Llama 3.1 405B; GPQA 76,0 %                                  |
| **Nemotron‑H 8B**             | 2025      | 8B / 8B                     | —         | Hybride Mamba‑Transformer      | 15B tokens; 24 Mamba‑2 + 24 MLP + 4 Attn                         |
| **Nemotron‑H 56B**            | 2025      | 56B / 56B                   | —         | Hybride Mamba‑Transformer      | 20B tokens FP8; 54 Mamba‑2 + 54 MLP + 10 Attn                    |
| **Nemotron‑H 47B**            | 2025      | 47B / 47B                   | ~1M (FP4) | Hybride Mamba‑Transformer      | MiniPuzzle uit 56B; +20 % snelheid                               |
| **Nemotron Nano 2 (9B)**      | 2025      | 12B → 9B / 9B               | 128K      | Hybride Mamba‑Transformer      | 20B tokens; Minitron-distillatie                                 |
| **Nemotron 3 Nano**           | 2025      | 31,6B / 3,2B                | 1M        | Hybride Mamba‑Transformer MoE  | 25B tokens; 128 experts, 6 actief                                |
| **Nemotron 3 Super**          | 2025–2026 | ~100B / ~10B                | 1M        | Hybride LatentMoE              | MTP; NVFP4                                                       |
| **Nemotron 3 Ultra**          | 2025–2026 | ~500B / ~50B                | 1M        | Hybride LatentMoE              | MTP; NVFP4                                                       |

### Nemotron‑4 15B

Architectuur: 32 lagen, hidden dimension 6144, 48 aandachtskoppen, 8 KV‑koppen (GQA). Totaal aantal parameters — 15 miljard (3,2 miljard embedding + 12,5 miljard non‑embedding). Resultaten: MMLU — 64,2 % (5‑shot), BBH — 58,7 % (3‑shot), GSM8K — 46,0 % (8‑shot, maj@1), HumanEval — 31,6 % (0‑shot, pass@1). Op meertalige benchmarks (XCOPA, TyDiQA‑GoldP, MGSM) overtrof het model modellen die vier keer zo groot zijn.<sup>[\[8\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-N4_15B-8)</sup>

### Nemotron‑4 340B

Architectuur: 96 lagen, hidden dimension 18.432, 96 aandachtskoppen, 8 KV‑koppen.

**Nemotron‑4 340B Base** behaalde: MMLU — 81,1 % (5‑shot), BBH — 85,4 % (3‑shot), HumanEval — 57,3 % (0‑shot), ARC‑Challenge — 94,3 % (25‑shot).<sup>[\[3\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-N4_340B-3)</sup>

**Nemotron‑4 340B Instruct** doorliep meerfasige uitlijning: Code SFT → General SFT → DPO → RPO (3 rondes). Resultaten: MT‑Bench — 8,22 (GPT‑4‑Turbo judge), MMLU — 78,7 % (0‑shot), GSM8K — 92,3 % (0‑shot), HumanEval — 73,2 % (0‑shot), Arena Hard — 54,2, AlpacaEval 2.0 LC — 41,5 %.<sup>[\[3\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-N4_340B-3)</sup>

**Nemotron‑4 340B Reward** vervangt de finale softmax-laag door een lineaire projectie van de verborgen toestand van de laatste laag naar een vector van 5 HelpSteer2-attributen. De uiteindelijke beloning is een gewogen som van vijf attributen. De training werd uitgevoerd gedurende 2 epochs op HelpSteer2-data (batch size 128). Op RewardBench behaalde het model 92,0 %, waarmee het GPT‑4o (84,7 %), Gemini 1.5 Pro (88,1 %) en Claude‑3‑Opus (80,7 %) overtrof ten tijde van publicatie.<sup>[\[3\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-N4_340B-3)</sup>

| Model                    | Arena Hard | AlpacaEval 2.0 LC | MT‑Bench |
|--------------------------|------------|-------------------|----------|
| Nemotron‑4‑340B‑Instruct | 54,2       | 41,5 %            | 8,22     |
| Llama‑3‑70B‑Instruct     | 41,1       | 34,4 %            | 8,16     |
| Mixtral‑8x22B‑Instruct   | 36,4       | 30,9 %            | 7,63     |
| Qwen‑2‑72B‑Instruct      | 48,1       | 38,8 %            | 8,26     |

Bron: arXiv:2406.11704, tabel 5.<sup>[\[3\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-N4_340B-3)</sup>

### Llama‑3.1‑Nemotron‑70B

**Llama‑3.1‑Nemotron‑70B‑Reward** is getraind met een hybride methode die Bradley‑Terry (paarsgewijze voorkeuren) en SteerLM-regressie (meerkenmerksbeoordeling op Likert-schaal) combineert. De training werd uitsluitend uitgevoerd op HelpSteer2-data (CC‑BY‑4.0). Op RewardBench behaalde het model 94,1 %, waarmee het de eerste plaats innam ten tijde van publicatie.<sup>[\[10\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-LN70B_Reward-10)</sup>

**Llama‑3.1‑Nemotron‑70B‑Instruct** is uitgelijnd met RLHF met het REINFORCE-algoritme (niet PPO) en het Nemotron‑70B‑Reward-beloningsmodel. De infrastructuur — NeMo‑Aligner met TRT‑LLM-versnelling.<sup>[\[9\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-LN70B_Instruct-9)</sup>

### Llama‑Nemotron: Nano, Super, Ultra (2025)

Een familie van redenerende modellen, afgeleid van Llama 3.x via NAS, distillatie en uitgebreid post-training.<sup>[\[11\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-Llama_Nemotron-11)</sup>

| Model                     | Parameters  | Basismodel                     | Context |
|---------------------------|-------------|--------------------------------|---------|
| Llama‑Nemotron‑Nano 8B    | 8 miljard   | Llama‑3.1‑8B‑Instruct          | 128K    |
| Llama‑Nemotron‑Nano 4B    | 4 miljard   | Minitron 4B (uit Llama 3.1 8B) | 128K    |
| Llama‑Nemotron‑Super 49B  | 49 miljard  | Llama‑3.3‑70B → NAS (Puzzle)   | 128K    |
| Llama‑Nemotron‑Ultra 253B | 253 miljard | Llama‑3.1‑405B → NAS (Puzzle)  | 128K    |

Vijffasige trainingspipeline: (1) NAS + FFN Fusion, (2) distillatie + voortgezette voortraining, (3) SFT (inclusief redeneerreeksen van DeepSeek‑R1), (4) grootschalig RL (GRPO voor Ultra, REINFORCE/RLOO + Online RPO voor Nano), (5) uitlijning voor dialoog.<sup>[\[11\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-Llama_Nemotron-11)</sup>

De modellen zijn als eerste onder open LLM's uitgerust met **schakelbare redeneerstand** (reasoning toggle): wanneer ingeschakeld («detailed thinking on» in de systeemprompt) genereert het model een redeneerreeks in de tags `<think>...</think>` vóór het antwoord; wanneer uitgeschakeld — antwoordt het direct.<sup>[\[11\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-Llama_Nemotron-11)</sup>

Resultaten van Llama‑Nemotron‑Ultra 253B (reasoning ON): GPQA Diamond — 76,0 %, AIME 2024 — 80,8 %, MATH 500 — ~97 %. Ter vergelijking: DeepSeek‑R1 (671B totaal / 37B actief) — GPQA Diamond 71,5 %, AIME 2024 ~79,8 %. Ultra past op één knooppunt met 8×H100.<sup>[\[11\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-Llama_Nemotron-11)</sup>

### Nemotron‑H (april 2025)

Een familie van dense hybride modellen, van grond af getraind en ontworpen voor taken met lange generaties. Nemotron‑H‑56B is getraind op 20 biljoen tokens in FP8 — een van de grootste publieke experimenten met FP8-voortraining. Nemotron‑H‑47B is verkregen door compressie van 56B via MiniPuzzle (63 miljard distillatietokens) en past in het geheugen van één RTX 5090 (FP4) bij een context van ~1M tokens.<sup>[\[12\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-Nemotron_H-12)</sup>

| Benchmark             | Nemotron‑H‑56B | Nemotron‑H‑47B | Qwen‑2.5‑72B | Llama‑3.1‑70B |
|-----------------------|----------------|----------------|--------------|---------------|
| MMLU‑Pro (5‑shot CoT) | 60,5           | 61,8           | 58,8         | 51,3          |
| MMLU (5‑shot)         | 84,2           | 83,6           | 86,1         | 78,8          |
| GSM8K (8‑shot CoT)    | 93,7           | 93,3           | 90,9         | 83,9          |
| HumanEval (0‑shot)    | 60,4           | 61,0           | 56,7         | 57,3          |

Bron: arXiv:2504.03624.<sup>[\[12\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-Nemotron_H-12)</sup>

Bij generatie met 65.536 invoer- en 1024 uitvoertokens op H100 werkte het model Nemotron‑H‑47B 2,9 keer sneller dan Qwen‑2.5‑72B en Llama‑3.1‑70B.<sup>[\[12\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-Nemotron_H-12)</sup>

### Nemotron Nano 2 (augustus 2025)

Een hybride Mamba‑Transformer-model voor redeneren. **Nemotron‑Nano‑12B‑v2‑Base** — het oudermodel met 62 lagen (voornamelijk Mamba‑2 + MLP + 4 aandachtslagen), van grond af getraind op 20 biljoen tokens in FP8. Hieruit is via uitgebreide Minitron **Nemotron‑Nano‑9B‑v2** afgeleid, dat in staat is te werken in een context van 128K tokens op één GPU NVIDIA A10G (22 GB, BF16). Post-training: SFT (80 miljard tokens, inclusief DeepSeek‑R1‑0528-reeksen) → GRPO → DPO → RLHF. Op redeneer-benchmarks is het model vergelijkbaar met Qwen3‑8B qua nauwkeurigheid, maar bereikt het een 3–6-voudige versnelling van generatie bij lange uitvoerreeksen.<sup>[\[13\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-Nano2-13)</sup>

### Nemotron 3 Nano 30B‑A3B

Uitgebracht in december 2025. Parameters: 31,6 miljard totaal, 3,2 miljard actief. Architectuur: 52 lagen, hidden dimension 2688, expert dimension 1856, Mamba state dimension 128. MoE: 128 gerouteerde experts + 1 gedeelde, 6 actief per token. Context — tot 1.048.576 tokens. Training op 25 biljoen tokens (WSD-schema, batch size 3072, datumgrens gegevens — juni 2025) in twee fasen: Fase 1 — 23,5 biljoen (gevarieerde data), Fase 2 — 1,5 biljoen (hoogwaardige data) + 121 miljard tokens long‑context CPT.<sup>[\[2\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-N3_nano_report-2)</sup>

| Benchmark             | Nemotron 3 Nano 30B‑A3B | Qwen3‑30B‑A3B‑Thinking | GPT‑OSS‑20B |
|-----------------------|-------------------------|------------------------|-------------|
| MMLU‑Pro              | 78,30                   | 80,9                   | 75,0        |
| AIME25 (no tools)     | 89,06                   | 85,0                   | 91,7        |
| AIME25 (with tools)   | 99,17                   | —                      | 98,7        |
| GPQA (no tools)       | 73,04                   | 73,4                   | 71,5        |
| LiveCodeBench v6      | 68,25                   | 66,0                   | 61,0        |
| SWE‑Bench (OpenHands) | 38,76                   | 22,0\*                 | 34,0        |
| TauBench V2 Avg       | 49,04                   | 47,7                   | 47,5        |
| Arena‑Hard‑V2 Avg     | 67,65                   | 57,8                   | 48,55       |
| RULER‑100 @ 1M        | 86,34                   | 77,5                   | —           |

\* — waarden uit secundaire bronnen; reproductiescondities — NeMo Evaluator SDK. Bron: arXiv:2512.20848.<sup>[\[2\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-N3_nano_report-2)[\[28\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-NeMo_eval-28)</sup>

Doorvoer (8K invoer / 16K uitvoer, single H200): 3,3 keer hoger dan Qwen3‑30B‑A3B‑Thinking en 2,2 keer hoger dan GPT‑OSS‑20B.<sup>[\[2\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-N3_nano_report-2)</sup>

### Aanvullende modellen en richtingen

- **OpenReasoning‑Nemotron** (juli 2025) — een serie modellen van 1,5B–32B parameters, gebaseerd op Qwen 2.5 en verfijnd op DeepSeek‑R1‑0528-data. De 32B-variant met GenSelect@64 overtreft o3 (high) op een aantal wiskundige benchmarks.<sup>[\[29\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-OpenReasoning-29)</sup>
- **Nemotron Elastic** (arXiv:2511.16664) — een framework van geneste submodellen (6B, 9B, 12B) binnen één oudermodel, waarmee de trainingskosten 360 keer worden verlaagd ten opzichte van training vanaf nul.<sup>[\[30\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-Elastic-30)</sup>
- **Nemotron‑CrossThink** — een framework voor multi-domein versterkend leren buiten wiskundige taken, dat +30,1 % op MATH‑500 en +12,8 % op MMLU‑PRO opleverde.<sup>[\[31\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-CrossThink-31)</sup>
- **Nemotron‑UltraLong** — contextuitbreiding tot 4 miljoen tokens.<sup>[\[32\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-UltraLong-32)</sup>
- **Jet‑Nemotron** — post-training NAS voor compacte hybride modellen van 2B/4B.<sup>[\[33\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-JetNemotron-33)</sup>

### Gespecialiseerde modellen

In het Nemotron-ecosysteem zijn ook modellen aanwezig voor gespecialiseerde taken:

- **Nemotron Nano V2 VL** — een vision‑language-model voor OCR, tabelverwerking en video.<sup>[\[34\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-NanoV2_VL-34)</sup>
- **Nemotron Parse 1.1** — een model voor OCR en extractie van documentstructuur.<sup>[\[35\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-Parse-35)</sup>
- **Nemotron ColEmbed V2** — een embedding-model voor visueel zoeken in documenten (late interaction).<sup>[\[36\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-ColEmbed-36)</sup>
- **Nemotron Speech** — modellen voor automatische spraakherkenning (ASR), spraaksynthese (TTS) en machinevertaling (NMT).<sup>[\[1\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-N3_whitepaper-1)</sup>
- **Llama 3.1 Nemotron Safety Guard 8B V3** — een model voor classificatie van onveilige inhoud in 23 categorieën in 9 talen met een nauwkeurigheid van 84,2 %.<sup>[\[37\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-Safety-37)</sup>

## Voortrainingsdata

### Samenstelling van de Nemotron‑4-data

Het voortrainingscorpus van Nemotron‑4 15B en 340B bestaat uit drie hoofdcategorieën: Engelstalige teksten (70 %), meertalige teksten in 53 talen (15 %) en broncode in 43 programmeertalen (15 %). Er werd deduplicatie toegepast op documentniveau (exact en near‑duplicate), evenals filtering via taalmodellen en heuristieken. Het 15B-model is getraind op 8 biljoen tokens, het 340B-model op 9 biljoen (8 biljoen voortraining + 1 biljoen voortgezette training met herwogen hoogwaardige bronnen).<sup>[\[8\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-N4_15B-8)[\[3\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-N4_340B-3)</sup>

### Nemotron‑CC

De dataset **Nemotron‑CC** is samengesteld uit Common Crawl met behulp van een ensemble van kwaliteitsclassificatoren en synthetisch herformuleren van teksten. De totale omvang bedraagt **6,3 biljoen tokens** (4,4 biljoen gededupliceerd + 1,9 biljoen synthetische herformuleringen). De pipeline is geïntegreerd in het NeMo Curator-hulpprogramma. Het werk is geaccepteerd als Long Paper op de ACL 2025-conferentie.<sup>[\[38\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-NemotronCC-38)</sup>

### Nemotron‑CC‑Math

Een wiskundig georiënteerde deelverzameling van **133 miljard tokens**, geëxtraheerd uit Common Crawl met behulp van de Lynx + LLM-pipeline, met behoud van de structuur van wiskundige formules en code in LaTeX.<sup>[\[39\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-NemotronCC_Math-39)</sup>

### Data voor Nemotron 3 Nano

De voortraining van Nemotron 3 Nano is uitgevoerd op 25 biljoen tokens. De verzameling omvat: Nemotron‑CC‑v2.1, Nemotron‑CC‑Code‑v1, Nemotron‑CC‑Math‑v1 en gespecialiseerde STEM-datasets. Voor long‑context training werden aanvullend 121 miljard tokens CPT gebruikt.<sup>[\[2\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-N3_nano_report-2)</sup>

### Nemotron Pretraining Dataset v1

Een synthetisch gegenereerde dataset die STEM, academische teksten, redeneertaken en meertalige domeinen omvat; bevat meer dan 10 biljoen taaltokens en 18 miljoen SFT-voorbeelden. Alle datasets worden verspreid onder open, permissieve licenties.<sup>[\[40\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-NemotronNano2_page-40)</sup>

### Pipeline voor synthetische datageneratie (SDG)

De pipeline beschreven in het rapport over Nemotron‑4 340B omvat de volgende stappen<sup>[\[3\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-N4_340B-3)</sup>:

- **Promptgeneratie**: synthetische één- en tweebeurtsvragen, gegenereerd met Mixtral‑8x7B‑Instruct‑v0.1 (Apache 2.0-licentie) op basis van sjablonen voor Open QA, Writing, Closed QA, Math & Coding.
- **Antwoordgeneratie**: meerdere antwoorden van tussentijdse modelversies voor diversiteit.
- **Kwaliteitsevaluatie**: drie benaderingen — Ground‑Truth‑as‑a‑Judge (voor taken met verifieerbare antwoorden), LLM‑as‑Judge (paarsgewijze antwoordvergelijking door een taalmodel), Reward‑Model‑as‑Judge (gebruik van Nemotron‑4‑340B‑Reward).
- **Iteratieve uitlijning** van zwakke naar sterke modellen (weak‑to‑strong).

Uit ~20.000 door mensen geannoteerde voorbeelden (10.000 voor SFT, 10.000 HelpSteer2 voor het beloningsmodel) werd al het overige uitlijndata synthetisch gegenereerd.<sup>[\[3\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-N4_340B-3)</sup>

## Kwantisering en inferentie-optimalisatie

De Nemotron 3 Nano-modellen ondersteunen Post‑Training Quantization (PTQ) in FP8 met behulp van ModelOpt en Megatron‑LM. Selectieve kwantisering wordt toegepast — aandachtslagen en een deel van Mamba worden in BF16 gehouden voor kwaliteitsbehoud; de overige worden gekwantiseerd naar FP8. Volgens het technisch rapport levert dit ~99 % nauwkeurigheidsretentie op.<sup>[\[2\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-N3_nano_report-2)</sup> Nano ondersteunt ook inferentie in NVFP4-formaat.<sup>[\[1\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-N3_whitepaper-1)</sup>

## Overzichtstabel van benchmarks per generatie

| Model                           | MMLU   | GSM8K  | HumanEval | Arena Hard | MT‑Bench | Condities                 |
|---------------------------------|--------|--------|-----------|------------|----------|---------------------------|
| Nemotron‑4 15B                  | 64,2 % | 46,0 % | 31,6 %    | —          | —        | base; 5‑/8‑/0‑shot        |
| Nemotron‑4 340B Base            | 81,1 % | —      | 57,3 %    | —          | —        | 5‑/—/0‑shot               |
| Nemotron‑4 340B Instruct        | 78,7 % | 92,3 % | 73,2 %    | 54,2       | 8,22     | 0‑shot                    |
| Llama‑3.1‑Nemotron‑70B‑Instruct | —      | —      | —         | 85,0       | 8,98     | okt. 2024                 |
| LN‑Ultra 253B (reasoning ON)    | —      | —      | —         | —          | —        | GPQA 76,0 %, AIME 80,8 %  |
| Nemotron‑H‑47B                  | 83,6 % | 93,3 % | 61,0 %    | —          | —        | 5‑/8‑CoT/0‑shot           |
| Nemotron 3 Nano (30B‑A3B)       | —      | —      | —         | 67,7 (v2)  | —        | AIME25 89,1 %, LCB 68,3 % |

De vergelijking dient met voorzichtigheid te worden geïnterpreteerd: de evaluatiecondities, benchmarkversies en testdata verschillen per generatie. Resultaten zijn afkomstig uit technische rapporten van NVIDIA; onafhankelijke evaluaties kunnen afwijken (zie paragraaf «Beperkingen»).<sup>[\[8\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-N4_15B-8)[\[3\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-N4_340B-3)[\[9\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-LN70B_Instruct-9)[\[11\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-Llama_Nemotron-11)[\[12\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-Nemotron_H-12)[\[2\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-N3_nano_report-2)</sup>

## Toepassingen

### Implementatie via NVIDIA NIM

NVIDIA NIM — een set geoptimaliseerde container-microservices voor inferentie met een OpenAI-compatibele API. Nemotron-modellen zijn beschikbaar als NIM-microservices op het platform build.nvidia.com. NIM ondersteunt FP8-, BF16- en NVFP4-formaten en implementatie via Helm-charts op Kubernetes. De modellen zijn ook compatibel met onafhankelijke frameworks: vLLM, SGLang, Ollama, llama.cpp, TensorRT‑LLM.<sup>[\[4\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-NV_developer-4)[\[1\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-N3_whitepaper-1)</sup>

### Synthetische datageneratie

Nemotron‑4 340B is ontworpen voor gebruik als generator van synthetische data: het Instruct-model genereert antwoorden, het Reward-model beoordeelt en filtert ze. De NVIDIA Open Model License staat uitdrukkelijk toe dat modeloutput wordt gebruikt voor het trainen van andere modellen. Volgens ServiceNow is 15 % van de voortrainingsdata van hun model Apriel 1.6 afkomstig uit Nemotron-datasets.<sup>[\[3\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-N4_340B-3)[\[15\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-NVIDIA_news-15)[\[41\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-NeMo_synth-41)</sup>

### Agentische systemen

De modellen van de Nemotron 3-familie (Nano, Super, Ultra) zijn bestemd voor multi-agent-werklast: planning, retrieval, gereedschapsaanroep (tool use). Nemotron 3 Nano behaalt 38,76 % op SWE‑Bench (met OpenHands) en 49,04 % (gemiddeld) op TauBench V2.<sup>[\[2\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-N3_nano_report-2)</sup>

### Bedrijfsimplementaties

Onder de genoemde partners: ServiceNow (gezamenlijk model Apriel Nemotron 15B voor workflowautomatisering), CrowdStrike (Charlotte AI-platform), Perplexity (verzoekroutering), Accenture, Cadence, Deloitte, Oracle, Palantir, Siemens, Synopsys, Zoom. Modellen zijn beschikbaar op AWS Amazon Bedrock; ondersteuning voor Google Cloud, CoreWeave en Nebius is gepland.<sup>[\[15\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-NVIDIA_news-15)</sup>

## Beperkingen en open problemen

### Onafhankelijke evaluaties en afwijkingen van NVIDIA-gegevens

Onafhankelijke evaluaties onthullen afwijkingen ten opzichte van de door NVIDIA vermelde resultaten. Volgens Artificial Analysis (februari 2026) kreeg Llama‑Nemotron‑Ultra 253B (reasoning) een Intelligence Index van 15, bij een mediaan van 26 voor modellen van vergelijkbare grootte. Op het Chatbot Arena-platform (LMArena) bezetten Nemotron-modellen middenposities in de ranglijst: Llama 3.3 Nemotron Super 49B — Elo 1327 ±12 (~147e positie), Nemotron 3 Nano — 1317 ±6 (~164e positie).<sup>[\[42\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-AA-42)[\[43\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-LMArena-43)</sup>

### Uitvoerigheid

Nemotron-modellen vertonen verhoogde uitvoerigheid: bij de Artificial Analysis-evaluatie genereerde het Nemotron 3 Nano-model 140 miljoen tokens (bij een mediaan van 12 miljoen voor andere modellen), wat de inferentiekosten verhoogt. Een analyse van DEV Community toonde aan dat Llama‑3.1‑Nemotron‑70B‑Instruct, ondanks formeel leiderschap op automatische benchmarks, onvoldoende nauwkeurigheid vertoonde bij een aantal praktische taken, en dat hoge scores op Arena-achtige benchmarks mogelijk verklaard werden door een voorkeur voor uitvoerige en zelfverzekerde, maar niet noodzakelijk nauwkeurige antwoorden.<sup>[\[42\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-AA-42)[\[44\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-Saplin-44)</sup>

### Hallucinaties en bias

NVIDIA vermeldt in de modelkaarten dat de modellen «vooroordelen kunnen versterken en toxische antwoorden kunnen genereren, met name bij toxische prompts», en ook «onnauwkeurige informatie kunnen produceren en cruciale details kunnen weglaten». De openheid van gewichten en data maakt externe audits mogelijk.<sup>[\[11\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-Llama_Nemotron-11)[\[13\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-Nano2-13)</sup>

### Rekenvereisten

Dense modellen (Nemotron‑4 340B) vereisen een cluster van 768 DGX H100-knooppunten voor volledige training, wat de reproduceerbaarheid beperkt voor academische groepen zonder toegang tot vergelijkbare infrastructuur. Lange context (1M) bij inferentie vereist aanzienlijke VRAM.<sup>[\[3\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-N4_340B-3)[\[2\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-N3_nano_report-2)</sup>

### Degradatie bij contextuitbreiding

Bij uitbreiding van de context naar zeer lange reeksen werd een verslechtering van de resultaten op benchmarks met korte context waargenomen.<sup>[\[32\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-UltraLong-32)</sup>

### Kwantisering van hybride architecturen

Het gebruik van extreem lage bitbreedte (FP8/NVFP4) in Mamba‑Transformer-architecturen leidt tot een geringe daling van de nauwkeurigheid (~1 % op een aantal benchmarks). Dit probleem wordt opgelost door selectieve kwantisering toe te passen, wat de architectuur van compilers en inferentie-servers compliceert.<sup>[\[2\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-N3_nano_report-2)[\[1\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-N3_whitepaper-1)</sup>

### Overige open problemen

- Afhankelijkheid van benchmarks met mogelijke contaminatie van trainingsdata.
- Gebrek aan gestandaardiseerde vergelijkingsprotocollen voor hybride SSM‑Transformer-architecturen ten opzichte van pure Transformer-modellen.
- Vraagstuk van de schaalbaarheid van de MoE-aanpak bij taken die diepgaand redeneren vereisen met een klein aantal actieve experts per token.
- Gegevens over Super/Ultra waren per december 2025 voorlopig; volledige benchmarks worden verwacht na de release.<sup>[\[1\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-N3_whitepaper-1)</sup>

## Ethische en regelgevende aspecten

### Veiligheid tijdens de ontwikkelingsfase

Tijdens de ontwikkeling worden toegepast: evaluatie via het AEGIS-framework (13 categorieën voor inhoudsveiligheid), kwetsbaarheidsscanner garak, handmatig «rood testen» (red‑teaming).<sup>[\[37\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-Safety-37)</sup>

### Veiligheid tijdens inferentie

**NeMo Guardrails** — een open toolset (Apache 2.0-licentie), die programmeerbare barrières toevoegt aan LLM-systemen. De microservice onderschept invoer en uitvoer en past configureerbare controles toe: onderwerpbewaking, PII-detectie, verificatie van RAG-«aarding», detectie van jailbreak-aanvallen (inclusief bescherming tegen code-injectie op basis van YARA), meertalige multimodale veiligheidsclassificatie.<sup>[\[45\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-Guardrails-45)</sup>

Voor Nemotron 3 is een set van ~11.000 gelabelde OpenTelemetry-reeksen gepubliceerd uit realistische scenario's met toolgebruik, bedoeld voor evaluatie van agentische veiligheid.<sup>[\[1\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-N3_whitepaper-1)</sup>

### Licenties

| Modellen                               | Licentie                                       |
|----------------------------------------|------------------------------------------------|
| Nemotron 3 (Nano, Super, Ultra)        | NVIDIA Nemotron Open Model License             |
| Nemotron‑4 340B (Base/Instruct/Reward) | NVIDIA Open Model License Agreement            |
| Llama‑Nemotron (Nano/Super/Ultra)      | Llama Community License (overgenomen van Meta) |
| Nemotron‑3 8B (2023)                   | NVIDIA AI Foundation Models Community License  |

De **NVIDIA Nemotron Open Model License** staat commercieel gebruik toe, het maken en verspreiden van afgeleide werken, vereist geen naamsvermelding (bij behoud van het NOTICE-bestand), stelt geen beperkingen aan het aantal gebruikers en staat uitdrukkelijk toe dat modeloutput wordt gebruikt voor het trainen van andere modellen.<sup>[\[46\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-License-46)</sup> Op Llama gebaseerde modellen erven de beperkingen van Meta, waaronder de drempel van 700 miljoen maandelijks actieve gebruikers.<sup>[\[11\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-Llama_Nemotron-11)</sup>

Voor Nemotron 3 Nano heeft NVIDIA gepubliceerd: modelgewichten, meer dan 10 biljoen tokens aan trainingsdata, trainingsrecepten, codebases (NeMo, Megatron‑LM, NeMo‑Aligner), meer dan 10 versterkend-leerromgevingen met 900.000+ taken.<sup>[\[1\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-N3_whitepaper-1)[\[2\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-N3_nano_report-2)</sup>

## Vooruitzichten en onderzoeksrichtingen

De komende releases omvatten **Nemotron 3 Super** (~100 miljard parameters, ~10 miljard actief) en **Nemotron 3 Ultra** (~500 miljard, ~50 miljard actief), verwacht in de eerste helft van 2026. Beide modellen zullen gebruik maken van LatentMoE, MTP en training in NVFP4-formaat op de Blackwell-architectuur.<sup>[\[1\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-N3_whitepaper-1)</sup>

NVIDIA's strategische richting is het positioneren van Nemotron niet als een enkel model, maar als een infrastructurele laag voor multi-agentsystemen, waarbij verschillende modellen van de familie worden gebruikt voor verschillende taken met routering op basis van complexiteit.<sup>[\[1\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-N3_whitepaper-1)[\[15\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-NVIDIA_news-15)</sup>

Belangrijkste onderzoeksrichtingen:

- Ontwikkeling van hybride architecturen — verdere integratie van SSM-lagen met het aandachtsmechanisme voor schaalbare lange context.<sup>[\[12\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-Nemotron_H-12)</sup>
- Meerlagige compressiemethoden — verdere ontwikkeling van Minitron, Puzzle, MiniPuzzle en Nemotron Elastic.<sup>[\[20\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-Minitron-20)[\[21\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-Puzzle-21)[\[30\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-Elastic-30)</sup>
- Multi-domein versterkend leren — Nemotron‑CrossThink voor uitbreiding van RL buiten wiskundige taken.<sup>[\[31\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-CrossThink-31)</sup>
- Contextuitbreiding — Nemotron‑UltraLong tot 4 miljoen tokens.<sup>[\[32\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-UltraLong-32)</sup>
- Multimodaliteit — uitbreiding naar vision‑language (Nemotron VL), spraak (Nemotron Speech), visueel zoeken in documenten (Nemotron ColEmbed V2).<sup>[\[34\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-NanoV2_VL-34)[\[35\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-Parse-35)[\[36\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-ColEmbed-36)</sup>
- Compacte hybride modellen — Jet‑Nemotron (NAS voor modellen van 2B/4B).<sup>[\[33\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-JetNemotron-33)</sup>
- Open recepten — publicatie van volledige trainingsrecepten en data voor reproductie en aanpassing door de gemeenschap.<sup>[\[14\]](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_note-N3_blog-14)</sup>

## Zie ook

- Transformer-architectuur
- Reinforcement Learning from Human Feedback (RLHF)
- Llama (modellfamilie van Meta)

## Verwijzingen

- NVIDIA Research — Nemotron 3-pagina: <a href="https://research.nvidia.com/labs/nemotron/Nemotron-3/" class="external free" rel="nofollow">https://research.nvidia.com/labs/nemotron/Nemotron-3/</a>
- NVIDIA Developer — Nemotron AI Models: <a href="https://developer.nvidia.com/nemotron" class="external free" rel="nofollow">https://developer.nvidia.com/nemotron</a>
- Hugging Face — Nemotron-documentatie: <a href="https://huggingface.co/docs/transformers/main/en/model_doc/nemotron" class="external free" rel="nofollow">https://huggingface.co/docs/transformers/main/en/model_doc/nemotron</a>
- NeMo Framework Documentation: <a href="https://docs.nvidia.com/nemo/" class="external free" rel="nofollow">https://docs.nvidia.com/nemo/</a>

## Literatuur

- NVIDIA. *Nemotron‑3‑8B‑Base‑4k — Model Card*. Hugging Face, 2023. <a href="https://huggingface.co/nvidia/nemotron-3-8b-base-4k" class="external free" rel="nofollow">https://huggingface.co/nvidia/nemotron-3-8b-base-4k</a>
- Parmar, J. et al. (NVIDIA). *Nemotron‑4 15B Technical Report*. arXiv:2402.16819, februari 2024. <a href="https://arxiv.org/abs/2402.16819" class="external free" rel="nofollow">https://arxiv.org/abs/2402.16819</a>
- Adler, B. et al. (NVIDIA). *Nemotron‑4 340B Technical Report*. arXiv:2406.11704, juni 2024. <a href="https://arxiv.org/abs/2406.11704" class="external free" rel="nofollow">https://arxiv.org/abs/2406.11704</a>
- Wang, Z. et al. (NVIDIA). *HelpSteer2: Open‑source Dataset for Training Top‑Performing Reward Models*. arXiv:2406.08673, juni 2024. <a href="https://arxiv.org/abs/2406.08673" class="external free" rel="nofollow">https://arxiv.org/abs/2406.08673</a>
- Muralidharan, S. et al. (NVIDIA). *Compact Language Models via Pruning and Knowledge Distillation*. arXiv:2407.14679, juli 2024. <a href="https://arxiv.org/abs/2407.14679" class="external free" rel="nofollow">https://arxiv.org/abs/2407.14679</a>
- Bercovich, A. et al. (NVIDIA). *Puzzle: Distillation‑Based NAS for Inference‑Optimized LLMs*. arXiv:2411.19146, november 2024. <a href="https://arxiv.org/abs/2411.19146" class="external free" rel="nofollow">https://arxiv.org/abs/2411.19146</a>
- Su, D. et al. (NVIDIA). *Nemotron‑CC: Transforming Common Crawl into a Refined Long‑Horizon Pretraining Dataset*. ACL 2025. arXiv:2412.02595. <a href="https://arxiv.org/abs/2412.02595" class="external free" rel="nofollow">https://arxiv.org/abs/2412.02595</a>
- Sun, S. et al. (NVIDIA). *Reward‑aware Preference Optimization*. arXiv:2502.00203, januari 2025. <a href="https://arxiv.org/abs/2502.00203" class="external free" rel="nofollow">https://arxiv.org/abs/2502.00203</a>
- Blakeman, A. et al. (NVIDIA). *Nemotron‑H: A Family of Accurate and Efficient Hybrid Mamba‑Transformer Models*. arXiv:2504.03624, april 2025. <a href="https://arxiv.org/abs/2504.03624" class="external free" rel="nofollow">https://arxiv.org/abs/2504.03624</a>
- Bercovich, A. et al. (NVIDIA). *Llama‑Nemotron: Efficient Reasoning Models*. arXiv:2505.00949, mei 2025. <a href="https://arxiv.org/abs/2505.00949" class="external free" rel="nofollow">https://arxiv.org/abs/2505.00949</a>
- Basant, A. et al. (NVIDIA). *NVIDIA Nemotron Nano 2*. arXiv:2508.14444, augustus 2025. <a href="https://arxiv.org/abs/2508.14444" class="external free" rel="nofollow">https://arxiv.org/abs/2508.14444</a>
- Karimi Mahabadi, R. et al. (NVIDIA). *Nemotron‑CC‑Math*. arXiv:2508.15096, augustus 2025. <a href="https://arxiv.org/abs/2508.15096" class="external free" rel="nofollow">https://arxiv.org/abs/2508.15096</a>
- Taghibakhshi, A. et al. (NVIDIA). *Nemotron Elastic*. arXiv:2511.16664, november 2025. <a href="https://arxiv.org/abs/2511.16664" class="external free" rel="nofollow">https://arxiv.org/abs/2511.16664</a>
- Blakeman, A. et al. (NVIDIA). *NVIDIA Nemotron 3: Efficient and Open Intelligence*. arXiv:2512.20856, december 2025. <a href="https://arxiv.org/abs/2512.20856" class="external free" rel="nofollow">https://arxiv.org/abs/2512.20856</a>
- Blakeman, A. et al. (NVIDIA). *Nemotron 3 Nano*. arXiv:2512.20848, december 2025. <a href="https://arxiv.org/abs/2512.20848" class="external free" rel="nofollow">https://arxiv.org/abs/2512.20848</a>
- Dao, T.; Gu, A. *Transformers are SSMs*. ICML 2024. arXiv:2405.21060. <a href="https://arxiv.org/abs/2405.21060" class="external free" rel="nofollow">https://arxiv.org/abs/2405.21060</a>
- Ainslie, J. et al. *GQA*. arXiv:2305.13245, 2023. <a href="https://arxiv.org/abs/2305.13245" class="external free" rel="nofollow">https://arxiv.org/abs/2305.13245</a>
- Su, J. et al. *RoFormer (RoPE)*. Neurocomputing, 2024. arXiv:2104.09864. <a href="https://arxiv.org/abs/2104.09864" class="external free" rel="nofollow">https://arxiv.org/abs/2104.09864</a>
- Zhang, B.; Sennrich, R. *RMSNorm*. arXiv:1910.07467, 2019. <a href="https://arxiv.org/abs/1910.07467" class="external free" rel="nofollow">https://arxiv.org/abs/1910.07467</a>
- Rafailov, R. et al. *Direct Preference Optimization*. NeurIPS 2023. arXiv:2305.18290. <a href="https://arxiv.org/abs/2305.18290" class="external free" rel="nofollow">https://arxiv.org/abs/2305.18290</a>

## Noten

1.  <span id="cite_note-N3_whitepaper-1">↑ <sup>[1.00](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-N3_whitepaper_1-0)</sup> <sup>[1.01](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-N3_whitepaper_1-1)</sup> <sup>[1.02](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-N3_whitepaper_1-2)</sup> <sup>[1.03](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-N3_whitepaper_1-3)</sup> <sup>[1.04](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-N3_whitepaper_1-4)</sup> <sup>[1.05](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-N3_whitepaper_1-5)</sup> <sup>[1.06](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-N3_whitepaper_1-6)</sup> <sup>[1.07](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-N3_whitepaper_1-7)</sup> <sup>[1.08](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-N3_whitepaper_1-8)</sup> <sup>[1.09](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-N3_whitepaper_1-9)</sup> <sup>[1.10](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-N3_whitepaper_1-10)</sup> <sup>[1.11](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-N3_whitepaper_1-11)</sup> <sup>[1.12](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-N3_whitepaper_1-12)</sup> <sup>[1.13](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-N3_whitepaper_1-13)</sup> <sup>[1.14](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-N3_whitepaper_1-14)</sup> Blakeman, A. et al. (NVIDIA). *NVIDIA Nemotron 3: Efficient and Open Intelligence*. arXiv:2512.20856, декабрь 2025. <a href="https://arxiv.org/abs/2512.20856" class="external free" rel="nofollow">https://arxiv.org/abs/2512.20856</a></span>
2.  <span id="cite_note-N3_nano_report-2">↑ <sup>[2.00](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-N3_nano_report_2-0)</sup> <sup>[2.01](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-N3_nano_report_2-1)</sup> <sup>[2.02](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-N3_nano_report_2-2)</sup> <sup>[2.03](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-N3_nano_report_2-3)</sup> <sup>[2.04](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-N3_nano_report_2-4)</sup> <sup>[2.05](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-N3_nano_report_2-5)</sup> <sup>[2.06](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-N3_nano_report_2-6)</sup> <sup>[2.07](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-N3_nano_report_2-7)</sup> <sup>[2.08](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-N3_nano_report_2-8)</sup> <sup>[2.09](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-N3_nano_report_2-9)</sup> <sup>[2.10](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-N3_nano_report_2-10)</sup> <sup>[2.11](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-N3_nano_report_2-11)</sup> <sup>[2.12](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-N3_nano_report_2-12)</sup> <sup>[2.13](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-N3_nano_report_2-13)</sup> <sup>[2.14](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-N3_nano_report_2-14)</sup> <sup>[2.15](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-N3_nano_report_2-15)</sup> <sup>[2.16](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-N3_nano_report_2-16)</sup> Blakeman, A. et al. (NVIDIA). *Nemotron 3 Nano: Open, Efficient Mixture‑of‑Experts Hybrid Mamba‑Transformer Model for Agentic Reasoning*. arXiv:2512.20848, декабрь 2025. <a href="https://arxiv.org/abs/2512.20848" class="external free" rel="nofollow">https://arxiv.org/abs/2512.20848</a></span>
3.  <span id="cite_note-N4_340B-3">↑ <sup>[3.00](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-N4_340B_3-0)</sup> <sup>[3.01](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-N4_340B_3-1)</sup> <sup>[3.02](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-N4_340B_3-2)</sup> <sup>[3.03](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-N4_340B_3-3)</sup> <sup>[3.04](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-N4_340B_3-4)</sup> <sup>[3.05](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-N4_340B_3-5)</sup> <sup>[3.06](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-N4_340B_3-6)</sup> <sup>[3.07](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-N4_340B_3-7)</sup> <sup>[3.08](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-N4_340B_3-8)</sup> <sup>[3.09](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-N4_340B_3-9)</sup> <sup>[3.10](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-N4_340B_3-10)</sup> <sup>[3.11](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-N4_340B_3-11)</sup> <sup>[3.12](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-N4_340B_3-12)</sup> <sup>[3.13](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-N4_340B_3-13)</sup> <sup>[3.14](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-N4_340B_3-14)</sup> Adler, B. et al. (NVIDIA). *Nemotron‑4 340B Technical Report*. arXiv:2406.11704, июнь 2024. <a href="https://arxiv.org/abs/2406.11704" class="external free" rel="nofollow">https://arxiv.org/abs/2406.11704</a></span>
4.  <span id="cite_note-NV_developer-4">↑ <sup>[4.0](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-NV_developer_4-0)</sup> <sup>[4.1](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-NV_developer_4-1)</sup> NVIDIA Developer. *NVIDIA Nemotron AI Models*. <a href="https://developer.nvidia.com/nemotron" class="external free" rel="nofollow">https://developer.nvidia.com/nemotron</a></span>
5.  <span id="cite_note-NeMo_docs-5">[↑](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-NeMo_docs_5-0) NVIDIA. *NeMo Framework Documentation*. <a href="https://docs.nvidia.com/nemo/" class="external free" rel="nofollow">https://docs.nvidia.com/nemo/</a></span>
6.  <span id="cite_note-N3_8B_card-6">[↑](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-N3_8B_card_6-0) NVIDIA. *Nemotron‑3‑8B‑Base‑4k — Model Card*. Hugging Face, 2023. <a href="https://huggingface.co/nvidia/nemotron-3-8b-base-4k" class="external free" rel="nofollow">https://huggingface.co/nvidia/nemotron-3-8b-base-4k</a></span>
7.  <span id="cite_note-MS_nemotron3-7">[↑](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-MS_nemotron3_7-0) Microsoft. *Introducing NVIDIA Nemotron‑3 8B LLMs on the Model Catalog*. Microsoft Tech Community, 14 ноября 2023. <a href="https://techcommunity.microsoft.com/blog/machinelearningblog/introducing-nvidia-nemotron-3-8b-llms-on-the-model-catalog/3983569" class="external free" rel="nofollow">https://techcommunity.microsoft.com/blog/machinelearningblog/introducing-nvidia-nemotron-3-8b-llms-on-the-model-catalog/3983569</a></span>
8.  <span id="cite_note-N4_15B-8">↑ <sup>[8.0](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-N4_15B_8-0)</sup> <sup>[8.1](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-N4_15B_8-1)</sup> <sup>[8.2](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-N4_15B_8-2)</sup> <sup>[8.3](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-N4_15B_8-3)</sup> <sup>[8.4](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-N4_15B_8-4)</sup> <sup>[8.5](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-N4_15B_8-5)</sup> <sup>[8.6](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-N4_15B_8-6)</sup> Parmar, J., Prabhumoye, S., Jennings, J. et al. (NVIDIA). *Nemotron‑4 15B Technical Report*. arXiv:2402.16819, февраль 2024. <a href="https://arxiv.org/abs/2402.16819" class="external free" rel="nofollow">https://arxiv.org/abs/2402.16819</a></span>
9.  <span id="cite_note-LN70B_Instruct-9">↑ <sup>[9.0](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-LN70B_Instruct_9-0)</sup> <sup>[9.1](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-LN70B_Instruct_9-1)</sup> <sup>[9.2](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-LN70B_Instruct_9-2)</sup> NVIDIA. *Llama‑3.1‑Nemotron‑70B‑Instruct — Model Card*. Hugging Face, октябрь 2024. <a href="https://huggingface.co/nvidia/Llama-3.1-Nemotron-70B-Instruct-HF" class="external free" rel="nofollow">https://huggingface.co/nvidia/Llama-3.1-Nemotron-70B-Instruct-HF</a></span>
10. <span id="cite_note-LN70B_Reward-10">↑ <sup>[10.0](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-LN70B_Reward_10-0)</sup> <sup>[10.1](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-LN70B_Reward_10-1)</sup> NVIDIA. *Llama‑3.1‑Nemotron‑70B‑Reward — Model Card*. Hugging Face, октябрь 2024. <a href="https://huggingface.co/nvidia/Llama-3.1-Nemotron-70B-Reward" class="external free" rel="nofollow">https://huggingface.co/nvidia/Llama-3.1-Nemotron-70B-Reward</a></span>
11. <span id="cite_note-Llama_Nemotron-11">↑ <sup>[11.0](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-Llama_Nemotron_11-0)</sup> <sup>[11.1](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-Llama_Nemotron_11-1)</sup> <sup>[11.2](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-Llama_Nemotron_11-2)</sup> <sup>[11.3](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-Llama_Nemotron_11-3)</sup> <sup>[11.4](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-Llama_Nemotron_11-4)</sup> <sup>[11.5](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-Llama_Nemotron_11-5)</sup> <sup>[11.6](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-Llama_Nemotron_11-6)</sup> <sup>[11.7](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-Llama_Nemotron_11-7)</sup> <sup>[11.8](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-Llama_Nemotron_11-8)</sup> Bercovich, A. et al. (NVIDIA). *Llama‑Nemotron: Efficient Reasoning Models*. arXiv:2505.00949, май 2025. <a href="https://arxiv.org/abs/2505.00949" class="external free" rel="nofollow">https://arxiv.org/abs/2505.00949</a></span>
12. <span id="cite_note-Nemotron_H-12">↑ <sup>[12.0](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-Nemotron_H_12-0)</sup> <sup>[12.1](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-Nemotron_H_12-1)</sup> <sup>[12.2](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-Nemotron_H_12-2)</sup> <sup>[12.3](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-Nemotron_H_12-3)</sup> <sup>[12.4](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-Nemotron_H_12-4)</sup> <sup>[12.5](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-Nemotron_H_12-5)</sup> <sup>[12.6](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-Nemotron_H_12-6)</sup> <sup>[12.7](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-Nemotron_H_12-7)</sup> <sup>[12.8](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-Nemotron_H_12-8)</sup> Blakeman, A. et al. (NVIDIA). *Nemotron‑H: A Family of Accurate and Efficient Hybrid Mamba‑Transformer Models*. arXiv:2504.03624, апрель 2025. <a href="https://arxiv.org/abs/2504.03624" class="external free" rel="nofollow">https://arxiv.org/abs/2504.03624</a></span>
13. <span id="cite_note-Nano2-13">↑ <sup>[13.0](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-Nano2_13-0)</sup> <sup>[13.1](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-Nano2_13-1)</sup> <sup>[13.2](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-Nano2_13-2)</sup> Basant, A. et al. (NVIDIA). *NVIDIA Nemotron Nano 2: An Accurate and Efficient Hybrid Mamba‑Transformer Reasoning Model*. arXiv:2508.14444, август 2025. <a href="https://arxiv.org/abs/2508.14444" class="external free" rel="nofollow">https://arxiv.org/abs/2508.14444</a></span>
14. <span id="cite_note-N3_blog-14">↑ <sup>[14.0](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-N3_blog_14-0)</sup> <sup>[14.1](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-N3_blog_14-1)</sup> <sup>[14.2](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-N3_blog_14-2)</sup> Blakeman, A. et al. (NVIDIA). *Inside NVIDIA Nemotron 3: Techniques, Tools, and Data That Make It Efficient and Accurate*. NVIDIA Technical Blog, 15 декабря 2025. <a href="https://developer.nvidia.com/blog/inside-nvidia-nemotron-3-techniques-tools-and-data-that-make-it-efficient-and-accurate/" class="external free" rel="nofollow">https://developer.nvidia.com/blog/inside-nvidia-nemotron-3-techniques-tools-and-data-that-make-it-efficient-and-accurate/</a></span>
15. <span id="cite_note-NVIDIA_news-15">↑ <sup>[15.0](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-NVIDIA_news_15-0)</sup> <sup>[15.1](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-NVIDIA_news_15-1)</sup> <sup>[15.2](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-NVIDIA_news_15-2)</sup> <sup>[15.3](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-NVIDIA_news_15-3)</sup> NVIDIA. *NVIDIA Debuts Nemotron 3 Family of Open Models*. NVIDIA Newsroom, 15 декабря 2025. <a href="https://nvidianews.nvidia.com/news/nvidia-debuts-nemotron-3-family-of-open-models" class="external free" rel="nofollow">https://nvidianews.nvidia.com/news/nvidia-debuts-nemotron-3-family-of-open-models</a></span>
16. <span id="cite_note-GQA-16">[↑](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-GQA_16-0) Ainslie, J. et al. *GQA: Training Generalized Multi‑Query Transformer Models from Multi‑Head Checkpoints*. arXiv:2305.13245, 2023. <a href="https://arxiv.org/abs/2305.13245" class="external free" rel="nofollow">https://arxiv.org/abs/2305.13245</a></span>
17. <span id="cite_note-RoPE-17">[↑](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-RoPE_17-0) Su, J. et al. *RoFormer: Enhanced Transformer with Rotary Position Embedding*. Neurocomputing, 2024. arXiv:2104.09864. <a href="https://arxiv.org/abs/2104.09864" class="external free" rel="nofollow">https://arxiv.org/abs/2104.09864</a></span>
18. <span id="cite_note-Mamba2-18">[↑](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-Mamba2_18-0) Dao, T.; Gu, A. *Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality*. ICML 2024. arXiv:2405.21060. <a href="https://arxiv.org/abs/2405.21060" class="external free" rel="nofollow">https://arxiv.org/abs/2405.21060</a></span>
19. <span id="cite_note-RMSNorm-19">[↑](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-RMSNorm_19-0) Zhang, B.; Sennrich, R. *Root Mean Square Layer Normalization*. arXiv:1910.07467, 2019. <a href="https://arxiv.org/abs/1910.07467" class="external free" rel="nofollow">https://arxiv.org/abs/1910.07467</a></span>
20. <span id="cite_note-Minitron-20">↑ <sup>[20.0](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-Minitron_20-0)</sup> <sup>[20.1](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-Minitron_20-1)</sup> Muralidharan, S. et al. (NVIDIA). *Compact Language Models via Pruning and Knowledge Distillation*. arXiv:2407.14679, июль 2024. <a href="https://arxiv.org/abs/2407.14679" class="external free" rel="nofollow">https://arxiv.org/abs/2407.14679</a></span>
21. <span id="cite_note-Puzzle-21">↑ <sup>[21.0](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-Puzzle_21-0)</sup> <sup>[21.1](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-Puzzle_21-1)</sup> Bercovich, A. et al. (NVIDIA). *Puzzle: Distillation‑Based NAS for Inference‑Optimized LLMs*. arXiv:2411.19146, ноябрь 2024. <a href="https://arxiv.org/abs/2411.19146" class="external free" rel="nofollow">https://arxiv.org/abs/2411.19146</a></span>
22. <span id="cite_note-HelpSteer-22">[↑](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-HelpSteer_22-0) Wang, Z. et al. (NVIDIA). *HelpSteer: Multi‑attribute Helpfulness Dataset for SteerLM*. arXiv:2311.09528, ноябрь 2023. <a href="https://arxiv.org/abs/2311.09528" class="external free" rel="nofollow">https://arxiv.org/abs/2311.09528</a></span>
23. <span id="cite_note-HelpSteer2-23">[↑](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-HelpSteer2_23-0) Wang, Z. et al. (NVIDIA). *HelpSteer2: Open‑source Dataset for Training Top‑Performing Reward Models*. arXiv:2406.08673, июнь 2024. <a href="https://arxiv.org/abs/2406.08673" class="external free" rel="nofollow">https://arxiv.org/abs/2406.08673</a></span>
24. <span id="cite_note-HelpSteer2_Pref-24">[↑](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-HelpSteer2_Pref_24-0) Wang, Z. et al. (NVIDIA). *HelpSteer2‑Preference: Complementing Ratings with Preferences*. arXiv:2410.01257, октябрь 2024. <a href="https://arxiv.org/abs/2410.01257" class="external free" rel="nofollow">https://arxiv.org/abs/2410.01257</a></span>
25. <span id="cite_note-SteerLM-25">[↑](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-SteerLM_25-0) Dong, Y. et al. (NVIDIA). *SteerLM: Attribute Conditioned SFT as an (User‑Steerable) Alternative to RLHF*. arXiv:2310.05344, октябрь 2023. <a href="https://arxiv.org/abs/2310.05344" class="external free" rel="nofollow">https://arxiv.org/abs/2310.05344</a></span>
26. <span id="cite_note-DPO-26">[↑](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-DPO_26-0) Rafailov, R. et al. *Direct Preference Optimization: Your Language Model is Secretly a Reward Model*. NeurIPS 2023. arXiv:2305.18290. <a href="https://arxiv.org/abs/2305.18290" class="external free" rel="nofollow">https://arxiv.org/abs/2305.18290</a></span>
27. <span id="cite_note-RPO-27">↑ <sup>[27.0](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-RPO_27-0)</sup> <sup>[27.1](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-RPO_27-1)</sup> Sun, S. et al. (NVIDIA). *Reward‑aware Preference Optimization: A Unified Mathematical Framework for Model Alignment*. arXiv:2502.00203, январь 2025. <a href="https://arxiv.org/abs/2502.00203" class="external free" rel="nofollow">https://arxiv.org/abs/2502.00203</a></span>
28. <span id="cite_note-NeMo_eval-28">[↑](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-NeMo_eval_28-0) NVIDIA. *NeMo Evaluator SDK — Evaluation Recipe for Nemotron 3 Nano*. Hugging Face Blog, декабрь 2025. <a href="https://huggingface.co/blog/nvidia/nemotron-3-nano-evaluation-recipe" class="external free" rel="nofollow">https://huggingface.co/blog/nvidia/nemotron-3-nano-evaluation-recipe</a></span>
29. <span id="cite_note-OpenReasoning-29">[↑](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-OpenReasoning_29-0) NVIDIA. *OpenReasoning‑Nemotron*. Hugging Face Collection, июль 2025. <a href="https://huggingface.co/collections/nvidia/openreasoning-nemotron-685824d42db3e24d8b8f5e39" class="external free" rel="nofollow">https://huggingface.co/collections/nvidia/openreasoning-nemotron-685824d42db3e24d8b8f5e39</a></span>
30. <span id="cite_note-Elastic-30">↑ <sup>[30.0](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-Elastic_30-0)</sup> <sup>[30.1](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-Elastic_30-1)</sup> Taghibakhshi, A. et al. (NVIDIA). *Nemotron Elastic: Towards Efficient Many‑in‑One Reasoning LLMs*. arXiv:2511.16664, ноябрь 2025. <a href="https://arxiv.org/abs/2511.16664" class="external free" rel="nofollow">https://arxiv.org/abs/2511.16664</a></span>
31. <span id="cite_note-CrossThink-31">↑ <sup>[31.0](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-CrossThink_31-0)</sup> <sup>[31.1](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-CrossThink_31-1)</sup> Akter, S. N. et al. (NVIDIA). *Nemotron‑CrossThink: Scaling Self‑Learning beyond Math Reasoning*. arXiv:2504.13941, апрель 2025. <a href="https://arxiv.org/abs/2504.13941" class="external free" rel="nofollow">https://arxiv.org/abs/2504.13941</a></span>
32. <span id="cite_note-UltraLong-32">↑ <sup>[32.0](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-UltraLong_32-0)</sup> <sup>[32.1](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-UltraLong_32-1)</sup> <sup>[32.2](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-UltraLong_32-2)</sup> Xu, C. et al. (NVIDIA). *From 128K to 4M: Efficient Training of Ultra‑Long Context Large Language Models*. arXiv:2504.06214, апрель 2025. <a href="https://arxiv.org/abs/2504.06214" class="external free" rel="nofollow">https://arxiv.org/abs/2504.06214</a></span>
33. <span id="cite_note-JetNemotron-33">↑ <sup>[33.0](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-JetNemotron_33-0)</sup> <sup>[33.1](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-JetNemotron_33-1)</sup> Gu, Y. et al. (NVIDIA). *Jet‑Nemotron: Efficient Language Model with Post Neural Architecture Search*. arXiv:2508.15884, август 2025. <a href="https://arxiv.org/abs/2508.15884" class="external free" rel="nofollow">https://arxiv.org/abs/2508.15884</a></span>
34. <span id="cite_note-NanoV2_VL-34">↑ <sup>[34.0](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-NanoV2_VL_34-0)</sup> <sup>[34.1](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-NanoV2_VL_34-1)</sup> Deshmukh, A. S. et al. (NVIDIA). *NVIDIA Nemotron Nano V2 VL*. arXiv:2511.03929, ноябрь 2025. <a href="https://arxiv.org/abs/2511.03929" class="external free" rel="nofollow">https://arxiv.org/abs/2511.03929</a></span>
35. <span id="cite_note-Parse-35">↑ <sup>[35.0](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-Parse_35-0)</sup> <sup>[35.1](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-Parse_35-1)</sup> Chumachenko, K. et al. (NVIDIA). *NVIDIA Nemotron Parse 1.1*. arXiv:2511.20478, ноябрь 2025. <a href="https://arxiv.org/abs/2511.20478" class="external free" rel="nofollow">https://arxiv.org/abs/2511.20478</a></span>
36. <span id="cite_note-ColEmbed-36">↑ <sup>[36.0](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-ColEmbed_36-0)</sup> <sup>[36.1](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-ColEmbed_36-1)</sup> de Souza P. Moreira, G. et al. (NVIDIA). *Nemotron ColEmbed V2: Top‑Performing Late Interaction Embedding Models for Visual Document Retrieval*. arXiv:2602.03992, февраль 2026. <a href="https://arxiv.org/abs/2602.03992" class="external free" rel="nofollow">https://arxiv.org/abs/2602.03992</a></span>
37. <span id="cite_note-Safety-37">↑ <sup>[37.0](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-Safety_37-0)</sup> <sup>[37.1](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-Safety_37-1)</sup> NVIDIA Developer Blog. *Safeguard Agentic AI Systems with the NVIDIA Safety Recipe*. 2025. <a href="https://developer.nvidia.com/blog/safeguard-agentic-ai-systems-with-the-nvidia-safety-recipe/" class="external free" rel="nofollow">https://developer.nvidia.com/blog/safeguard-agentic-ai-systems-with-the-nvidia-safety-recipe/</a></span>
38. <span id="cite_note-NemotronCC-38">[↑](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-NemotronCC_38-0) Su, D. et al. (NVIDIA). *Nemotron‑CC: Transforming Common Crawl into a Refined Long‑Horizon Pretraining Dataset*. ACL 2025 (Long Paper). arXiv:2412.02595. <a href="https://arxiv.org/abs/2412.02595" class="external free" rel="nofollow">https://arxiv.org/abs/2412.02595</a></span>
39. <span id="cite_note-NemotronCC_Math-39">[↑](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-NemotronCC_Math_39-0) Karimi Mahabadi, R. et al. (NVIDIA). *Nemotron‑CC‑Math: A 133 Billion‑Token‑Scale High Quality Math Pretraining Dataset*. arXiv:2508.15096, август 2025. <a href="https://arxiv.org/abs/2508.15096" class="external free" rel="nofollow">https://arxiv.org/abs/2508.15096</a></span>
40. <span id="cite_note-NemotronNano2_page-40">[↑](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-NemotronNano2_page_40-0) NVIDIA Research. *NVIDIA Nemotron Nano 2 and the Nemotron Pretraining Dataset v1*. <a href="https://research.nvidia.com/labs/adlr/NVIDIA-Nemotron-Nano-2/" class="external free" rel="nofollow">https://research.nvidia.com/labs/adlr/NVIDIA-Nemotron-Nano-2/</a></span>
41. <span id="cite_note-NeMo_synth-41">[↑](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-NeMo_synth_41-0) NVIDIA. *Synthetic Data Generation — NVIDIA NeMo Framework User Guide*. <a href="https://docs.nvidia.com/nemo-framework/user-guide/24.12/datacuration/syntheticdata.html" class="external free" rel="nofollow">https://docs.nvidia.com/nemo-framework/user-guide/24.12/datacuration/syntheticdata.html</a></span>
42. <span id="cite_note-AA-42">↑ <sup>[42.0](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-AA_42-0)</sup> <sup>[42.1](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-AA_42-1)</sup> Artificial Analysis. *NVIDIA Nemotron 3 Nano 30B‑A3B — Intelligence Index*. Февраль 2026. <a href="https://artificialanalysis.ai/models/nvidia-nemotron-3-nano-30b-a3b-reasoning" class="external free" rel="nofollow">https://artificialanalysis.ai/models/nvidia-nemotron-3-nano-30b-a3b-reasoning</a></span>
43. <span id="cite_note-LMArena-43">[↑](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-LMArena_43-0) LMArena (Chatbot Arena). Рейтинги моделей, по состоянию на февраль 2026. <a href="https://lmarena.ai/" class="external free" rel="nofollow">https://lmarena.ai/</a></span>
44. <span id="cite_note-Saplin-44">[↑](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-Saplin_44-0) Saplin, M. *Llama 3.1 Nemotron 70B: Quirks and Features*. DEV Community, 2024. <a href="https://dev.to/maximsaplin/llama-31-nemotron-70b-quirks-and-features-4nbg" class="external free" rel="nofollow">https://dev.to/maximsaplin/llama-31-nemotron-70b-quirks-and-features-4nbg</a></span>
45. <span id="cite_note-Guardrails-45">[↑](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-Guardrails_45-0) NVIDIA. *NeMo Guardrails*. GitHub, 2023–2025. <a href="https://github.com/NVIDIA/NeMo-Guardrails" class="external free" rel="nofollow">https://github.com/NVIDIA/NeMo-Guardrails</a> (arXiv:2310.10501).</span>
46. <span id="cite_note-License-46">[↑](https://systems-analysis.info/int/Nemotron_(NVIDIA)_(NL)#cite_ref-License_46-0) NVIDIA. *NVIDIA Nemotron Open Model License*. <a href="https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-nemotron-open-model-license/" class="external free" rel="nofollow">https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-nemotron-open-model-license/</a></span>
