Nemotron 3 Super:開源高效混合專家 Mamba-Transformer 模型,面向智能體推理
原始論文:Nemotron 3 Super: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning 作者:NVIDIA(數百位貢獻者) arXiv ID:2604.12374v1 日期:2025-04-17 標籤:LLM MoE Mamba Agentic RLHF Quantization 開源模型
1. 引言
Nemotron 3 Super 是 NVIDIA 推出的開源大型語言模型,採用混合專家 (Mixture-of-Experts, MoE) 架構,結合 Mamba-2 與 Transformer 的混合設計。模型總參數量 120B,活躍參數僅 12B,在智能體 (agentic) 任務上表現特別突出。
這個模型的核心賣點:
- 混合架構:Mamba-2 狀態空間模型 (SSM) + Transformer 注意力層 + LatentMoE,三者在 88 層中週期性交錯排列
- NVFP4 預訓練:首個在此規模以 NVFP4 低精度從頭預訓練的模型,用 25 兆 token
- 多 token 預測 (MTP):內建 MTP 頭支援推測解碼 (speculative decoding),在 SPEED-Bench 上平均接受長度達 3.45
- 強大的智能體能力:後訓練流程特別針對軟體工程、工具使用、終端操作等長程互動任務優化
- 高效推理:吞吐量比 GPT-OSS-120B 高 2.2 倍,同時維持更高的準確率
模型在 HuggingFace 上開放預訓練、後訓練及量化版本的 checkpoint。

2. 預訓練
2.1 架構
Nemotron 3 Super 的架構要點:
整體規格:
- 總參數:120.6B,活躍參數:12.7B
- 88 層,隱藏維度 6144
- 詞彙表大小:131,072
- 最大上下文長度:1M tokens
混合層設計:88 層以週期性模式交錯三種類型的層:
- Mamba-2 層:狀態空間模型,線性複雜度,適合長序列。SSM 狀態維度 256,卷積核大小 4
- 注意力層 (Attention):分組查詢注意力 (GQA),32 個查詢頭、2 個 KV 頭、頭維度 128。使用 RoPE 位置編碼(基頻 10,000)
- LatentMoE 層:這是一個關鍵創新。傳統 MoE 的 router 直接在完整隱藏維度 $d$ 上運作,LatentMoE 先把 token 從隱藏維度 $d$ 投影到一個較低的潛在維度 $\ell$,再由 router 從 512 個專家中選出 top-22 個。這讓每個專家可以做得很小(因為輸入維度只有 $\ell$),大幅減少 MoE 的 FLOP。共享專家 (shared expert) 也在潛在空間中運作

交錯模式為:每 8 層中有 6 層 Mamba-2、1 層 Attention、1 層 LatentMoE(具體比例隨深度略有調整)。


多 token 預測 (MTP):在模型頂部加入 MTP 頭,每一步預測未來多個 token。MTP 頭共享主模型的權重,因此不需要額外的 draft model。在推理時直接用 MTP 頭做推測解碼,SPEED-Bench 上平均接受長度 3.45 tokens。


2.2 NVFP4 預訓練
Nemotron 3 Super 是第一個在此規模以 NVFP4 精度從頭預訓練的模型。NVFP4 是 NVIDIA 的 4-bit 浮點格式,可在 Blackwell GPU 上加速。
混合精度策略:
- 大多數層用 NVFP4:包含 Mamba GEMM、MoE 路由專家 GEMM 等
- 最後 15% 的層用 BF16:因為最後幾層對精度最敏感
- 潛在投影 (latent projection) 用 BF16:LatentMoE 的投影矩陣保持高精度
- MTP、QKV、注意力計算用 BF16

一個有趣的發現:到訓練末期,零值權重梯度元素達到 7%。這是因為 NVFP4 量化的下溢 (underflow) 導致的,小梯度被量化為零後就無法更新權重。



2.3 預訓練資料
總訓練量:25 兆 (25T) tokens,採用兩階段課程 (curriculum)。
合成資料集(這些是新貢獻):
- 程式碼概念:1,500 萬題,從程式碼概念出發生成問題
- 無條件演算法:純演算法推理題
- 經濟學 MCQ:經濟學領域選擇題
- 形式邏輯:邏輯推理題
- 選擇題:350 萬題,MMLU 風格的多領域選擇題
資料混合遵循 Nemotron-CC 的分類法,依品質進行加權。


2.4 超參數
- 學習率排程:WSD (Warmup-Stable-Decay)
- 優化器:AdamW
- 批次大小:3,072 序列
- 序列長度:8,192
- MTP 損失權重:0.3
- 負載均衡:無輔助損失 (auxiliary-loss-free)
2.5 Checkpoint 合併
在穩定學習率 (stable LR) 階段,離線合併多個 checkpoint 可以帶來 2-4 分的提升。從三個時間窗口(125B、250B、500B tokens)中各產生候選合併結果,選最好的。

2.6 長上下文擴展
預訓練尾端加入長上下文 (LC) 階段:
- 先做 LC-Phase,逐步擴展上下文
- 接著做繼續預訓練 (CPT),把上下文推到 1M tokens
- 用了 34B + 17B tokens 的長上下文資料
2.7 基礎模型結果
基礎模型在多個 benchmark 上超越 Ling-flash-Base-2.0 和 GLM-4.5-Air-Base:
- 數學:AIME25 等數學推理任務表現強勁
- 程式碼:LiveCodeBench 表現優異
- 推理:GPQA 等高難度推理任務有競爭力
- 長上下文:RULER 在 256k、512k、1M 都維持 90%+ 準確率
3. 後訓練
後訓練流程如下(按順序執行):
SFT → RLVR(3 輪)→ RLHF(18B tokens)→ SWE-RL(20B tokens)→ MTP Healing

3.1 監督式微調 (SFT)
SFT 分兩階段:
- 第一階段:token-level 平均損失
- 第二階段:sample-level 平均損失(讓模型對不同長度的樣本給予同等重視)
SFT 資料組成(以圓餅圖比例):
- 智能體 (Agent):36%
- 推理 (Reasoning):31%
- 對話 (Chat):23%
- 長上下文 (Long Context):8%
- 其他 (Misc):2%
主要資料來源:
- 軟體工程:SWE-Gym、R2E-Gym、SWE-rebench 等真實 GitHub issue 解決資料
- 智能體 CLI 程式設計:15k 合成任務 + 3k SWE 任務 + 10k 網頁開發任務,從 Qwen3-Coder-480B 和 Minimax M2.5 蒸餾而來

- 對話式工具使用:6 階段生成管線,279K 對話涵蓋 838 個領域

- 通用工具呼叫 (function calling):150 萬條軌跡

- 長上下文、金融推理、CUDA、安全、搜尋、終端操作、多語言、結構化查詢語言等
推理控制:模型支援三種推理模式:
- 推理關閉 (reasoning-off):不做思考鏈
- 正常模式 (regular):完整推理
- 低算力模式 (low-effort):簡短推理,用 GPT-OSS-120B 的樣本(佔 SFT 資料的 2%)訓練
推理關閉模式的做法是從 3% 的樣本中移除推理過程。之後再加入 350 步的半線上策略 SFT 階段,收集模型自身的 rollout 並隨機截斷 12% 的推理痕跡,用於推理時間預算控制。
3.2 強化學習
RL 階段包含四個子階段:
3.2.1 第一階段:多環境 RLVR
採用統一的 RLVR 策略,在 21 個環境、37 個不同的 RL 資料集上同時訓練。同時訓練所有環境可以讓每次 RL 更新都涵蓋完整的環境分佈,避免在個別任務上退化。
環境涵蓋:
- 數學:競賽數學題,有 / 無 Python 執行工具,新增形式化證明驗證
- 程式碼:競賽風格程式碼 + 單步 patch 生成(為後續 SWE-RL 做準備)
- STEM:含新策劃的更有挑戰性的科學問題
- 指令遵循:多挑戰資料集,agent 需遵循複雜的使用者指令,reward 由預定 rubric 計算
- 安全:兩個環境,分別針對過度拒絕和越獄韌性。越獄資料用 PAIR 攻擊管線生成
- 長上下文:同 Nemotron 3 Nano
- 智能體工具使用:新增對話式工具使用和終端操作環境
- Reasoning Gym:多樣推理任務套件
低算力推理:在多環境 RL 過程中,部分 prompt 被轉換為低算力模式。reward 根據正確率和生成 token 數量兩者調整。低算力 prompt 的比例從 2% 逐步降到 1%。
3.2.2 第二階段:SWE-RL(軟體工程端到端 RL)
單獨做為一個階段,因為 SWE rollout 顯著比其他環境慢,需要更長的上下文,與短 horizon 環境混訓會造成吞吐量瓶頸。
每個 rollout 的流程:
- 啟動一個 Apptainer 容器(目標 repo 已在其中)
- 運行 OpenHands agent loop 產生程式碼 patch
- 對 ground-truth 測試進行評估,回傳 binary reward
為了增加工具多樣性,在 OpenHands 中實作了 OpenCode 和 Codex agent 類別,模擬 Claude Code 和 Codex CLI 的工具輸入/輸出格式。這樣多 harness 訓練提升了模型在不同 harness 下的泛化能力。
基礎設施要點:
- 用 Apptainer(而非 Docker)做容器隔離,因為叢集沒有 root 權限
- 記憶體 watchdog 監控 tmux process tree 的 RSS,超限時主動 kill
- 命令黑名單攔截
killall、pkill等危險命令 - 用
orjson(Rust 實作)替代 Python 標準json,加速大 payload 序列化
3.2.3 第三階段:RLHF
訓練一個大型 GenRM(生成式獎勵模型)來提供 RL 監督信號。GenRM 以 Qwen3-235B-A22B-Thinking-2507 為初始化,用 Helpsteer 3 資料集和 lmarena-140k 的商用友好子集訓練。
與 Nemotron 3 Nano 不同,這裡的 GenRM 在整個多環境 RL 過程中持續使用,並在最後加入一個純 RLHF 階段。
3.2.4 演算法
使用非同步 GRPO 設定:
- 訓練和推理解耦到不同的 GPU 裝置上
- 推理 worker 持續生成軌跡,存入 rollout buffer
- 收集足夠軌跡後送入訓練引擎
- 推理 worker 會在新權重可用時立即載入,不重算 KV cache
- 限制推理 worker 最多落後訓練一步,避免過度的策略滯後 (policy lag)
- 遮蔽 importance sampling ratio 以穩定訓練
多環境 RLVR 的具體設定:每步 256 個 prompt,每個 prompt 生成 16 個回應,batch size 4096(等於單次梯度更新),最大生成長度從 49K 開始後增至 64K。
PivotRL(智能體 RL):長 horizon 智能體任務面臨效率與準確性的兩難。SFT 便宜但容易 OOD,端到端 RL 準確但 rollout 成本極高。PivotRL 透過重用離線 SFT 專家軌跡中的「pivot」(模型不確定下一步行動的 informative turn)來訓練,reward 基於與專家行動的相似度而非完全匹配。應用於智能體程式設計、搜尋、終端操作和對話式工具使用。
3.2.5 基礎設施
- NeMo Gym + NeMo RL:NeMo Gym 是 rollout kernel,NeMo RL 是訓練 loop controller
- 非同步 RL:所有 RL 階段都用非同步架構,訓練和生成可獨立進行
- In-flight weight updates:訓練完成後不等推理結束就更新權重
- 韌性:擴展到 1k GPU 時遇到多種間歇性失敗,包括:
- 硬體故障需要完整重啟
- 埠衝突:1K GPU 規模下多組件(Ray、vLLM、OpenAI server、NeMo Gym server)搶埠,TOCTOU 競態條件頻繁發生
- 平行初始化、預取虛擬環境和二進位檔快取等優化加速啟動
第四階段:MTP Healing
在 RL 之後,凍結主模型權重,只訓練 MTP 頭。重用 RLVR 的 prompt,在模型生成的回應上用負對數似然損失(類似 SFT)訓練 MTP 頭。這個階段顯著提升了 MTP 的準確率。
3.3 後訓練模型評估
與 Qwen-3.5-122B-A10B 和 GPT-OSS-120B 的比較(Table 5 摘要):
| 類別 | Benchmark | N-3-Super | Qwen3.5-122B | GPT-OSS-120B |
|---|---|---|---|---|
| 知識 | MMLU-Pro | 83.73 | 86.70 | 81.00 |
| 推理 | AIME25 (no tools) | 90.21 | 90.36 | 92.50 |
| 推理 | HMMT Feb25 (no tools) | 93.67 | 91.40 | 90.00 |
| 推理 | GPQA (no tools) | 79.23 | 86.60 | 80.10 |
| 推理 | LiveCodeBench v5 | 81.19 | 78.93 | 88.00 |
| 推理 | HLE (with tools) | 22.82 | - | 19.0 |
| 智能體 | SWE-Bench (OpenHands) | 60.47 | 66.40 | 41.9 |
| 智能體 | SWE-Bench (Codex) | 53.73 | 61.20 | - |
| 智能體 | SWE-Bench Multilingual | 45.78 | - | 30.80 |
| 智能體 | TauBench V2 Average | 61.15 | 74.53 | 61.0 |
| 智能體 | Terminal Bench (hard) | 25.78 | 26.80 | 24.00 |
| 智能體 | BrowseComp with Search | 31.28 | - | 33.89 |
| 智能體 | BIRD Bench | 41.80 | - | 38.25 |
| 對話 | IFBench | 72.56 | 73.77 | 68.32 |
| 對話 | Arena-Hard-V2 | 73.88 | 75.15 | 90.26 |
| 長上下文 | RULER 256k | 96.83 | 96.74 | 52.30 |
| 長上下文 | RULER 1M | 91.64 | 91.33 | 22.30 |
| 多語言 | MMLU-ProX | 79.36 | 85.06 | 76.59 |
重點觀察:
- 推理能力與 Qwen3.5-122B 和 GPT-OSS-120B 處於同一水準,但 Nemotron 3 Super 的活躍參數只有 12.7B
- 智能體任務全面優於或接近 GPT-OSS-120B,部分 harness 接近 Qwen 3.5 122B
- 長上下文(RULER 1M: 91.64)遠超 GPT-OSS-120B (22.30)
- 多語言能力(WMT24++ en→xx: 86.67)超越兩個 baseline
4. 推理量化
使用 Model-Optimizer 進行訓後量化 (PTQ),產生兩種部署版本:
- FP8 (W8A8):用於 Hopper GPU
- NVFP4 (W4A4):用於 Blackwell GPU
4.1 FP8 Checkpoint
用 256 個樣本(65536 上下文長度,來自 SFT 資料集)做校準。
精度配置:
| 組件 | FP8 Checkpoint | BF16 基線 |
|---|---|---|
| Embedding | BF16 | BF16 |
| Attention QKV/Out Projection | BF16 | BF16 |
| KV Cache + Attention BMM1 | FP8 | FP8 |
| Attention BMM2 | BF16 | BF16 |
| MoE GEMM (稀疏 + 共享專家) | FP8 | BF16 |
| MoE 潛在投影 GEMM | BF16 | BF16 |
| Router | FP32 | FP32 |
| Mamba GEMM | FP8 | BF16 |
| Mamba SSM Cache | FP16 | BF16 |
4.2 NVFP4 Checkpoint
NVFP4 是比 FP8 更激進的量化格式,特別適合 prefill-heavy 的推理負載(如程式碼智能體),因為線性層和 MoE GEMM 是主要效能瓶頸。NVFP4 在 Blackwell GPU 上原生加速,且比 MXFP4 等替代方案有更好的精度。
NVFP4 PTQ 配方的三個要素:
- 校準的逐 block 權重縮放:用最小化權重 MSE(而非 max-based)來選擇縮放因子
- 動態逐 block max-based 啟動值縮放:啟動值必須在執行時高效計算,所以用 max-based
- Model-Optimizer AutoQuantize 選擇性精度提升:受神經架構搜索 (NAS) 啟發,AutoQuantize 估計每個算子的效能成本和精度敏感度,用背包問題 (knapsack) 式的最佳化在有效精度預算 4.75 bits 下求解最優配置。靈敏度指標基於 Optimal Brain Surgeon 的二階 Taylor 近似
最終配置:稀疏專家 GEMM 全部 NVFP4,注意力和 Mamba 投影用 FP8 或 BF16 混合,只有少量層保留在 FP8/BF16 以保精度。
混合精度 PTQ 在單台 B200 節點(8 GPU)上不到 2 小時完成,使用 512 個樣本。結果模型達到 BF16 基線的 99.8% 中位精度。
量化後評估(Table 8 摘要):
| Benchmark | BF16 | FP8 | NVFP4 |
|---|---|---|---|
| MMLU-Pro | 83.73 | 83.63 | 83.33 |
| HMMT Feb25 (with tools) | 94.73 | 94.38 | 95.36 |
| GPQA (no tools) | 79.23 | 79.36 | 79.42 |
| SWE-Bench (OpenHands) | 60.47 | - | 59.90 |
| TauBench V2 Average | 61.15 | 61.07 | 60.46 |
| IFBench | 72.58 | 72.32 | 73.30 |
| RULER 256k | 96.83 | 96.84 | 96.81 |
| RULER 1M | 91.64 | 91.43 | 91.60 |
FP8 和 NVFP4 的精度損失都非常小,NVFP4 甚至在某些 benchmark(如 HMMT、GPQA)上略有提升。
4.3 Mamba 狀態量化
Mamba 解碼是記憶體頻寬受限的,SSM cache 的 DRAM 讀取是解碼速度的主要瓶頸。
核心挑戰:Mamba 解碼是遞迴的,量化誤差會從前一步傳播到後續所有步驟,而且會累積。設遞迴狀態更新為 $h_t = A_t h_{t-1} + B_t x_t$,cache 量化在第 $t$ 步引入加性誤差 $e_t$,展開遞迴得:
$$h_{q,t} = h_t + \sum_{i=0}^{t} \left(\prod_{j=i+1}^{t} A_j\right) e_i$$
量化誤差通過後續遞迴轉移累積,這跟 Transformer 的 KV cache 量化有本質區別。
解法:
- 直接將 SSM cache 從 FP32 降到 FP16 會導致高達 40% 的冗長度 (verbosity) 增加
- 嘗試 INT16 逐 block 縮放,可以維持精度,但分析發現 SSM cache 動態範圍大,naive INT16 沒效果
- 引入 FP32 逐 block 縮放(block 大小 128),提升有效動態範圍
- 發現誤差累積與四捨五入模式有關:Round-to-Nearest-Even (RTNE) 有偏差(將值映射到同一個四捨五入值時,量化誤差的期望值非零),在遞迴設定下偏差會一致累積。而隨機四捨五入 (Stochastic Rounding, SR) 是無偏的
- 最終選擇 FP16 + Philox<5> 偽隨機數隨機四捨五入做為 SSM cache 配方,原因:
- 不需要計算、儲存和載入 block 縮放因子
- Blackwell 有專用的 PTX 指令支援隨機四捨五入
- Blackwell 透過 cuRAND 支援 Philox 偽隨機數生成
Philox 的輪數會影響品質:輪數越多統計品質越好但開銷越大。Philox<5> 是精度與開銷的平衡點。
5. 結論
Nemotron 3 Super 是一個 12B 活躍/120B 總參數的 MoE 混合 Mamba-Attention 模型,具備強大的智能體能力。它採用 LatentMoE 提升精度,內建 MTP 層透過推測解碼加速推理。在 25 兆 token 上以低精度 NVFP4 預訓練,再經多樣化 RL 環境的後訓練。量化到 FP8 和 NVFP4 後,推理吞吐量大幅提升而幾乎不損失準確率。Nemotron 3 Super 的吞吐量比 GPT-OSS-120B 高 2.2 倍,同時在廣泛任務上維持更高的準確率。所有 checkpoint 均在 HuggingFace 上開放。
參考文獻
Talor Abramovich, Maor Ashkenazi, Carl (Izzy) Putterman, Benjamin Chislett, Tiyasa Mitra, Bita Darvish Rouhani, Ran Zilberstein, and Yonatan Geifman. SPEED-Bench: A Unified and Diverse Benchmark for Speculative Decoding, 2026.
Syeda Nahida Akter, Shrimai Prabhumoye, John Kamalu, Sanjeev Satheesh, Eric Nyberg, Mostofa Patwary, Mohammad Shoeybi, and Bryan Catanzaro. MIND: Math Informed syNthetic Dialogues for Pretraining LLMs, 2024.
Syeda Nahida Akter, Shrimai Prabhumoye, Eric Nyberg, Mostofa Patwary, Mohammad Shoeybi, Yejin Choi, and Bryan Catanzaro. Front-loading reasoning: The synergy between pretraining and post-training data. In ICLR, 2026.
Jacob Austin, Augustus Odena, Maxwell Nye, et al. Program Synthesis with Large Language Models, 2021.
Ibragim Badertdinov, Alexander Golubev, et al. Swe-rebench: An automated pipeline for task collection and decontaminated evaluation of software engineering agents, 2025.
Victor Barres, Honghua Dong, et al. τ²-Bench: Evaluating Conversational Agents in a Dual-Control Environment, 2025.
Noga BenYoash, Menachem Brief, et al. Secque: A benchmark for evaluating real-world financial analysis capabilities, 2025.
Yonatan Bisk, Rowan Zellers, et al. PIQA: Reasoning about Physical Commonsense in Natural Language, 2019.
Patrick Chao, Alexander Robey, et al. Jailbreaking black box large language models in twenty queries, 2024.
Mark Chen, Jerry Tworek, et al. Evaluating Large Language Models Trained on Code, 2021.
Wei-Lin Chiang, Lianmin Zheng, et al. Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference, 2024.
Peter Clark, Isaac Cowhey, et al. Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge, 2018.
Karl Cobbe, Vineet Kosaraju, et al. Training Verifiers to Solve Math Word Problems, 2021.
Damai Dai, Chengqi Deng, et al. DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models, 2024.
Tri Dao and Albert Gu. Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality, 2024.
DeepSeek-AI. Deepseek-v3.2-exp, 2025a.
DeepSeek-AI. DeepSeek-R1, 2025b.
DeepSeek-AI. DeepSeek-V3 Technical Report, 2025c.
DeepSeek-AI, Aixin Liu, et al. Deepseek-v3.2, 2025.
Daniel Deutsch, Eleftheria Briakou, et al. WMT24++: Expanding the Language Coverage of WMT24 to 55 Languages & Dialects. ACL 2025.
Wei Du, Shubham Toshniwal, et al. Nemotron-math: Efficient long-context distillation of mathematical reasoning from multi-mode supervision, 2025.
Venmugil Elango, Nidhi Bhatia, et al. LatentMoE: Toward Optimal Accuracy per FLOP and Parameter in Mixture of Experts, 2026.
Steven Feng, Shrimai Prabhumoye, et al. Maximize Your Data's Potential: Enhancing LLM Accuracy with Two-Phase Pretraining, 2024.
Elias Frantar, Saleh Ashkboos, et al. Gptq: Accurate post-training quantization for generative pre-trained transformers. ICLR, 2023.
Shaona Ghosh, Prasoon Varshney, et al. AEGIS2.0: A Diverse AI Safety Dataset and Risks Taxonomy for Alignment of LLM Guardrails. NAACL 2025.
Glaive AI. glaiveai/glaive-function-calling-v2, 2025.
GLM-4.5-Team. GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models, 2025.
Fabian Gloeckle, Badr Youbi Idrissi, et al. Better & Faster Large Language Models via Multi-token Prediction. ICML, 2024.
Gretel. Gretel Synthetic Safety Alignment Dataset, 2024.
Alex Gu, Baptiste Rozière, et al. Cruxeval: A benchmark for code reasoning, understanding and execution, 2024.
Melody Y. Guan, Manas Joglekar, et al. Deliberative alignment: Reasoning enables safer language models, 2025.
Steve Harris and Andy Seaborne. SPARQL 1.1 Query Language. W3C Recommendation, 2013.
Adish Hasan, Ileana Rugina, and Alex Wang. Pruning for Protection, 2024.
B. Hassibi, D.G. Stork, and G.J. Wolff. Optimal brain surgeon and general network pruning. IEEE ICNN, 1993.
Dan Hendrycks, Collin Burns, et al. Measuring Massive Multitask Language Understanding, 2021a.
Dan Hendrycks, Collin Burns, et al. Measuring massive multitask language understanding. ICLR, 2021b.
Dan Hendrycks, Collin Burns, et al. Measuring Mathematical Problem Solving With the MATH Dataset, 2021c.
Cheng-Ping Hsieh, Simeng Sun, et al. RULER: What's the Real Context Size of Your Long-Context Language Models? 2024.
Shengding Hu, Yuge Tu, et al. MiniCPM: Unveiling the potential of small language models with scalable training strategies, 2024.
Shijue Huang, Wanjun Zhong, et al. Planning, creation, usage: Benchmarking llms for comprehensive tool utilization, 2024.
Zongle Huang, Lei Zhu, et al. Moesd: Unveil speculative decoding's potential for accelerating sparse moe, 2025.
Naman Jain, King Han, et al. Livecodebench: Holistic and contamination free evaluation of large language models for code, 2024.
Naman Jain, Jaskirat Singh, et al. R2E-Gym: Procedural Environments and Hybrid Verifiers for Scaling Open-Weights SWE Agents, 2025.
Carlos E Jimenez, John Yang, et al. SWE-bench: Can Language Models Resolve Real-World GitHub Issues? 2023.
Woosuk Kwon, Zhuohan Li, et al. Efficient Memory Management for Large Language Model Serving with PagedAttention. ACM SIGOPS, 2023.
Hynek Kydlíček, Guilherme Penedo, and Leandro von Werra. Finepdfs, 2025.
Guokun Lai, Qizhe Xie, et al. RACE: Large-scale ReAding comprehension dataset from examinations. EMNLP, 2017.
Hugo Laurençon, Lucile Saulnier, et al. The bigscience roots corpus, 2023.
Hieu Tran, Ian Yu, et al. Terminal-bench: Benchmarking agents on hard, realistic tasks in command line interfaces, 2026.
Dmitry Lepikhin, HyoukJoong Lee, et al. GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding, 2020.
Jinyang Li, Binyuan Hui, et al. Can LLM Already Serve As A Database Interface? NeurIPS, 2023a.
Minghao Li, Yingxiu Zhao, et al. Api-bank: A comprehensive benchmark for tool-augmented llms. EMNLP 2023, 2023b.
Shiyao Li, Xuefei Ning, et al. Llm-mq: Mixed-precision quantization for efficient llm deployment. NeurIPS, 2023c.
Xuehai Li, Zi Ye, et al. WildChat: 1M ChatGPT Interaction Logs in the Wild, 2024.
Ling Team. Every activation boosted: Scaling general reasoner to 1 trillion open language foundation, 2025.
Jiawei Liu, Chunqiu Steven Xia, et al. Is Your Code Generated by ChatGPT Really Correct? 2023.
Zuxin Liu et al. Toolace: Winning the points of llm function calling, 2024.
Ilya Loshchilov and Frank Hutter. Decoupled Weight Decay Regularization, 2017.
Weidi Luo, Siyuan Ma, et al. JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks, 2024.
Rabeeh Karimi Mahabadi, Sanjeev Satheesh, et al. Nemotron-CC-Math, 2025.
Mike A. Merrill, Alexander G. Shaw, et al. Terminal-bench, 2026.
Yev Meyer and Dane Corneil. Nemotron-Personas-USA, 2025.
Todor Mihaylov, Peter Clark, et al. Can a suit of armor conduct electricity? EMNLP, 2018a.
MiniMax AI. MiniMax-M2, 2025.
NVIDIA. NeMo Gym, 2025a.
NVIDIA. NeMo RL, 2025b.
NVIDIA. Nemotron Nano 2, 2025c.
NVIDIA. Fine-Tuning gpt-oss for Accuracy and Performance with Quantization Aware Training, 2025.
NVIDIA. Nemotron 3 nano, 2025a.
NVIDIA. Nemotron-H, 2025b.
NVIDIA. NVIDIA Nemotron 3, 2025c.
NVIDIA. Pretraining large language models with nvfp4, 2025d.
NVIDIA. Introducing nvfp4 for efficient and accurate low-precision inference, 2025.
術語對照表
| 英文 | 中文 |
|---|---|
| Mixture-of-Experts (MoE) | 混合專家 |
| State Space Model (SSM) | 狀態空間模型 |
| LatentMoE | 潛在混合專家 |
| Multi-Token Prediction (MTP) | 多 token 預測 |
| Speculative Decoding | 推測解碼 |
| Grouped Query Attention (GQA) | 分組查詢注意力 |
| Rotary Position Embedding (RoPE) | 旋轉位置編碼 |
| Post-Training Quantization (PTQ) | 訓後量化 |
| Reinforcement Learning from Human Feedback (RLHF) | 基於人類回饋的強化學習 |
| Reinforcement Learning from Verifiable Rewards (RLVR) | 基於可驗證獎勵的強化學習 |
| Supervised Fine-Tuning (SFT) | 監督式微調 |
| Generative Reward Model (GenRM) | 生成式獎勵模型 |
| AutoQuantize | 自動量化 |
| Stochastic Rounding | 隨機四捨五入 |
| Round-to-Nearest-Even (RTNE) | 四捨六入五取偶 |
| Warmup-Stable-Decay (WSD) | 預熱-穩定-衰減 |
| Rollout | 展開(策略採樣) |
| Policy Lag | 策略滯後 |
| Importance Sampling | 重要性採樣 |
| In-flight Weight Updates | 飛行中權重更新 |
| Checkpoint Merging | 檢查點合併 |
| Long Context Extension | 長上下文擴展 |
| Verbosity | 冗長度 |
| Agentic | 智能體式 |
| PivotRL | 樞紐強化學習 |
| Harness | 測試框架 |
| Apptainer | 應用容器(前身 Singularity) |
| Knapsack Optimization | 背包最佳化 |
| Optimal Brain Surgeon | 最優腦外科手術 |