中文

Nemotron 3 Nano:面向智能体推理的开源、高效 Mixture-of-Experts 混合 Mamba-Transformer 模型

机器人学 2025-12-25 v1 人机交互

摘要

我们介绍了 Nemotron 3 Nano 30B-A3B,一种 Mixture-of-Experts 混合 Mamba-Transformer 语言模型。Nemotron 3 Nano 在 25 万亿文本 token 上进行预训练,其中包括超过 3 万亿个新的唯一 token,随后进行监督微调和大规模 RL 训练。Nemotron 3 Nano 在准确率上优于我们上一代的 Nemotron 2 Nano,但每前向传播仅激活不到一半的参数。其推理吞吐量比同样规模的开源模型如 GPT-OSS-20B 和 Qwen3-30B-A3B-Thinking-2507 最高可达 3.3 倍,同时在流行基准测试上也更准确。Nemotron 3 Nano 显示出增强的智能体、推理和聊天能力,并支持最长 100 万 token 的上下文长度。我们在 Hugging Face 上发布了我们的预训练 Nemotron 3 Nano 30B-A3B Base 以及经过微调的 Nemotron 3 Nano 30B-A3B checkpoint。

关键词

引用

@article{arxiv.2512.20847,
  title  = {YCB-Handovers Dataset: Analyzing Object Weight Impact on Human Handovers to Adapt Robotic Handover Motion},
  author = {Parag Khanna and Karen Jane Dsouza and Chunyu Wang and Mårten Björkman and Christian Smith},
  journal= {arXiv preprint arXiv:2512.20847},
  year   = {2025}
}

备注

Paper presented at the IEEE International Conference on Robot and Human Interactive Communication (RO-MAN), 2025