Nemotron, every mention
19 scenes, the whole family · ← back to Nemotron
tap a year for its mentions
every year anyone Pratyush Maini 4Ronak Malde 3Shawn Wang 2Philip Kiely 2Ethan He 2Ari Morcos 2Eiso Kant 1Andy Beam 1Alessio Fanelli 1
Verbatim, from the transcripts: the passages where Nemotron comes up
⏭️ Forward Deployed: Voice AI on what works in 2026
- ▶ 33:41 unnamed speaker The whole point of that is that, as I think today, Pneumotron three-point-five launched, and there's already an ASR, um, benchmark out, did very well, the NVIDIA one.
Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
- ▶ 57:03 Philip Kiely If you look at, for example, NVIDIA Nemotron models, they run very, very well on Blackwell.
- ▶ 1:13:51 Philip Kiely Um, so yeah, it's, it's, it's mostly in my mind about, uh, model size, and then about matching the architecture and the native quantization to the, the target hardware, like we talked about with, like, you know, all NemoTron models or NVFP…
The AI Frontier: from open weights to open research — Eiso Kant, Poolside AI
🔬 RL with Verifiable Rewards, but the Verifier is a Lab — Lila Sciences
⚡️Every product of the future will be a living system — Ronak Malde, Trajectory.ai
- ▶ 11:29 Ronak Malde Harvey and NVIDIA in order to train Nimachan three, super, in order to get to the Pareto frontier. 2 times in the scene
- ▶ 15:06 Ronak Malde Like, since then, Nematron has gotten significantly better.
Agent Inference at the "Speed of Light" — How NVIDIA moves like a $4.3 Trillion Startup
- ▶ 22:18 unnamed speaker You expect foundation models with Nemo Tron voice just randomly popped here.
- ▶ 37:59 unnamed speaker It's exciting when you guys put out something like Nemotron, because I remember the paper on this, Nemotron III, the amount of, like, post-training, the amount of tokens that the GPU rich can just train on, and it was a hybrid state-space… 3 times in the scene
- ▶ 37:59 unnamed speaker It's exciting when you guys put out something like Nemotron, because I remember the paper on this, Nemotron III, the amount of, like, post-training, the amount of tokens that the GPU rich can just train on, and it was a hybrid state-space…
⚡️ Reverse Engineering OpenAI's Training Data — Pratyush Maini, Datology
- ▶ 21:22 Pratyush Maini So, the NemoTron dataset is, like, one of the top datasets today, which is, like, based with a lot of synthetic data, uh, so the, and the Datology, uh, data that we, a model that we released called BeyondWeb, uh, is the blue line over here. 4 times in the scene
Captaining IMO Gold, Deep Think, On-Policy RL, Feeling the AGI in Singapore — Yi Tay
- ▶ 53:25 unnamed speaker Would you say that the ideas that I see there, Nvidia has NemoTron, OpenAI has GPT-OSS, these are all basically checkpoints on what's publicly known about training models as of this year.
Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
- ▶ 56:10 Shawn Wang And also a special shout out to Nvidia Nemo Tron, which doesn't get enough credit for the amount of stuff that they do. 2 times in the scene
Better Data is All You Need — Ari Morcos, Datology
- ▶ 29:12 Ari Morcos Like, like, Nematron is actually pretty similar in quality to DCLM.
- ▶ 1:07:19 Ari Morcos We started with a combination of DCLM, Nemetron, and FineWeb.
[Paper Club] Upcycling Large Language Models into Mixture of Experts
[Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu
- ▶ 59:41 unnamed speaker So, for next week, right, for next week, we will be going, uh, we have, let me see, Ethan from NVIDIA who's going through the Megatron, uh, NemoTron distillation or Amoe, um, I'm not so sure which one, um, and yep, look forward to it, and…
The Winds of AI Winter (Q2 Four Wars of the AI Stack Recap)
- ▶ 44:02 unnamed speaker Yeah, so we can skip Nemo.
Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
- ▶ 26:01 Alessio Fanelli Do you have any other thoughts on the more synthetic data-focused models, kind of like a Mnemotron?