Insight certainty 4/5 debate potential 2/5

Mascorro: High-quality LLMs can be built purely with SFT data

Marco Mascorro · DeepSeek, Reasoning Models, and the Future of LLMs · Mar 5, 2025 · at 4:02

Marco Mascorro, partner at a16z, discusses language model post-training techniques and data sources like Stack Overflow and Reddit.

0:00 / 0:05exact quote · 5.2s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“Now, the reality is like you can get to really good models purely with like SFT data.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Marco Mascorro

Insight
Mascorro: DeepSeek-R1 proved reinforcement learning improves models without human feedback
“And I think the big thing in, in R-one, or generally with these reasoning models is, We were doing before there was a human in the loop always, right? Like when we have this SFT training and these other techniques that we're doing after like RLHF and having R …”
Marco Mascorro Mar 5, 2025 ▶ 6:42 DeepSeek, Reasoning Models, and the Future of LLMs
Assertion Supported
Mascorro: Distillations from DeepSeek-R1 Outperformed Direct RL on Smaller Models
“So it turns out in their experiments, they took Lama's EV and some of these are QN models, and they basically apply RL straight the same way they did it with R one on these base models. And it turns out that it improved in some fields, but it was not a signifi…”
Marco Mascorro Mar 5, 2025 ▶ 25:26 DeepSeek, Reasoning Models, and the Future of LLMs
Assertion Supported
Mascorro: DeepSeek consistently open-sources its model weights and training techniques
“So, one of the good things about DeepSeek is basically they open source their weights, their techniques, and how they build these models, and they've been doing that for a while.”
Marco Mascorro Mar 5, 2025 ▶ 0:19 DeepSeek, Reasoning Models, and the Future of LLMs
Assertion Supported
Mascorro: DeepSeek-R1-Zero Improved Math Scores but Struggled with Readability and Language Switching
“R one zero, which in a way was a very interesting model because it showed that it improved in some reasoning benchmarks and math benchmarks. But eventually didn't do really well on other things, right? Like it was switching between languages. I think that was …”
Marco Mascorro Mar 5, 2025 ▶ 7:53 DeepSeek, Reasoning Models, and the Future of LLMs
Assertion Supported
Mascorro: DeepSeek-V3 features 256 experts, far exceeding typical open-source models
“We talk about it as 256 experts, which is a large, a relative large number of experts in terms of at least open source models that we've seen out there.”
Marco Mascorro Mar 5, 2025 ▶ 9:00 DeepSeek, Reasoning Models, and the Future of LLMs
Assertion Supported
Mascorro: DeepSeek-R1 post-training used two SFT and two RL phases
“So basically the way they did that, trying to fix R one zero, is it added a couple more phases in the post-training. That included two supervised fine tuning phases and two reinforcement learning phases. And these reinforcement learning phases, they were a lar…”
Marco Mascorro Mar 5, 2025 ▶ 11:25 DeepSeek, Reasoning Models, and the Future of LLMs
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.