Jan 29, 2026 · 1h 8m · mad

State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka

Sebastian Raschka · 57m spoken Matt Turck · 7m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of The MAD Podcast, AI researcher Sebastian Raschka joins Matt Turck to discuss the state of large language models in 2026. They cover the shift from pre-training to post-training methods like RLVR and GRPO, inference-time scaling, Transformer optimizations, benchmark saturation, and enterprise AI strategies.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 11.2% of the talking time here. How this is scored →

Matt as informed peer 4.3 Guest teaching 4.9 Guest disagreement 1.3 Matt pushing back 1.3
05100:0015:0030:0045:001:00:000:59–4:05 · Matt as informed peer 5/10 Evaluating the Transformer Architecture and Modern Alternatives The host opens with an informed question about whether the transformer architecture's days are numbered given developments in SSMs and diffusion models. The guest explains that transformers remain state-of-the-art while modern variants mostly focus on cost efficiency.4:05–9:46 · Matt as informed peer 4/10 World Models and Internal State Prediction for Code Generation The host asks targeted questions about world models and small recursive models. The guest educates the host on internal state prediction for code LLMs and how recursive models target specific logic benchmarks like ARC.9:46–13:24 · Matt as informed peer 5/10 Text Diffusion Models versus Auto-Regressive Transformers The host demonstrates strong knowledge by bringing up DeepMind's Gemini Diffusion announcement. The guest explains the architectural differences between sequential auto-regressive generation and parallel diffusion denoising.13:24–18:03 · Matt as informed peer 5/10 LLM Architectural Optimizations, MoE Adoption, and Pre-Training Limits The host pushes the guest on whether architecture progress is just superficial tuning and checks if pre-training is dead. The guest re-frames pre-training as 'boring' rather than dead and highlights MoE adoption trends.18:03–24:42 · Matt as informed peer 5/10 Post-Training Evolution: RLVR, GRPO, and Reasoning Unlocking The host cites the guest's technical blog post to prompt an explanation of RLVR and GRPO. The guest provides an in-depth breakdown of how GRPO eliminates memory overhead by getting rid of separate reward and value models.24:42–30:21 · Matt as informed peer 5/10 Process Reward Models (PRMs) and Verifiable Rewards Beyond Code The host prompts the guest regarding why Process Reward Models (PRMs) haven't been successful yet. The guest outlines reward hacking risks in PRMs and explains multi-model grader setups.30:27–35:17 · Matt as informed peer 4/10 Cost and Complexity of Scaling RLVR vs. Pre-Training The host suggests RL scaling is notoriously finicky and complex based on the multi-model architecture. The guest counters that it is not overly complicated, citing his own 39-page Jupyter notebook implementation and cost differences.35:17–38:40 · Matt as informed peer 4/10 Meta-Lessons on AI Progress and Infrastructure Specialization The host asks about the meta-lessons from 2025 AI progress. The guest explains that progress is an aggregation of small engineering micro-optimizations across specialized sub-teams rather than a single magic bullet.38:40–43:11 · Matt as informed peer 4/10 "Bench-Maxing", Leaderboard Gaming, and Evaluation Reality The host brings up 'bench-maxing' and asks if corporate economic motives drive leaderboard gaming. The guest gives a nuanced breakdown of how evaluation style bias distorts leaderboards.43:11–46:46 · Matt as informed peer 4/10 Inference-Time Scaling and Recursive Prompt Chunking The host steers the discussion toward non-post-training factors like inference scaling. The guest explains parallel sampling, vote aggregation, and prompt chunking.46:46–50:11 · Matt as informed peer 3/10 System Engineering, Tool Calling, and Local Execution Security The host listens as the guest explains how tool calling and prompt wrappers create a gap between raw open-weight models and cloud LLM platforms.50:11–55:13 · Matt as informed peer 5/10 Generalist Model Parity and In-House Enterprise LLM Training The host links generalist model parity with private enterprise data as a moat, asking if companies are resuming in-house training. The guest clarifies that large enterprises are indeed building data-center scale models in-house.55:13–59:25 · Matt as informed peer 4/10 Evaluating Continual Learning and Iterative Model Updates The host brings up NeurIPS buzz around continual learning and asks about the guest's skeptical timeline. The guest explains why catastrophic forgetting and hosted API paradigms delay implementation.59:25–1:04:55 · Matt as informed peer 3/10 Sebastian Raschka's Code-First Research and Writing Methodology The host asks the guest about his research methodology, book authoring, and use of LLMs. The guest describes his code-first approach of implementing architectures from scratch to verify paper claims.0:59–4:05 · Guest teaching 4/10 Evaluating the Transformer Architecture and Modern Alternatives The host opens with an informed question about whether the transformer architecture's days are numbered given developments in SSMs and diffusion models. The guest explains that transformers remain state-of-the-art while modern variants mostly focus on cost efficiency.4:05–9:46 · Guest teaching 5/10 World Models and Internal State Prediction for Code Generation The host asks targeted questions about world models and small recursive models. The guest educates the host on internal state prediction for code LLMs and how recursive models target specific logic benchmarks like ARC.9:46–13:24 · Guest teaching 5/10 Text Diffusion Models versus Auto-Regressive Transformers The host demonstrates strong knowledge by bringing up DeepMind's Gemini Diffusion announcement. The guest explains the architectural differences between sequential auto-regressive generation and parallel diffusion denoising.13:24–18:03 · Guest teaching 5/10 LLM Architectural Optimizations, MoE Adoption, and Pre-Training Limits The host pushes the guest on whether architecture progress is just superficial tuning and checks if pre-training is dead. The guest re-frames pre-training as 'boring' rather than dead and highlights MoE adoption trends.18:03–24:42 · Guest teaching 6/10 Post-Training Evolution: RLVR, GRPO, and Reasoning Unlocking The host cites the guest's technical blog post to prompt an explanation of RLVR and GRPO. The guest provides an in-depth breakdown of how GRPO eliminates memory overhead by getting rid of separate reward and value models.24:42–30:21 · Guest teaching 5/10 Process Reward Models (PRMs) and Verifiable Rewards Beyond Code The host prompts the guest regarding why Process Reward Models (PRMs) haven't been successful yet. The guest outlines reward hacking risks in PRMs and explains multi-model grader setups.30:27–35:17 · Guest teaching 6/10 Cost and Complexity of Scaling RLVR vs. Pre-Training The host suggests RL scaling is notoriously finicky and complex based on the multi-model architecture. The guest counters that it is not overly complicated, citing his own 39-page Jupyter notebook implementation and cost differences.35:17–38:40 · Guest teaching 4/10 Meta-Lessons on AI Progress and Infrastructure Specialization The host asks about the meta-lessons from 2025 AI progress. The guest explains that progress is an aggregation of small engineering micro-optimizations across specialized sub-teams rather than a single magic bullet.38:40–43:11 · Guest teaching 5/10 "Bench-Maxing", Leaderboard Gaming, and Evaluation Reality The host brings up 'bench-maxing' and asks if corporate economic motives drive leaderboard gaming. The guest gives a nuanced breakdown of how evaluation style bias distorts leaderboards.43:11–46:46 · Guest teaching 5/10 Inference-Time Scaling and Recursive Prompt Chunking The host steers the discussion toward non-post-training factors like inference scaling. The guest explains parallel sampling, vote aggregation, and prompt chunking.46:46–50:11 · Guest teaching 5/10 System Engineering, Tool Calling, and Local Execution Security The host listens as the guest explains how tool calling and prompt wrappers create a gap between raw open-weight models and cloud LLM platforms.50:11–55:13 · Guest teaching 5/10 Generalist Model Parity and In-House Enterprise LLM Training The host links generalist model parity with private enterprise data as a moat, asking if companies are resuming in-house training. The guest clarifies that large enterprises are indeed building data-center scale models in-house.55:13–59:25 · Guest teaching 5/10 Evaluating Continual Learning and Iterative Model Updates The host brings up NeurIPS buzz around continual learning and asks about the guest's skeptical timeline. The guest explains why catastrophic forgetting and hosted API paradigms delay implementation.59:25–1:04:55 · Guest teaching 4/10 Sebastian Raschka's Code-First Research and Writing Methodology The host asks the guest about his research methodology, book authoring, and use of LLMs. The guest describes his code-first approach of implementing architectures from scratch to verify paper claims.0:59–4:05 · Guest disagreement 1/10 Evaluating the Transformer Architecture and Modern Alternatives The host opens with an informed question about whether the transformer architecture's days are numbered given developments in SSMs and diffusion models. The guest explains that transformers remain state-of-the-art while modern variants mostly focus on cost efficiency.4:05–9:46 · Guest disagreement 1/10 World Models and Internal State Prediction for Code Generation The host asks targeted questions about world models and small recursive models. The guest educates the host on internal state prediction for code LLMs and how recursive models target specific logic benchmarks like ARC.9:46–13:24 · Guest disagreement 1/10 Text Diffusion Models versus Auto-Regressive Transformers The host demonstrates strong knowledge by bringing up DeepMind's Gemini Diffusion announcement. The guest explains the architectural differences between sequential auto-regressive generation and parallel diffusion denoising.13:24–18:03 · Guest disagreement 2/10 LLM Architectural Optimizations, MoE Adoption, and Pre-Training Limits The host pushes the guest on whether architecture progress is just superficial tuning and checks if pre-training is dead. The guest re-frames pre-training as 'boring' rather than dead and highlights MoE adoption trends.18:03–24:42 · Guest disagreement 1/10 Post-Training Evolution: RLVR, GRPO, and Reasoning Unlocking The host cites the guest's technical blog post to prompt an explanation of RLVR and GRPO. The guest provides an in-depth breakdown of how GRPO eliminates memory overhead by getting rid of separate reward and value models.24:42–30:21 · Guest disagreement 1/10 Process Reward Models (PRMs) and Verifiable Rewards Beyond Code The host prompts the guest regarding why Process Reward Models (PRMs) haven't been successful yet. The guest outlines reward hacking risks in PRMs and explains multi-model grader setups.30:27–35:17 · Guest disagreement 2/10 Cost and Complexity of Scaling RLVR vs. Pre-Training The host suggests RL scaling is notoriously finicky and complex based on the multi-model architecture. The guest counters that it is not overly complicated, citing his own 39-page Jupyter notebook implementation and cost differences.35:17–38:40 · Guest disagreement 1/10 Meta-Lessons on AI Progress and Infrastructure Specialization The host asks about the meta-lessons from 2025 AI progress. The guest explains that progress is an aggregation of small engineering micro-optimizations across specialized sub-teams rather than a single magic bullet.38:40–43:11 · Guest disagreement 2/10 "Bench-Maxing", Leaderboard Gaming, and Evaluation Reality The host brings up 'bench-maxing' and asks if corporate economic motives drive leaderboard gaming. The guest gives a nuanced breakdown of how evaluation style bias distorts leaderboards.43:11–46:46 · Guest disagreement 1/10 Inference-Time Scaling and Recursive Prompt Chunking The host steers the discussion toward non-post-training factors like inference scaling. The guest explains parallel sampling, vote aggregation, and prompt chunking.46:46–50:11 · Guest disagreement 1/10 System Engineering, Tool Calling, and Local Execution Security The host listens as the guest explains how tool calling and prompt wrappers create a gap between raw open-weight models and cloud LLM platforms.50:11–55:13 · Guest disagreement 2/10 Generalist Model Parity and In-House Enterprise LLM Training The host links generalist model parity with private enterprise data as a moat, asking if companies are resuming in-house training. The guest clarifies that large enterprises are indeed building data-center scale models in-house.55:13–59:25 · Guest disagreement 1/10 Evaluating Continual Learning and Iterative Model Updates The host brings up NeurIPS buzz around continual learning and asks about the guest's skeptical timeline. The guest explains why catastrophic forgetting and hosted API paradigms delay implementation.59:25–1:04:55 · Guest disagreement 1/10 Sebastian Raschka's Code-First Research and Writing Methodology The host asks the guest about his research methodology, book authoring, and use of LLMs. The guest describes his code-first approach of implementing architectures from scratch to verify paper claims.0:59–4:05 · Matt pushing back 1/10 Evaluating the Transformer Architecture and Modern Alternatives The host opens with an informed question about whether the transformer architecture's days are numbered given developments in SSMs and diffusion models. The guest explains that transformers remain state-of-the-art while modern variants mostly focus on cost efficiency.4:05–9:46 · Matt pushing back 1/10 World Models and Internal State Prediction for Code Generation The host asks targeted questions about world models and small recursive models. The guest educates the host on internal state prediction for code LLMs and how recursive models target specific logic benchmarks like ARC.9:46–13:24 · Matt pushing back 1/10 Text Diffusion Models versus Auto-Regressive Transformers The host demonstrates strong knowledge by bringing up DeepMind's Gemini Diffusion announcement. The guest explains the architectural differences between sequential auto-regressive generation and parallel diffusion denoising.13:24–18:03 · Matt pushing back 2/10 LLM Architectural Optimizations, MoE Adoption, and Pre-Training Limits The host pushes the guest on whether architecture progress is just superficial tuning and checks if pre-training is dead. The guest re-frames pre-training as 'boring' rather than dead and highlights MoE adoption trends.18:03–24:42 · Matt pushing back 1/10 Post-Training Evolution: RLVR, GRPO, and Reasoning Unlocking The host cites the guest's technical blog post to prompt an explanation of RLVR and GRPO. The guest provides an in-depth breakdown of how GRPO eliminates memory overhead by getting rid of separate reward and value models.24:42–30:21 · Matt pushing back 1/10 Process Reward Models (PRMs) and Verifiable Rewards Beyond Code The host prompts the guest regarding why Process Reward Models (PRMs) haven't been successful yet. The guest outlines reward hacking risks in PRMs and explains multi-model grader setups.30:27–35:17 · Matt pushing back 2/10 Cost and Complexity of Scaling RLVR vs. Pre-Training The host suggests RL scaling is notoriously finicky and complex based on the multi-model architecture. The guest counters that it is not overly complicated, citing his own 39-page Jupyter notebook implementation and cost differences.35:17–38:40 · Matt pushing back 1/10 Meta-Lessons on AI Progress and Infrastructure Specialization The host asks about the meta-lessons from 2025 AI progress. The guest explains that progress is an aggregation of small engineering micro-optimizations across specialized sub-teams rather than a single magic bullet.38:40–43:11 · Matt pushing back 2/10 "Bench-Maxing", Leaderboard Gaming, and Evaluation Reality The host brings up 'bench-maxing' and asks if corporate economic motives drive leaderboard gaming. The guest gives a nuanced breakdown of how evaluation style bias distorts leaderboards.43:11–46:46 · Matt pushing back 1/10 Inference-Time Scaling and Recursive Prompt Chunking The host steers the discussion toward non-post-training factors like inference scaling. The guest explains parallel sampling, vote aggregation, and prompt chunking.46:46–50:11 · Matt pushing back 1/10 System Engineering, Tool Calling, and Local Execution Security The host listens as the guest explains how tool calling and prompt wrappers create a gap between raw open-weight models and cloud LLM platforms.50:11–55:13 · Matt pushing back 2/10 Generalist Model Parity and In-House Enterprise LLM Training The host links generalist model parity with private enterprise data as a moat, asking if companies are resuming in-house training. The guest clarifies that large enterprises are indeed building data-center scale models in-house.55:13–59:25 · Matt pushing back 1/10 Evaluating Continual Learning and Iterative Model Updates The host brings up NeurIPS buzz around continual learning and asks about the guest's skeptical timeline. The guest explains why catastrophic forgetting and hosted API paradigms delay implementation.59:25–1:04:55 · Matt pushing back 1/10 Sebastian Raschka's Code-First Research and Writing Methodology The host asks the guest about his research methodology, book authoring, and use of LLMs. The guest describes his code-first approach of implementing architectures from scratch to verify paper claims.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 40.3% · guest 59.7%0:00 · Matt 40.3% · guest 59.7%3:00 · Matt 1.5% · guest 98.5%3:00 · Matt 1.5% · guest 98.5%6:00 · Matt 3.6% · guest 96.4%6:00 · Matt 3.6% · guest 96.4%9:00 · Matt 8.3% · guest 91.7%9:00 · Matt 8.3% · guest 91.7%12:00 · Matt 14.6% · guest 85.4%12:00 · Matt 14.6% · guest 85.4%15:00 · Matt 2.7% · guest 97.3%15:00 · Matt 2.7% · guest 97.3%18:00 · Matt 23.8% · guest 76.2%18:00 · Matt 23.8% · guest 76.2%21:00 · Matt 0% · guest 100%21:00 · Matt 0% · guest 100%24:00 · Matt 9.4% · guest 90.6%24:00 · Matt 9.4% · guest 90.6%27:00 · Matt 6.5% · guest 93.5%27:00 · Matt 6.5% · guest 93.5%30:00 · Matt 13.5% · guest 86.5%30:00 · Matt 13.5% · guest 86.5%33:00 · Matt 12.4% · guest 87.6%33:00 · Matt 12.4% · guest 87.6%36:00 · Matt 7.3% · guest 92.7%36:00 · Matt 7.3% · guest 92.7%39:00 · Matt 6.6% · guest 93.4%39:00 · Matt 6.6% · guest 93.4%42:00 · Matt 13.9% · guest 86.1%42:00 · Matt 13.9% · guest 86.1%45:00 · Matt 0% · guest 100%45:00 · Matt 0% · guest 100%48:00 · Matt 6.1% · guest 93.9%48:00 · Matt 6.1% · guest 93.9%51:00 · Matt 8.8% · guest 91.2%51:00 · Matt 8.8% · guest 91.2%54:00 · Matt 15.7% · guest 84.3%54:00 · Matt 15.7% · guest 84.3%57:00 · Matt 24% · guest 76%57:00 · Matt 24% · guest 76%1:00:00 · Matt 6.7% · guest 93.3%1:00:00 · Matt 6.7% · guest 93.3%1:03:00 · Matt 3.6% · guest 96.4%1:03:00 · Matt 3.6% · guest 96.4%1:06:00 · Matt 33.9% · guest 66.1%1:06:00 · Matt 33.9% · guest 66.1%
Sharpest disagreement ▶ 30:46 Guest pushes back on RL complexity premise

When the host implies RLVR is finicky and complex to scale, the guest directly counters that it is not super complicated, noting he implemented GRPO in a 39-page notebook and that it is ten times cheaper than pre-training.

Hardest push from Matt ▶ 17:26 Host challenges guest on pre-training viability

The host explicitly interrupts the guest's focus on post-training to press whether he truly believes pre-training is dead or if there is still room for progress.

Biggest teaching moment ▶ 22:00 Guest explains GRPO model memory reductions

The guest breaks down the exact technical difference between PPO and GRPO, explaining how replacing separate reward and value models reduces memory overhead from three model copies down to one.

Matt holds his own ▶ 50:11 Host synthesizes private data moats in enterprise LLMs

The host demonstrates deep domain insight by tying together generalist model parity, private data moats, and the strategic rationale for enterprise in-house training.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Evaluating the Transformer Architecture and Modern Alternatives 5411 The host opens with an informed question about whether the transformer architecture's days are numbered given developments in SSMs and diffusion models. The guest explains that transformers remain state-of-the-art while modern variants mostly focus on cost efficiency.
World Models and Internal State Prediction for Code Generation 4511 The host asks targeted questions about world models and small recursive models. The guest educates the host on internal state prediction for code LLMs and how recursive models target specific logic benchmarks like ARC.
Text Diffusion Models versus Auto-Regressive Transformers 5511 The host demonstrates strong knowledge by bringing up DeepMind's Gemini Diffusion announcement. The guest explains the architectural differences between sequential auto-regressive generation and parallel diffusion denoising.
LLM Architectural Optimizations, MoE Adoption, and Pre-Training Limits 5522 The host pushes the guest on whether architecture progress is just superficial tuning and checks if pre-training is dead. The guest re-frames pre-training as 'boring' rather than dead and highlights MoE adoption trends.
Post-Training Evolution: RLVR, GRPO, and Reasoning Unlocking 5611 The host cites the guest's technical blog post to prompt an explanation of RLVR and GRPO. The guest provides an in-depth breakdown of how GRPO eliminates memory overhead by getting rid of separate reward and value models.
Process Reward Models (PRMs) and Verifiable Rewards Beyond Code 5511 The host prompts the guest regarding why Process Reward Models (PRMs) haven't been successful yet. The guest outlines reward hacking risks in PRMs and explains multi-model grader setups.
Cost and Complexity of Scaling RLVR vs. Pre-Training 4622 The host suggests RL scaling is notoriously finicky and complex based on the multi-model architecture. The guest counters that it is not overly complicated, citing his own 39-page Jupyter notebook implementation and cost differences.
Meta-Lessons on AI Progress and Infrastructure Specialization 4411 The host asks about the meta-lessons from 2025 AI progress. The guest explains that progress is an aggregation of small engineering micro-optimizations across specialized sub-teams rather than a single magic bullet.
"Bench-Maxing", Leaderboard Gaming, and Evaluation Reality 4522 The host brings up 'bench-maxing' and asks if corporate economic motives drive leaderboard gaming. The guest gives a nuanced breakdown of how evaluation style bias distorts leaderboards.
Inference-Time Scaling and Recursive Prompt Chunking 4511 The host steers the discussion toward non-post-training factors like inference scaling. The guest explains parallel sampling, vote aggregation, and prompt chunking.
System Engineering, Tool Calling, and Local Execution Security 3511 The host listens as the guest explains how tool calling and prompt wrappers create a gap between raw open-weight models and cloud LLM platforms.
Generalist Model Parity and In-House Enterprise LLM Training 5522 The host links generalist model parity with private enterprise data as a moat, asking if companies are resuming in-house training. The guest clarifies that large enterprises are indeed building data-center scale models in-house.
Evaluating Continual Learning and Iterative Model Updates 4511 The host brings up NeurIPS buzz around continual learning and asks about the guest's skeptical timeline. The guest explains why catastrophic forgetting and hosted API paradigms delay implementation.
Sebastian Raschka's Code-First Research and Writing Methodology 3411 The host asks the guest about his research methodology, book authoring, and use of LLMs. The guest describes his code-first approach of implementing architectures from scratch to verify paper claims.

Statements from this episode (25)

Assertion Supported
Transformer remains state of the art for LLM performance
“I would say right now, yes, because it's still the state of the art. So there is nothing really better in terms of state of the art performance, getting better quality results.”
Sebastian Raschka Jan 29, 2026 ▶ 2:10
Assertion Not checkable as stated
DeepSeek's architecture is still built on a GPT-2 scaffold
“And you can actually, in fact, Take a GPT one or two model and with a few, I mean, few lines of code almost, you can transform it into the latest let's say deep seek version, 3.2 architecture. It's not like a big leap. It's still the same as scaffold.”
Sebastian Raschka Jan 29, 2026 ▶ 2:54
Prediction Not checkable as stated
Predicting internal states might push state-of-the-art for code LLMs
“And I do think that is something that is maybe more expensive to do, but it is also something that might push the state of the art a little bit.”
Sebastian Raschka Jan 29, 2026 ▶ 5:53
Assertion Supported
Hierarchical Reasoning Model matches larger LLMs on ARC using Transformer architecture
“Hierarchical reasoning model, it became like popular because it performed relatively well on that benchmark compared to very expensive models like Gemini, Chachupiti, and so forth. And it is a transformer architecture.”
Sebastian Raschka Jan 29, 2026 ▶ 7:10
Insight
Enterprises should transition from generalist LLMs to cheaper task-specific models
“If you have a business problem, you are maybe manufacturing something, maybe you can start with a generalist model, but then once you know exactly what the task is and you want to hone in on it, maybe it makes sense to replace that expensive thing by something…”
Sebastian Raschka Jan 29, 2026 ▶ 9:18
Opinion
Text diffusion models will not replace autoregressive Transformers at state-of-the-art
“So it is a interesting direction to go into these diffusion, diffusion models as alternative to the auto regressive transformers, but it is not I would say the replacement at the state of the art.”
Sebastian Raschka Jan 29, 2026 ▶ 12:57
Prediction Held up
Major AI firm will launch a frontier text diffusion model in 2026
“I think one company will launch a big diffusion model this year.”
Sebastian Raschka Jan 29, 2026 ▶ 13:07
Prediction Not checkable as stated
Future LLMs will prioritize architectural efficiency over larger model sizes
“I wouldn't expect bigger architectures. I would expect a more efficient architectures tweaks getting, The same modeling performance for less compute”
Sebastian Raschka Jan 29, 2026 ▶ 17:09
Insight
Pre-training is no longer where the low-hanging AI gains lie
“Pre-training is not dead, but pre-training is boring. So it's not where the low hanging fruit is anymore.”
Sebastian Raschka Jan 29, 2026 ▶ 17:32
Assertion Partly supported
50 RLVR steps boosted Qwen 3 MATH-500 score from 15% to 50%
“I took the Quinn three model as part of my book, the reasoning from scratch book. And I trained it just for 50 steps with RLVR, and it goes from 15%, so one five percent accuracy on math 500 to 50% on math 500.”
Sebastian Raschka Jan 29, 2026 ▶ 23:58
Insight
RLVR unlocks pre-training knowledge rather than teaching LLMs new math
“The knowledge is already there in the pre-training, and this just unlocks it. It's just like a step that maybe shows the model how to use its own knowledge, basically.”
Sebastian Raschka Jan 29, 2026 ▶ 24:33
Prediction Not checkable as stated
Process Reward Models will eventually become standard in LLM post-training
“I think it is promising and we will see it working at some point. I think it's just like right now it's still Tricky to make it work, but I am quite sure we'll see it as part of the standard repertoire at some point.”
Sebastian Raschka Jan 29, 2026 ▶ 26:39
Insight
Bigger LLM gains will come from multi-model process refinement, not scaling
“That's where you make the bigger gains rather than scaling the model size. I think that's one of those things where you will see more progress coming from.”
Sebastian Raschka Jan 29, 2026 ▶ 28:04
Assertion Supported
DeepSeek-R1 cost roughly $300,000 to train, 10x cheaper than DeepSeek-V3
“Deep seek version three, they had like a five million dollar price tag on that given the, I think two dollars per GPU, they assume whether that's a correct assumption or not it's a different question, but if you compare it relative to the cost of R one, I thin…”
Sebastian Raschka Jan 29, 2026 ▶ 31:06
Insight
Original GRPO algorithm is flaky but stabilizes with practical engineering tricks
“Vanilla GRP or the original algorithm, it is pretty flaky. Like where it is, you have to babysit it. Over the course of the year, many people had these tips and tricks where some people were saying, remove the KL divergence term. Like if you just drop it for m…”
Sebastian Raschka Jan 29, 2026 ▶ 32:28
Prediction Not checkable as stated
RLVR and GRPO will stabilize into canonical algorithms similar to AdamW
“It's similar to, you know, optimizers with Adam. So Adam is, I mean, right now there's Adam W. There was SGD and all the other RMS prop and how they were called. And they kind of all converge to Adam W by adding more and more tricks. And I think that's the sam…”
Sebastian Raschka Jan 29, 2026 ▶ 34:51
Assertion Not checkable as stated
The AI industry is rapidly running out of challenging evaluation benchmarks
“The only thing is we are running out of is really benchmarks. So the improvement on benchmarks, it's kind of like harder to measure.”
Sebastian Raschka Jan 29, 2026 ▶ 37:56
Prediction Not checkable as stated
AI evaluation will shift from single-shot answers to multi-step agentic execution
“Maybe it's not the one shot problem anymore where it's not really answering knowledge question. That's not really solving math problems in, in one iteration of the benchmark. It is maybe more like the agentic cycle, like where you have like a more like a objec…”
Sebastian Raschka Jan 29, 2026 ▶ 38:04
Insight
Current AI model leaderboards reward response style over actual factual correctness
“It rewards the style more, more than the correctness because there is no correctness check.”
Sebastian Raschka Jan 29, 2026 ▶ 40:08
Insight
Tool calling reduces LLM hallucinations by outsourcing memory retrieval tasks
“And that is very, very powerful because I think this is one of the Ways you can mitigate not totally mitigate, but let's say reduce hallucinations because then the LM suddenly doesn't have to remember everything anymore.”
Sebastian Raschka Jan 29, 2026 ▶ 48:09
Assertion Supported
GPT OSS benchmarks demonstrate 1.2x capability jump when tool calling is enabled
“And also you can actually go to the GPT OSS release block, and they did have benchmarks to show how the performance on the benchmarks is with the same model with tool called enabled and disabled. And you can actually see there is, I mean, it's not like two tim…”
Sebastian Raschka Jan 29, 2026 ▶ 49:36
Opinion
Leading generalist models like ChatGPT, Gemini, and Claude show functional parity
“Like if you use or compare ChatGPT, Gemini Claude, Grock. I think they are all pretty much on the same level. Like, and I think that's because they're trying to do everything. Like the generalist models for a general person to do a lot of things. I mean, Claud…”
Sebastian Raschka Jan 29, 2026 ▶ 50:24
Assertion Supported
Large enterprises are secretly hiring teams to train ChatGPT-scale LLMs in-house
“I know for a fact that big companies are training now LLMs in-house. Really, like, big companies who have the financial means to train chat to be like model are hiring people who train LLMs.”
Sebastian Raschka Jan 29, 2026 ▶ 51:54
Prediction Not checkable as stated
Self-improving AI and continual learning agents will not be feasible in 2026
“If you have an LLM that self improves or like an agent that does something fails and learns, I don't think anything like that is feasible this year.”
Sebastian Raschka Jan 29, 2026 ▶ 56:10
Disclosure
Raschka does not use LLMs to write his books or blog posts
“For blog writing on book writing, not so much because honestly, I, for fun, I tried it out. It's just, it generates okay text, but it's, I don't know, it does not I can ask it to generate text like me, but it's almost like, then I don't like it and I end up ed…”
Sebastian Raschka Jan 29, 2026 ▶ 1:04:34
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.