AI researcher Sebastian Raschka compares the training cost of DeepSeek's base model (V3) against its reasoning model (R1) to illustrate the cost-effectiveness of RL post-training.
“Deep seek version three, they had like a five million dollar price tag on that given the, I think two dollars per GPU, they assume whether that's a correct assumption or not it's a different question, but if you compare it relative to the cost of R one, I think R one was about 300,000 dollars when they trained it, they had a number in the nature version of the paper. So it's basically more than 10 times cheaper than pre-training.”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Sebastian Raschka
Opinion
Text diffusion models will not replace autoregressive Transformers at state-of-the-art
“So it is a interesting direction to go into these diffusion, diffusion models as alternative to the auto regressive transformers, but it is not I would say the replacement at the state of the art.”
Sebastian RaschkaJan 29, 2026▶ 12:57State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
PredictionNot checkable as stated
Future LLMs will prioritize architectural efficiency over larger model sizes
“I wouldn't expect bigger architectures. I would expect a more efficient architectures tweaks getting, The same modeling performance for less compute”
Sebastian RaschkaJan 29, 2026▶ 17:09State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
Insight
Pre-training is no longer where the low-hanging AI gains lie
“Pre-training is not dead, but pre-training is boring. So it's not where the low hanging fruit is anymore.”
Sebastian RaschkaJan 29, 2026▶ 17:32State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
Insight
RLVR unlocks pre-training knowledge rather than teaching LLMs new math
“The knowledge is already there in the pre-training, and this just unlocks it. It's just like a step that maybe shows the model how to use its own knowledge, basically.”
Sebastian RaschkaJan 29, 2026▶ 24:33State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
PredictionNot checkable as stated
Process Reward Models will eventually become standard in LLM post-training
“I think it is promising and we will see it working at some point. I think it's just like right now it's still Tricky to make it work, but I am quite sure we'll see it as part of the standard repertoire at some point.”
Sebastian RaschkaJan 29, 2026▶ 26:39State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
Insight
Bigger LLM gains will come from multi-model process refinement, not scaling
“That's where you make the bigger gains rather than scaling the model size. I think that's one of those things where you will see more progress coming from.”
Sebastian RaschkaJan 29, 2026▶ 28:04State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
Made with StarZero
Turn any episode into a week of clips.
This entire site, over 400 conversations transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.