Sebastian Raschka

1 appearance on the record.

computed by AI from the episodes · how this works → · full disclaimer →

25statements → 15claims → 7claims resolved → 86%fully supported → 3.64/5average certainty → 2.36/5average debate potential → 4.1/5argument clarity · the sources →

6 supported 1 partly supported 0 contradicted 8 not checkable as stated how the 15 claims stand · each chip opens the sources

7 predictions · 8 assertions · 2 opinions · 7 insights · 1 disclosure · every statement was checked. The predictions and assertions are the 15 claims: statements the public record can support or contradict. 7 are resolved, and 8 name no date, number or outcome precise enough to check. Everything else (opinions, insights, what ifs, disclosures) can never be settled by the record, so it carries no assessment.

The record, in short

What the tape says about how Sebastian argues and how the claims held up. Everything they said, and everything said about them, is in the tabs below.

Their most notable supported claim

Assertion Supported
Large enterprises are secretly hiring teams to train ChatGPT-scale LLMs in-house
“I know for a fact that big companies are training now LLMs in-house. Really, like, big companies who have the financial means to train chat to be like model are hiring people who train LLMs.”
Sebastian Raschka Jan 29, 2026 ▶ 51:54 State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka

Expressed certainty vs assessment result

none yet certainty 1
none yet certainty 2
100% certainty 3
100% certainty 4
75% certainty 5

weighted support: a fully supported claim counts one, a partly supported claim counts half. Each filled bar is clickable and opens exactly those claims; "none yet" means nothing said at that certainty level has resolved yet

Argument clarity: do they answer the question? how? →

4.1 / 5 directness 4.3 · coherence 4.2 · precision 3.9 · compression 3.6

redirected or did not address 1 of 13 assessed questions (8%). Watch them ▸

This is a score against a rubric. It is not a rank. Every host question → answer exchange is scored with names hidden on directness, coherence, precision and compression, 1–5 each, on meaning alone: disfluencies are ignored, and only raw unedited episodes count. This is the score that measures thought. Every scored exchange, scores shown → · The rubric and its checks →

How they sound: speaking style how? →

259 words/min while actually speaking · 24.1 um and uh per 1k words

Measured by listening to the audio itself: 12,098 words across 1 episode of raw-level tape, transcribed verbatim with every um and uh kept, each one attributed only where the alignment onto our timed stream is unambiguous. These are measurements of speaking style. We do not rank them: across this corpus, fluency and argument quality are nearly uncorrelated (ρ≈0.2), and smooth talking does not signal clear thinking. How it's measured →

Everything Sebastian Raschka said on the MAD Podcast that made the record, most notable first. Filter by type, assessment or year in the ledger →

Opinion
Text diffusion models will not replace autoregressive Transformers at state-of-the-art
“So it is a interesting direction to go into these diffusion, diffusion models as alternative to the auto regressive transformers, but it is not I would say the replacement at the state of the art.”
Sebastian Raschka Jan 29, 2026 ▶ 12:57 State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
Prediction Not checkable as stated
Future LLMs will prioritize architectural efficiency over larger model sizes
“I wouldn't expect bigger architectures. I would expect a more efficient architectures tweaks getting, The same modeling performance for less compute”
Sebastian Raschka Jan 29, 2026 ▶ 17:09 State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
Insight
Pre-training is no longer where the low-hanging AI gains lie
“Pre-training is not dead, but pre-training is boring. So it's not where the low hanging fruit is anymore.”
Sebastian Raschka Jan 29, 2026 ▶ 17:32 State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
Insight
RLVR unlocks pre-training knowledge rather than teaching LLMs new math
“The knowledge is already there in the pre-training, and this just unlocks it. It's just like a step that maybe shows the model how to use its own knowledge, basically.”
Sebastian Raschka Jan 29, 2026 ▶ 24:33 State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
Prediction Not checkable as stated
Process Reward Models will eventually become standard in LLM post-training
“I think it is promising and we will see it working at some point. I think it's just like right now it's still Tricky to make it work, but I am quite sure we'll see it as part of the standard repertoire at some point.”
Sebastian Raschka Jan 29, 2026 ▶ 26:39 State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
Insight
Bigger LLM gains will come from multi-model process refinement, not scaling
“That's where you make the bigger gains rather than scaling the model size. I think that's one of those things where you will see more progress coming from.”
Sebastian Raschka Jan 29, 2026 ▶ 28:04 State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
Assertion Supported
Large enterprises are secretly hiring teams to train ChatGPT-scale LLMs in-house
“I know for a fact that big companies are training now LLMs in-house. Really, like, big companies who have the financial means to train chat to be like model are hiring people who train LLMs.”
Sebastian Raschka Jan 29, 2026 ▶ 51:54 State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
Assertion Supported
Transformer remains state of the art for LLM performance
“I would say right now, yes, because it's still the state of the art. So there is nothing really better in terms of state of the art performance, getting better quality results.”
Sebastian Raschka Jan 29, 2026 ▶ 2:10 State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
Assertion Not checkable as stated
DeepSeek's architecture is still built on a GPT-2 scaffold
“And you can actually, in fact, Take a GPT one or two model and with a few, I mean, few lines of code almost, you can transform it into the latest let's say deep seek version, 3.2 architecture. It's not like a big leap. It's still the same as scaffold.”
Sebastian Raschka Jan 29, 2026 ▶ 2:54 State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
Insight
Enterprises should transition from generalist LLMs to cheaper task-specific models
“If you have a business problem, you are maybe manufacturing something, maybe you can start with a generalist model, but then once you know exactly what the task is and you want to hone in on it, maybe it makes sense to replace that expensive thing by something…”
Sebastian Raschka Jan 29, 2026 ▶ 9:18 State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
Prediction Held up
Major AI firm will launch a frontier text diffusion model in 2026
“I think one company will launch a big diffusion model this year.”
Sebastian Raschka Jan 29, 2026 ▶ 13:07 State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
Assertion Partly supported
50 RLVR steps boosted Qwen 3 MATH-500 score from 15% to 50%
“I took the Quinn three model as part of my book, the reasoning from scratch book. And I trained it just for 50 steps with RLVR, and it goes from 15%, so one five percent accuracy on math 500 to 50% on math 500.”
Sebastian Raschka Jan 29, 2026 ▶ 23:58 State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
Assertion Supported
DeepSeek-R1 cost roughly $300,000 to train, 10x cheaper than DeepSeek-V3
“Deep seek version three, they had like a five million dollar price tag on that given the, I think two dollars per GPU, they assume whether that's a correct assumption or not it's a different question, but if you compare it relative to the cost of R one, I thin…”
Sebastian Raschka Jan 29, 2026 ▶ 31:06 State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
Prediction Not checkable as stated
AI evaluation will shift from single-shot answers to multi-step agentic execution
“Maybe it's not the one shot problem anymore where it's not really answering knowledge question. That's not really solving math problems in, in one iteration of the benchmark. It is maybe more like the agentic cycle, like where you have like a more like a objec…”
Sebastian Raschka Jan 29, 2026 ▶ 38:04 State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
Insight
Current AI model leaderboards reward response style over actual factual correctness
“It rewards the style more, more than the correctness because there is no correctness check.”
Sebastian Raschka Jan 29, 2026 ▶ 40:08 State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
Opinion
Leading generalist models like ChatGPT, Gemini, and Claude show functional parity
“Like if you use or compare ChatGPT, Gemini Claude, Grock. I think they are all pretty much on the same level. Like, and I think that's because they're trying to do everything. Like the generalist models for a general person to do a lot of things. I mean, Claud…”
Sebastian Raschka Jan 29, 2026 ▶ 50:24 State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
Prediction Not checkable as stated
Self-improving AI and continual learning agents will not be feasible in 2026
“If you have an LLM that self improves or like an agent that does something fails and learns, I don't think anything like that is feasible this year.”
Sebastian Raschka Jan 29, 2026 ▶ 56:10 State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
Disclosure
Raschka does not use LLMs to write his books or blog posts
“For blog writing on book writing, not so much because honestly, I, for fun, I tried it out. It's just, it generates okay text, but it's, I don't know, it does not I can ask it to generate text like me, but it's almost like, then I don't like it and I end up ed…”
Sebastian Raschka Jan 29, 2026 ▶ 1:04:34 State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
Prediction Not checkable as stated
Predicting internal states might push state-of-the-art for code LLMs
“And I do think that is something that is maybe more expensive to do, but it is also something that might push the state of the art a little bit.”
Sebastian Raschka Jan 29, 2026 ▶ 5:53 State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
Insight
Original GRPO algorithm is flaky but stabilizes with practical engineering tricks
“Vanilla GRP or the original algorithm, it is pretty flaky. Like where it is, you have to babysit it. Over the course of the year, many people had these tips and tricks where some people were saying, remove the KL divergence term. Like if you just drop it for m…”
Sebastian Raschka Jan 29, 2026 ▶ 32:28 State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
Prediction Not checkable as stated
RLVR and GRPO will stabilize into canonical algorithms similar to AdamW
“It's similar to, you know, optimizers with Adam. So Adam is, I mean, right now there's Adam W. There was SGD and all the other RMS prop and how they were called. And they kind of all converge to Adam W by adding more and more tricks. And I think that's the sam…”
Sebastian Raschka Jan 29, 2026 ▶ 34:51 State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
Assertion Not checkable as stated
The AI industry is rapidly running out of challenging evaluation benchmarks
“The only thing is we are running out of is really benchmarks. So the improvement on benchmarks, it's kind of like harder to measure.”
Sebastian Raschka Jan 29, 2026 ▶ 37:56 State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
Insight
Tool calling reduces LLM hallucinations by outsourcing memory retrieval tasks
“And that is very, very powerful because I think this is one of the Ways you can mitigate not totally mitigate, but let's say reduce hallucinations because then the LM suddenly doesn't have to remember everything anymore.”
Sebastian Raschka Jan 29, 2026 ▶ 48:09 State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
Assertion Supported
GPT OSS benchmarks demonstrate 1.2x capability jump when tool calling is enabled
“And also you can actually go to the GPT OSS release block, and they did have benchmarks to show how the performance on the benchmarks is with the same model with tool called enabled and disabled. And you can actually see there is, I mean, it's not like two tim…”
Sebastian Raschka Jan 29, 2026 ▶ 49:36 State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka

Show 1statements(1 left)

Appearances (1)

EpisodeDateSpeaking time
State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka Jan 29, 2026 57m
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.