The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

clear all ✕

why aren't all 15 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 0 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Prediction Not checkable as stated
Future LLMs will prioritize architectural efficiency over larger model sizes
“I wouldn't expect bigger architectures. I would expect a more efficient architectures tweaks getting, The same modeling performance for less compute”
Sebastian Raschka Jan 29, 2026 ▶ 17:09 State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
Prediction Not checkable as stated
Process Reward Models will eventually become standard in LLM post-training
“I think it is promising and we will see it working at some point. I think it's just like right now it's still Tricky to make it work, but I am quite sure we'll see it as part of the standard repertoire at some point.”
Sebastian Raschka Jan 29, 2026 ▶ 26:39 State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
Assertion Supported
Large enterprises are secretly hiring teams to train ChatGPT-scale LLMs in-house
“I know for a fact that big companies are training now LLMs in-house. Really, like, big companies who have the financial means to train chat to be like model are hiring people who train LLMs.”
Sebastian Raschka Jan 29, 2026 ▶ 51:54 State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
Assertion Supported
Transformer remains state of the art for LLM performance
“I would say right now, yes, because it's still the state of the art. So there is nothing really better in terms of state of the art performance, getting better quality results.”
Sebastian Raschka Jan 29, 2026 ▶ 2:10 State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
Assertion Not checkable as stated
DeepSeek's architecture is still built on a GPT-2 scaffold
“And you can actually, in fact, Take a GPT one or two model and with a few, I mean, few lines of code almost, you can transform it into the latest let's say deep seek version, 3.2 architecture. It's not like a big leap. It's still the same as scaffold.”
Sebastian Raschka Jan 29, 2026 ▶ 2:54 State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
Prediction Held up
Major AI firm will launch a frontier text diffusion model in 2026
“I think one company will launch a big diffusion model this year.”
Sebastian Raschka Jan 29, 2026 ▶ 13:07 State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
Assertion Partly supported
50 RLVR steps boosted Qwen 3 MATH-500 score from 15% to 50%
“I took the Quinn three model as part of my book, the reasoning from scratch book. And I trained it just for 50 steps with RLVR, and it goes from 15%, so one five percent accuracy on math 500 to 50% on math 500.”
Sebastian Raschka Jan 29, 2026 ▶ 23:58 State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
Assertion Supported
DeepSeek-R1 cost roughly $300,000 to train, 10x cheaper than DeepSeek-V3
“Deep seek version three, they had like a five million dollar price tag on that given the, I think two dollars per GPU, they assume whether that's a correct assumption or not it's a different question, but if you compare it relative to the cost of R one, I thin…”
Sebastian Raschka Jan 29, 2026 ▶ 31:06 State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
Prediction Not checkable as stated
AI evaluation will shift from single-shot answers to multi-step agentic execution
“Maybe it's not the one shot problem anymore where it's not really answering knowledge question. That's not really solving math problems in, in one iteration of the benchmark. It is maybe more like the agentic cycle, like where you have like a more like a objec…”
Sebastian Raschka Jan 29, 2026 ▶ 38:04 State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
Prediction Not checkable as stated
Self-improving AI and continual learning agents will not be feasible in 2026
“If you have an LLM that self improves or like an agent that does something fails and learns, I don't think anything like that is feasible this year.”
Sebastian Raschka Jan 29, 2026 ▶ 56:10 State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
Prediction Not checkable as stated
Predicting internal states might push state-of-the-art for code LLMs
“And I do think that is something that is maybe more expensive to do, but it is also something that might push the state of the art a little bit.”
Sebastian Raschka Jan 29, 2026 ▶ 5:53 State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
Prediction Not checkable as stated
RLVR and GRPO will stabilize into canonical algorithms similar to AdamW
“It's similar to, you know, optimizers with Adam. So Adam is, I mean, right now there's Adam W. There was SGD and all the other RMS prop and how they were called. And they kind of all converge to Adam W by adding more and more tricks. And I think that's the sam…”
Sebastian Raschka Jan 29, 2026 ▶ 34:51 State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
Assertion Not checkable as stated
The AI industry is rapidly running out of challenging evaluation benchmarks
“The only thing is we are running out of is really benchmarks. So the improvement on benchmarks, it's kind of like harder to measure.”
Sebastian Raschka Jan 29, 2026 ▶ 37:56 State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
Assertion Supported
GPT OSS benchmarks demonstrate 1.2x capability jump when tool calling is enabled
“And also you can actually go to the GPT OSS release block, and they did have benchmarks to show how the performance on the benchmarks is with the same model with tool called enabled and disabled. And you can actually see there is, I mean, it's not like two tim…”
Sebastian Raschka Jan 29, 2026 ▶ 49:36 State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
Assertion Supported
Hierarchical Reasoning Model matches larger LLMs on ARC using Transformer architecture
“Hierarchical reasoning model, it became like popular because it performed relatively well on that benchmark compared to very expensive models like Gemini, Chachupiti, and so forth. And it is a transformer architecture.”
Sebastian Raschka Jan 29, 2026 ▶ 7:10 State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.