Assertion Supported AI assessment confidence: 95% certainty 4/5 debate potential 2/5

OLMo 3 injects thinking traces during mid-training rather than post-training

Pratyush Maini · ⚡️ Reverse Engineering OpenAI's Training Data — Pratyush Maini, Datology · Feb 10, 2026 · at 14:48

Datology founding team member Pratyush Maini discusses AI2's OLMo model training reports to examine where reasoning traces are introduced in model training pipelines.

0:00 / 0:20exact quote · 20.0s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“So, with the instruct variant, they do not have any thinking data in the post-training phase, but they do mention that we are going to put, like this is the line that they write, there is some intentional addition of thinking traces in the mid-training phase of ALMO data, and so that's the information available from their report, and we can also see the actual data sets they use.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Pratyush Maini

Assertion Open · timeframe Feb 2027
Modern LLMs verbatim regurgitate JEE exam questions from two-word prompts
“We consistently saw how many of these, like, models today are being, like, massively, like, kind of fine-tuned on problems from... Like, oversight? Massively worked with. Like, even, like, imagine if I ask you the light bulb, what comes next in your mind? It w…”
Pratyush Maini Feb 10, 2026 ▶ 3:36 ⚡️ Reverse Engineering OpenAI's Training Data — Pratyush Maini, Datology
Assertion Not checkable as stated
Labs train LLMs on benchmark questions for multiple epochs late in training
“It's very clear how the last stage of training for many of these models does have a massive amount of example or examine. Benchmaxing. Because the model will not, like, behaviorally complete exam questions with options if they have not really seen it at the en…”
Pratyush Maini Feb 10, 2026 ▶ 4:17 ⚡️ Reverse Engineering OpenAI's Training Data — Pratyush Maini, Datology
Assertion Not checkable as stated
Self-reflection training data is now core to all frontier foundation models
“What this suggests about the GPT training data is that the self-reflection data has now actually become pretty much core to the training of all frontier models, because we're seeing that happen in non-instruct models across the board.”
Pratyush Maini Feb 10, 2026 ▶ 15:26 ⚡️ Reverse Engineering OpenAI's Training Data — Pratyush Maini, Datology
Insight
Core AI capabilities must be built during pre-training, not just fine-tuned
“If there is a core capability that you actually care about, that capability should be part of the foundation and not a fine-tuned artifact.”
Pratyush Maini Feb 10, 2026 ▶ 18:53 ⚡️ Reverse Engineering OpenAI's Training Data — Pratyush Maini, Datology
Prediction Not checkable as stated
Enterprises will widely adopt specialized AI pre-training in 2026 and 2027
“So I think like, 26 and 27 are going to be the years where different enterprises start doing specialized pre-training, because the cost of pre-training really amortizes itself very fast.”
Pratyush Maini Feb 10, 2026 ▶ 19:27 ⚡️ Reverse Engineering OpenAI's Training Data — Pratyush Maini, Datology
Insight
Small specialized pre-trained models can match capabilities of larger fine-tuned models
“When you think of the fact that by doing specialized pre-training, you can train a smaller model, which is as capable as a much larger model when fine-tuned.”
Pratyush Maini Feb 10, 2026 ▶ 19:38 ⚡️ Reverse Engineering OpenAI's Training Data — Pratyush Maini, Datology
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.