Feb 10, 2026 · 27m · latent-space

⚡️ Reverse Engineering OpenAI's Training Data — Pratyush Maini, Datology

Pratyush Maini · 21m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Datology founding team member and CMU PhD student Pratyush Maini discusses data-centric AI, exploring how empirical dataset analysis, reverse-engineered reasoning traces, and specialized pre-training paradigms optimize frontier LLM performance. He highlights key findings from memorization anomalies and synthetic data scaling to demonstrate why data composition remains the fundamental driver of modern AI capability.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 3.6 Guest teaching 6.0 Guest disagreement 1.4 The hosts pushing back 2.2
05100:0010:0020:002:04–6:05 · The hosts as informed peer 5/10 PhD Thesis and Exam Question Memorization in LLMs Host pushes back on the relevance of JEE exam memorization and poses technical questions on model size ablation and MOE index routing. Guest explains how recent models and active parameter architectures exhibit precise memorization.6:05–12:51 · The hosts as informed peer 3/10 The Seahorse Emoji Anomaly and Emergent Self-Correction Guest walks through his investigation into the seahorse emoji anomaly across frontier models and how output token lengths spiked after reasoning models appeared. Host asks a clarifying technical question regarding release dates versus cutoff dates.12:51–18:16 · The hosts as informed peer 4/10 Tracing Reasoning Traces in Mid-Training via OLMo Guest shows how OLMo Trace confirms intentional mid-training additions of thinking traces rather than web data poisoning. Host questions whether the data reached models indirectly through the open web, which the guest rejects.18:16–20:19 · The hosts as informed peer 3/10 The Fine Tuner's Fallacy and Specialized Pre-Training Guest introduces The Fine Tuner's Fallacy and the necessity of specialized pre-training over post-hoc fine-tuning. Host jokingly notes the argument aligns directly with Datology's product pitch.20:19–25:42 · The hosts as informed peer 3/10 Scaling Synthetic Data: Beyond Web and Source Rephrasing Guest breaks down synthetic data paradigms, contrasting generator-driven generation with source rephrasing at trillion-token scale. Host reacts to the historical evolution from TinyStories to modern models.2:04–6:05 · Guest teaching 5/10 PhD Thesis and Exam Question Memorization in LLMs Host pushes back on the relevance of JEE exam memorization and poses technical questions on model size ablation and MOE index routing. Guest explains how recent models and active parameter architectures exhibit precise memorization.6:05–12:51 · Guest teaching 7/10 The Seahorse Emoji Anomaly and Emergent Self-Correction Guest walks through his investigation into the seahorse emoji anomaly across frontier models and how output token lengths spiked after reasoning models appeared. Host asks a clarifying technical question regarding release dates versus cutoff dates.12:51–18:16 · Guest teaching 6/10 Tracing Reasoning Traces in Mid-Training via OLMo Guest shows how OLMo Trace confirms intentional mid-training additions of thinking traces rather than web data poisoning. Host questions whether the data reached models indirectly through the open web, which the guest rejects.18:16–20:19 · Guest teaching 5/10 The Fine Tuner's Fallacy and Specialized Pre-Training Guest introduces The Fine Tuner's Fallacy and the necessity of specialized pre-training over post-hoc fine-tuning. Host jokingly notes the argument aligns directly with Datology's product pitch.20:19–25:42 · Guest teaching 7/10 Scaling Synthetic Data: Beyond Web and Source Rephrasing Guest breaks down synthetic data paradigms, contrasting generator-driven generation with source rephrasing at trillion-token scale. Host reacts to the historical evolution from TinyStories to modern models.2:04–6:05 · Guest disagreement 2/10 PhD Thesis and Exam Question Memorization in LLMs Host pushes back on the relevance of JEE exam memorization and poses technical questions on model size ablation and MOE index routing. Guest explains how recent models and active parameter architectures exhibit precise memorization.6:05–12:51 · Guest disagreement 1/10 The Seahorse Emoji Anomaly and Emergent Self-Correction Guest walks through his investigation into the seahorse emoji anomaly across frontier models and how output token lengths spiked after reasoning models appeared. Host asks a clarifying technical question regarding release dates versus cutoff dates.12:51–18:16 · Guest disagreement 2/10 Tracing Reasoning Traces in Mid-Training via OLMo Guest shows how OLMo Trace confirms intentional mid-training additions of thinking traces rather than web data poisoning. Host questions whether the data reached models indirectly through the open web, which the guest rejects.18:16–20:19 · Guest disagreement 1/10 The Fine Tuner's Fallacy and Specialized Pre-Training Guest introduces The Fine Tuner's Fallacy and the necessity of specialized pre-training over post-hoc fine-tuning. Host jokingly notes the argument aligns directly with Datology's product pitch.20:19–25:42 · Guest disagreement 1/10 Scaling Synthetic Data: Beyond Web and Source Rephrasing Guest breaks down synthetic data paradigms, contrasting generator-driven generation with source rephrasing at trillion-token scale. Host reacts to the historical evolution from TinyStories to modern models.2:04–6:05 · The hosts pushing back 3/10 PhD Thesis and Exam Question Memorization in LLMs Host pushes back on the relevance of JEE exam memorization and poses technical questions on model size ablation and MOE index routing. Guest explains how recent models and active parameter architectures exhibit precise memorization.6:05–12:51 · The hosts pushing back 2/10 The Seahorse Emoji Anomaly and Emergent Self-Correction Guest walks through his investigation into the seahorse emoji anomaly across frontier models and how output token lengths spiked after reasoning models appeared. Host asks a clarifying technical question regarding release dates versus cutoff dates.12:51–18:16 · The hosts pushing back 3/10 Tracing Reasoning Traces in Mid-Training via OLMo Guest shows how OLMo Trace confirms intentional mid-training additions of thinking traces rather than web data poisoning. Host questions whether the data reached models indirectly through the open web, which the guest rejects.18:16–20:19 · The hosts pushing back 2/10 The Fine Tuner's Fallacy and Specialized Pre-Training Guest introduces The Fine Tuner's Fallacy and the necessity of specialized pre-training over post-hoc fine-tuning. Host jokingly notes the argument aligns directly with Datology's product pitch.20:19–25:42 · The hosts pushing back 1/10 Scaling Synthetic Data: Beyond Web and Source Rephrasing Guest breaks down synthetic data paradigms, contrasting generator-driven generation with source rephrasing at trillion-token scale. Host reacts to the historical evolution from TinyStories to modern models.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 0%27:00 · the hosts 0% · guest 0%
Sharpest disagreement ▶ 17:01 Rejection of passive web ingestion theory

When the host suggests reasoning data might have passively entered models via web pre-training data, the guest immediately disagrees, asserting labs made deliberate additions to mid-training datasets.

Hardest push from the hosts ▶ 4:05 Challenging JEE benchmark significance

The host pushes back against the guest's observation of exam memorization by stating that nobody uses JEE as a reported benchmark.

Biggest teaching moment ▶ 6:10 Reverse engineering frontier reasoning inclusion

The guest explains in detail how the seahorse emoji prompt triggers an internal self-correction loop in frontier models, demonstrating how lab mid-training data practices can be reverse-engineered.

The host holds their own ▶ 5:50 MOE routing as index lookup

The host showcases technical understanding of mixture-of-experts architectures by explaining how router mechanisms function as an index to route requests to memorized expert weights.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
PhD Thesis and Exam Question Memorization in LLMs 5523 Host pushes back on the relevance of JEE exam memorization and poses technical questions on model size ablation and MOE index routing. Guest explains how recent models and active parameter architectures exhibit precise memorization.
The Seahorse Emoji Anomaly and Emergent Self-Correction 3712 Guest walks through his investigation into the seahorse emoji anomaly across frontier models and how output token lengths spiked after reasoning models appeared. Host asks a clarifying technical question regarding release dates versus cutoff dates.
Tracing Reasoning Traces in Mid-Training via OLMo 4623 Guest shows how OLMo Trace confirms intentional mid-training additions of thinking traces rather than web data poisoning. Host questions whether the data reached models indirectly through the open web, which the guest rejects.
The Fine Tuner's Fallacy and Specialized Pre-Training 3512 Guest introduces The Fine Tuner's Fallacy and the necessity of specialized pre-training over post-hoc fine-tuning. Host jokingly notes the argument aligns directly with Datology's product pitch.
Scaling Synthetic Data: Beyond Web and Source Rephrasing 3711 Guest breaks down synthetic data paradigms, contrasting generator-driven generation with source rephrasing at trillion-token scale. Host reacts to the historical evolution from TinyStories to modern models.

Statements from this episode (12)

Assertion Open · timeframe Feb 2027
Modern LLMs verbatim regurgitate JEE exam questions from two-word prompts
“We consistently saw how many of these, like, models today are being, like, massively, like, kind of fine-tuned on problems from... Like, oversight? Massively worked with. Like, even, like, imagine if I ask you the light bulb, what comes next in your mind? It w…”
Pratyush Maini Feb 10, 2026 ▶ 3:36
Assertion Not checkable as stated
Labs train LLMs on benchmark questions for multiple epochs late in training
“It's very clear how the last stage of training for many of these models does have a massive amount of example or examine. Benchmaxing. Because the model will not, like, behaviorally complete exam questions with options if they have not really seen it at the en…”
Pratyush Maini Feb 10, 2026 ▶ 4:17
Assertion Open · timeframe Feb 2027
Qwen 3 memorizes benchmark questions significantly more than Qwen 1.5 or 2
“I don't, we don't see this phenomenon like the earlier versions of like QN 1.5 or even QN two, but start seeing it in QN three. So there's something about a more, like a significantly higher weight on benchmark or like a J or like any times of evaluation quest…”
Pratyush Maini Feb 10, 2026 ▶ 4:59
Assertion Supported
MoE models with 20B active parameters match 72B dense models in memorization
“With the coming of the MOE models, we're starting to see these phenomena, even like models which have a less number of active parameters. So a phenomena that we would observe at a 72 B model in the past is probably now visible at a 20 B active one 20 B MOE als…”
Pratyush Maini Feb 10, 2026 ▶ 5:24
Assertion Supported
Non-reasoning Grok, GPT, and Gemini models exhibit recursive self-correction loops
“And so then I tried this across models, and I saw consistently across Grok, and GPT, and Gemini, that you were seeing this phenomena where models will, like, self-correct themselves quite a bit, and these were, like, non-thinking models. They were, like, the s…”
Pratyush Maini Feb 10, 2026 ▶ 7:08
Assertion Supported
Pre-December 2024 OpenAI models did not exhibit seahorse emoji self-correction loops
“And so I like ran the OpenAI API across like models released from 23 to 25, and you would see like all the models until twenty-twenty-four December had very Terce and short responses to the question. Is there a seahorse emoji? They would either say that there …”
Pratyush Maini Feb 10, 2026 ▶ 9:04
Assertion Supported
OLMo 3 injects thinking traces during mid-training rather than post-training
“So, with the instruct variant, they do not have any thinking data in the post-training phase, but they do mention that we are going to put, like this is the line that they write, there is some intentional addition of thinking traces in the mid-training phase o…”
Pratyush Maini Feb 10, 2026 ▶ 14:48
Assertion Not checkable as stated
Self-reflection training data is now core to all frontier foundation models
“What this suggests about the GPT training data is that the self-reflection data has now actually become pretty much core to the training of all frontier models, because we're seeing that happen in non-instruct models across the board.”
Pratyush Maini Feb 10, 2026 ▶ 15:26
Insight
Core AI capabilities must be built during pre-training, not just fine-tuned
“If there is a core capability that you actually care about, that capability should be part of the foundation and not a fine-tuned artifact.”
Pratyush Maini Feb 10, 2026 ▶ 18:53
Prediction Not checkable as stated
Enterprises will widely adopt specialized AI pre-training in 2026 and 2027
“So I think like, 26 and 27 are going to be the years where different enterprises start doing specialized pre-training, because the cost of pre-training really amortizes itself very fast.”
Pratyush Maini Feb 10, 2026 ▶ 19:27
Insight
Small specialized pre-trained models can match capabilities of larger fine-tuned models
“When you think of the fact that by doing specialized pre-training, you can train a smaller model, which is as capable as a much larger model when fine-tuned.”
Pratyush Maini Feb 10, 2026 ▶ 19:38
Assertion Partly supported
Datology BeyondWeb 3B matches NVIDIA Nemotron 8B in 2.7x less training time
“As you can see that we achieved the same performance as the NVIDIA model in almost, like, 2.7 X, like, lesser time. And then much faster than anything that hugging face or pajama does. Very interestingly, our three B model is pretty much the same performance a…”
Pratyush Maini Feb 10, 2026 ▶ 21:35
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.