Pratyush Maini

Founding Member of Technical Staff, DatologyAI · 1 appearance on the record.

computed by AI from the episodes · how this works → · full disclaimer →

academicscientistengineer@pratyushmaini ↗LinkedIn ↗pratyushmaini.github.io ↗

Pratyush Maini conducts research on data-centric AI, exploring how pre-training data curation, synthetic data, and machine unlearning influence foundation models. He co-authored research on LLM unlearning benchmarks and serves as a founding member of technical staff at DatologyAI.

12statements → 10claims → 5claims resolved → 80%fully supported → 3.83/5average certainty → 2.5/5average debate potential → 1said about them ↓

4 supported 1 partly supported 0 contradicted 2 not yet assessed 3 not checkable as stated how the 10 claims stand · each chip opens the sources

1 prediction · 9 assertions · 2 insights · every statement was checked. The prediction and assertions are the 10 claims: statements the public record can support or contradict. 5 are resolved, 2 are not yet assessed, and 3 name no date, number or outcome precise enough to check. Everything else (opinions, insights, what ifs, disclosures) can never be settled by the record, so it carries no assessment.

The record, in short

What the tape says about how Pratyush argues and how the claims held up. Everything they said, and everything said about them, is in the tabs below.

Their most notable supported claim

Assertion Supported
MoE models with 20B active parameters match 72B dense models in memorization
“With the coming of the MOE models, we're starting to see these phenomena, even like models which have a less number of active parameters. So a phenomena that we would observe at a 72 B model in the past is probably now visible at a 20 B active one 20 B MOE als…”
Pratyush Maini Feb 10, 2026 ▶ 5:24 ⚡️ Reverse Engineering OpenAI's Training Data — Pratyush Maini, Datology

Expressed certainty vs assessment result

none yet certainty 1
none yet certainty 2
100% certainty 3
80% certainty 4
none yet certainty 5

weighted support: a fully supported claim counts one, a partly supported claim counts half. Each filled bar is clickable and opens exactly those claims; "none yet" means nothing said at that certainty level has resolved yet

Everything Pratyush Maini said on Latent Space that made the record, most notable first. Filter by type, assessment or year in the ledger →

Assertion Open · timeframe Feb 2027
Modern LLMs verbatim regurgitate JEE exam questions from two-word prompts
“We consistently saw how many of these, like, models today are being, like, massively, like, kind of fine-tuned on problems from... Like, oversight? Massively worked with. Like, even, like, imagine if I ask you the light bulb, what comes next in your mind? It w…”
Pratyush Maini Feb 10, 2026 ▶ 3:36 ⚡️ Reverse Engineering OpenAI's Training Data — Pratyush Maini, Datology
Assertion Not checkable as stated
Labs train LLMs on benchmark questions for multiple epochs late in training
“It's very clear how the last stage of training for many of these models does have a massive amount of example or examine. Benchmaxing. Because the model will not, like, behaviorally complete exam questions with options if they have not really seen it at the en…”
Pratyush Maini Feb 10, 2026 ▶ 4:17 ⚡️ Reverse Engineering OpenAI's Training Data — Pratyush Maini, Datology
Assertion Not checkable as stated
Self-reflection training data is now core to all frontier foundation models
“What this suggests about the GPT training data is that the self-reflection data has now actually become pretty much core to the training of all frontier models, because we're seeing that happen in non-instruct models across the board.”
Pratyush Maini Feb 10, 2026 ▶ 15:26 ⚡️ Reverse Engineering OpenAI's Training Data — Pratyush Maini, Datology
Insight
Core AI capabilities must be built during pre-training, not just fine-tuned
“If there is a core capability that you actually care about, that capability should be part of the foundation and not a fine-tuned artifact.”
Pratyush Maini Feb 10, 2026 ▶ 18:53 ⚡️ Reverse Engineering OpenAI's Training Data — Pratyush Maini, Datology
Prediction Not checkable as stated
Enterprises will widely adopt specialized AI pre-training in 2026 and 2027
“So I think like, 26 and 27 are going to be the years where different enterprises start doing specialized pre-training, because the cost of pre-training really amortizes itself very fast.”
Pratyush Maini Feb 10, 2026 ▶ 19:27 ⚡️ Reverse Engineering OpenAI's Training Data — Pratyush Maini, Datology
Insight
Small specialized pre-trained models can match capabilities of larger fine-tuned models
“When you think of the fact that by doing specialized pre-training, you can train a smaller model, which is as capable as a much larger model when fine-tuned.”
Pratyush Maini Feb 10, 2026 ▶ 19:38 ⚡️ Reverse Engineering OpenAI's Training Data — Pratyush Maini, Datology
Assertion Partly supported
Datology BeyondWeb 3B matches NVIDIA Nemotron 8B in 2.7x less training time
“As you can see that we achieved the same performance as the NVIDIA model in almost, like, 2.7 X, like, lesser time. And then much faster than anything that hugging face or pajama does. Very interestingly, our three B model is pretty much the same performance a…”
Pratyush Maini Feb 10, 2026 ▶ 21:35 ⚡️ Reverse Engineering OpenAI's Training Data — Pratyush Maini, Datology
Assertion Open · timeframe Feb 2027
Qwen 3 memorizes benchmark questions significantly more than Qwen 1.5 or 2
“I don't, we don't see this phenomenon like the earlier versions of like QN 1.5 or even QN two, but start seeing it in QN three. So there's something about a more, like a significantly higher weight on benchmark or like a J or like any times of evaluation quest…”
Pratyush Maini Feb 10, 2026 ▶ 4:59 ⚡️ Reverse Engineering OpenAI's Training Data — Pratyush Maini, Datology
Assertion Supported
MoE models with 20B active parameters match 72B dense models in memorization
“With the coming of the MOE models, we're starting to see these phenomena, even like models which have a less number of active parameters. So a phenomena that we would observe at a 72 B model in the past is probably now visible at a 20 B active one 20 B MOE als…”
Pratyush Maini Feb 10, 2026 ▶ 5:24 ⚡️ Reverse Engineering OpenAI's Training Data — Pratyush Maini, Datology
Assertion Supported
Non-reasoning Grok, GPT, and Gemini models exhibit recursive self-correction loops
“And so then I tried this across models, and I saw consistently across Grok, and GPT, and Gemini, that you were seeing this phenomena where models will, like, self-correct themselves quite a bit, and these were, like, non-thinking models. They were, like, the s…”
Pratyush Maini Feb 10, 2026 ▶ 7:08 ⚡️ Reverse Engineering OpenAI's Training Data — Pratyush Maini, Datology
Assertion Supported
Pre-December 2024 OpenAI models did not exhibit seahorse emoji self-correction loops
“And so I like ran the OpenAI API across like models released from 23 to 25, and you would see like all the models until twenty-twenty-four December had very Terce and short responses to the question. Is there a seahorse emoji? They would either say that there …”
Pratyush Maini Feb 10, 2026 ▶ 9:04 ⚡️ Reverse Engineering OpenAI's Training Data — Pratyush Maini, Datology
Assertion Supported
OLMo 3 injects thinking traces during mid-training rather than post-training
“So, with the instruct variant, they do not have any thinking data in the post-training phase, but they do mention that we are going to put, like this is the line that they write, there is some intentional addition of thinking traces in the mid-training phase o…”
Pratyush Maini Feb 10, 2026 ▶ 14:48 ⚡️ Reverse Engineering OpenAI's Training Data — Pratyush Maini, Datology

The other half of the tape: Pratyush Maini's own voice is left out of every number here. Other people bring the name up 1 time in 1 episode on Latent Space. every mention, with the transcript →

Who brings them up most Loubna Ben Allal 1

Every mention by year

tap a year for its mentions
0011112024episodesmentions
0112024episodes it came up in
000.50.5112024episodesmentions per episode

Appearances (1)

EpisodeDateSpeaking time
⚡️ Reverse Engineering OpenAI's Training Data — Pratyush Maini, Datology Feb 10, 2026 21m
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.