Pratyush Maini, researcher at Datology and CMU, discusses enterprise AI model training strategies.
Assertion Open · timeframe Feb 2027
Modern LLMs verbatim regurgitate JEE exam questions from two-word prompts
“We consistently saw how many of these, like, models today are being, like, massively, like, kind of fine-tuned on problems from... Like, oversight? Massively worked with. Like, even, like, imagine if I ask you the light bulb, what comes next in your mind? It w…”
Assertion Not checkable as stated
Labs train LLMs on benchmark questions for multiple epochs late in training
“It's very clear how the last stage of training for many of these models does have a massive amount of example or examine. Benchmaxing. Because the model will not, like, behaviorally complete exam questions with options if they have not really seen it at the en…”
Assertion Not checkable as stated
Self-reflection training data is now core to all frontier foundation models
“What this suggests about the GPT training data is that the self-reflection data has now actually become pretty much core to the training of all frontier models, because we're seeing that happen in non-instruct models across the board.”
Insight
Core AI capabilities must be built during pre-training, not just fine-tuned
“If there is a core capability that you actually care about, that capability should be part of the foundation and not a fine-tuned artifact.”
Insight
Small specialized pre-trained models can match capabilities of larger fine-tuned models
“When you think of the fact that by doing specialized pre-training, you can train a smaller model, which is as capable as a much larger model when fine-tuned.”
Assertion Partly supported
Datology BeyondWeb 3B matches NVIDIA Nemotron 8B in 2.7x less training time
“As you can see that we achieved the same performance as the NVIDIA model in almost, like, 2.7 X, like, lesser time. And then much faster than anything that hugging face or pajama does. Very interestingly, our three B model is pretty much the same performance a…”