“Some of the newer architectures don't actually employ it a lot. I think the last architecture that actually really employed it was the Mosaic MPT model class, and then almost all the models these days are all rope scaling, and then effectively you can use yarn with that as well.”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Mark Huang
Insight
Huang: True AI agents require measurable probability improvements per node
“It's like on each stage of the node, you're gonna have to see a marginal improvement in the probability of success for that particular workload because of non-determinism.”
Mark HuangMay 31, 2024▶ 5:16How to train a Million Context LLM — with Mark Huang of Gradient.ai
Opinion
Huang: Google's internal AI tooling was far superior to competitors
“Google was using AI for systems before everybody else too, right? They invented a transformer, and their internal set of tooling was just so far superior to everything else. Like, it's really hard for people to go back after seeing that.”
Mark HuangMay 31, 2024▶ 7:41How to train a Million Context LLM — with Mark Huang of Gradient.ai
Insight
Huang: RAG versus fine-tuning is fundamentally just meta-learning
“And like, at the end of the day, it's just all meta-learning, right? Like, all we want is, like, the best meta learning workflow or meta learning setup possible to be able to adapt the model to do anything.”
Mark HuangMay 31, 2024▶ 11:03How to train a Million Context LLM — with Mark Huang of Gradient.ai
AssertionOpen · timeframe May 2025
Huang: PoSE breaks down on needle-in-a-haystack at 500k tokens
“It does start to break down a little bit more on the longer, longer context. So, like, 500,000 to a million it appeared that it doesn't hold as well specifically for, like, needle in the haystack.”
Mark HuangMay 31, 2024▶ 24:11How to train a Million Context LLM — with Mark Huang of Gradient.ai
Insight
Huang: Adding one billion tokens cannot teach trillion-token models new knowledge
“All models these days are now double-digit trillions, right? So it's kind of a drop in the bucket if you really think I can just put, you know, a billion tokens in there, and I actually think that the model's gonna truly learn new Information.”
Mark HuangMay 31, 2024▶ 31:49How to train a Million Context LLM — with Mark Huang of Gradient.ai
AssertionSupported
Huang: Training CodeLlama on Llama 2 caused catastrophic language forgetting
“We do have historical precedent where CodeLlama was, you know, trained further from the original CodeLlama was trained further from Lama II, and it just lost, All its language capabilities, basically, right?”
Mark HuangMay 31, 2024▶ 33:03How to train a Million Context LLM — with Mark Huang of Gradient.ai
Made with StarZero
Turn any episode into a week of clips.
This entire site, over 200 episodes transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.