frontier models

also referred to as: frontier model

19 statements across 17 episodes · 7 bullish · 9 bearish · 17 people on the record · first statement Aug 4, 2025 by Stefano Ermon · across every show →

Everything said about frontier models, oldest first

Aug 4, 2025 bullish
Prediction Not checkable as stated
Ermon: Power constraints will drive diffusion models to replace frontier LLMs
“If it happens, it's gonna be driven by efficiency. Like we're all constrained by essentially power. And if you have, I mean, at the end of the day, it's all an inference game, right? Okay. Training is expensive, but then the thing that matters is being able to…”
Stefano Ermon Aug 4, 2025 ▶ 23:42 ⚡️Mercury: Ultra-Fast Diffusion LLMs — Estefano Ermon, CEO Inception Labs
Nov 22, 2025 bullish
Insight
Wagner: Fine-tuned open-source models can beat frontier models on quality
“I think you're now with, you know, we'll be seeing the trend going also with the coding agents, you know, fine-tuning open source models and like way faster as a result and more accurate. You can actually beat frontier models on quality that way.”
Matthias Wagner Nov 22, 2025 ▶ 27:28 ⚡️ Building the AI Hardware Engineer with Matthias Wagner, Co-founder of Flux
Dec 7, 2025 positive
Disclosure
Ubl: Vercel's Composite Models Are Faster Than Agentic Loops
“Basically what we do is we have this like composite model architecture. We run the frontier model and then we run the fine tune model after to fix its errors. That doesn't perform better than an agentic loop, but it's orders of magnitude faster, right?”
Malte Ubl Dec 7, 2025 ▶ 16:51 The Great Evals Debate — Ankur Goyal & Malte Ubl
Dec 18, 2025 bullish
Prediction Not checkable as stated
Zhang: AI models will handle simple vision natively, using tools for complexity
“I think at least I want to bet on, you know, running their work natively together, the future for simple, I would say for simple or even intermediate difficult vision tasks. For example, kind of counting with less than 20 objects. I think for this kind of simp…”
Pengchuan Zhang Dec 18, 2025 ▶ 54:11 SAM 3: The Eyes for AI — Nikhila & Pengchuan (Meta Superintelligence), ft. Joseph Nelson (Roboflow)
Jan 28, 2026 negative
Assertion Not checkable as stated
White: Frontier AI models cannot handle molecules well
“We noticed that the frontier models can't work with molecules very well. So let's make a model with intuition for medicinal chemistry. And that was what led to ether zero.”
Andrew White Jan 28, 2026 ▶ 35:49 🔬 From Red Teaming GPT-4 to Automating Drug Discovery: The Future of AI in Science — Andrew White
Feb 10, 2026 neutral
Assertion Not checkable as stated
Self-reflection training data is now core to all frontier foundation models
“What this suggests about the GPT training data is that the self-reflection data has now actually become pretty much core to the training of all frontier models, because we're seeing that happen in non-instruct models across the board.”
Pratyush Maini Feb 10, 2026 ▶ 15:26 ⚡️ Reverse Engineering OpenAI's Training Data — Pratyush Maini, Datology
Feb 12, 2026
Insight
Jeff Dean: Capable small models require first building frontier models
“Through distillation, which is a key technique for making the smaller models more capable, you know, you have to have the frontier model in order to then distill it into your smaller model. So it's not like an either or choice. You sort of need that in order t…”
Jeff Dean Feb 12, 2026 ▶ 3:06 The AI Frontier: from Gemini 3 Deep Think distilling to Flash — Jeff Dean
Mar 5, 2026 negative
Assertion Contradicted
Huber: Frontier AI models are not actually good at agentic search
“We've like sort of stress tested like frontier models and their ability to search. And they are not actually that good at searching.”
Jeff Huber Mar 5, 2026 ▶ 26:43 Why Every Agent Needs a Box — Aaron Levie, Box
Mar 5, 2026 negative
Insight
Huber: Frontier models repeat mistakes if failed actions remain in context
“A few of the insights is, like, everyone, frontier model is not good at search. Humans have this natural explore-exploit trade-off, where we kind of understand, like, when to stop doing something. Also, humans are pretty good at, like, forgetting, actually, li…”
Jeff Huber Mar 5, 2026 ▶ 29:56 Why Every Agent Needs a Box — Aaron Levie, Box
Apr 18, 2026 bearish
Prediction Open · timeframe Apr 2031
Swix: Frontier model context windows will not reach billion-token scales
“We took, you know, three years to go from 100 K context to one million context in every frontier model, but we're not going to a billion, you're not going to a trillion with context graphs, context lengths.”
Shawn Wang Apr 18, 2026 ▶ 25:12 ⚡️ How to turn Documents into Knowledge: Graphs in Modern AI — Emil Eifrem, CEO Neo4J
May 28, 2026 neutral
Disclosure
Yan: Devin requires orchestrating multiple frontier models for end-to-end app testing
“Well, in some cases we found that actually no one frontier model can actually do this full end-to-end task itself. We've seen cases where we actually had had to orchestrate different frontier models together to kind of solve this problem together.”
Walden Yan May 28, 2026 ▶ 22:07 Devin’s 80% Moment: Background Agents, 7x PRs, & End of Hand-Held Coding — Walden Yan & Cole Murray
Jun 4, 2026 positive
Assertion Supported
Petersson: Frontier AI models now survive the full year in VendingBench
“The models at the time were worse, so they crashed out earlier and now they survive the full year all the time.”
Lukas Petersson Jun 4, 2026 ▶ 9:34 When AI Agents Run Businesses — Lukas Petersson and Axel Backlund of Andon Labs
Jun 18, 2026 negative
Assertion Not checkable as stated
Midha: Frontier AI Models Were Terrible at Analyzing Condensed Matter Physics Data
“We had started benchmarking frontier models on physics and science capabilities, and they were not very good. They were good at, like, doing things like summarization of papers, but if you said, hey, could you, like, analyze the scientific data coming out of a…”
Anjney Midha Jun 18, 2026 ▶ 55:48 Why AI Labs With Unlimited GPUs Still Fail — Anjney Midha, AMP
Jun 21, 2026 bullish
Assertion Not checkable as stated
Malde: SWE-ONE beat frontier models via user-signal post-training
“And this was the kind of major unlock for the company as well, is we had all this massive data. We were able to post train on all of that user signal and now beat the frontier.”
Ronak Malde Jun 21, 2026 ▶ 3:06 ⚡️Every product of the future will be a living system — Ronak Malde, Trajectory.ai
Jun 22, 2026 negative
Insight
Kolter: Frontier models fail at red teaming due to safety refusals
“So generally speaking, the issue with this is that frontier models are extremely bad at automated red teaming because they have a lot of safeguards built into them. So if you try to use them to jailbreak another model, they will actually refuse their safety tr…”
Zico Kolter Jun 22, 2026 ▶ 11:06 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Jun 24, 2026 bullish
Assertion Not checkable as stated
Zaharia: Databricks' document vision model is ~100x cheaper than frontier models
“Our team built this document sort of vision model that takes a page and gives you back a nice JSON with all the components. And it's very competitive. It's like probably like a hundred X cheaper than those frontier models and still better.”
Matei Zaharia Jun 24, 2026 ▶ 1:01:56 The Agent Cloud: Databricks’ Bet on the Future of AI — Matei Zaharia and Reynold Xin
Jul 13, 2026 bearish
Prediction Not checkable as stated
Biderman: Model accuracy will still degrade at 10M context window scale
“But two is like, for the agentic tasks of 18 months from now, inside those major repositories of knowledge, and asking the models more and more things in underspecified ways, I suspect that the accuracy of the models would go down. The phenomenon of context fr…”
Dan Biderman Jul 13, 2026 ▶ 17:31 The AI Memory Problem: Why Long Context Isn’t Enough — Dan Biderman, Engram Co-founder & CEO
Jul 13, 2026 negative
Assertion Not checkable as stated
Biderman: Harmless enterprise queries on frontier models cost thousands of dollars
“And now you can solve these tasks with frontier models and compaction. And when you ask them to do so, they will consume thousands of dollars for queries that we think are harmless. That every employee in the company would be able to answer.”
Dan Biderman Jul 13, 2026 ▶ 25:50 The AI Memory Problem: Why Long Context Isn’t Enough — Dan Biderman, Engram Co-founder & CEO
Aug 21, 2026 bearish
Assertion Not checkable as stated
Park: Frontier models hit only 20-30% accuracy predicting niche human behavior
“Where in some cases, the model performance of frontier models go all the way down to 20, 30%. Especially if you go into that more niche population on topics that our customers will actually care about. On more gen pop, it might be around 50 to 60%.”
Joon Sung Park Aug 21, 2026 ▶ 31:17 Simulating Humanity: from Generative Agents to 8 Billion Digital Twins — Joon Sung Park, Simile AI
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.