model

also referred to as: models · the model

36 statements across 32 episodes · 16 bullish · 11 bearish · 32 people on the record · first statement Oct 21, 2023 by Kanjun Qiu · across every show →

Everything said about model, oldest first

Oct 21, 2023 positive
Assertion Supported
Kanjun Qiu: Larger models fine-tune faster and with higher sample efficiency
“As models get bigger, they fine tune faster. So they're more sample efficient as they get bigger.”
Kanjun Qiu Oct 21, 2023 ▶ 59:05 Why AI Agents Don't Work (yet) - with Kanjun Qiu of Imbue
Apr 11, 2024 neutral
Insight
Byun: Early LLMs prioritized answering questions over faithfulness to source text
“At the time, the models hadn't been trained at all to be faithful to a text. So they were just generating. So then when you ask them a question, they tried too hard to ask, answer the question, and didn't try hard enough to answer the question given the text o…”
Jungwon Byun Apr 11, 2024 ▶ 31:52 Supervise the Process of AI Research — with Jungwon Byun and Andreas Stuhlmüller of Elicit
Jun 11, 2024 bullish
Insight
Conover: AI models can parse documents to identify second-order derivative bets
“It is very straightforward to take a model and say, parse through all of these documents and find second order derivative bets and say, oh, it turns out that energy is like very, very adjacent to investments in AI and may not be priced in the same way that GPU…”
Mike Conover Jun 11, 2024 ▶ 55:12 How AI is Eating Finance - with Mike Conover of Brightwave
Aug 28, 2024 positive
Assertion Not checkable as stated
Carlini: LLMs decompile obscure binaries into readable Python code
“It can turn the compiled source code, which is impossible for any human to understand into the Python code that is entirely reasonable to understand. And, you know, it doesn't run. It has a bunch of problems, but like, it's so much nicer that it's immediately …”
Nicholas Carlini Aug 28, 2024 ▶ 29:52 Personal benchmarks vs HumanEval - with Nicholas Carlini of DeepMind
Oct 11, 2024
Insight
Goyal: Engineering around LLM limitations guarantees technical obsolescence
“If you make assumptions about the capabilities of models, and you engineer around them, you're almost, like, guaranteed to be screwed.”
Ankur Goyal Oct 11, 2024 ▶ 1:31:45 Production AI Engineering starts with Evals
Oct 11, 2024 bearish
Prediction Not checkable as stated
Goyal: Agent control flow and graph routing will move into models
“It feels very clear to me that this type of logic is going to be built into the model. Anytime there is control flow complexity or uncertainty complexity, I think the history of AI has been to push more and more into the model.”
Ankur Goyal Oct 11, 2024 ▶ 1:33:10 Production AI Engineering starts with Evals
Nov 11, 2024 negative
Opinion
Polu: Airbyte's Notion connector output is not useful for AI models
“And the reality is that if you look at Notion, Airby does the job of taking Notion and putting it in a structured way, but that's a way that is not really usable to actually make it available to models in a useful way. Because you get all the blocks, details, …”
Stanislas Polu Nov 11, 2024 ▶ 45:15 Agents @ Work: Dust.tt — with Stanislas Polu
Nov 15, 2024
Assertion Supported
Crivello: Lindy lets users configure different AI models per workflow node
“And you can see here for every node, you can configure which model you want to power the node.”
Florent Crivello Nov 15, 2024 ▶ 10:42 Agents @ Work: Lindy.ai (with live demo!)
Dec 24, 2024 neutral
Insight
Fu: Efficient AI Architectures Are Dead on Arrival Without Hardware Co-Design
“Even if your model is theoretically more efficient, if somebody goes and runs it and it's two times slower one of the things that, that we've learned is that if you're in that situation, it's just going to be dead on arrival. So you want to be designing your a…”
Dan Fu Dec 24, 2024 ▶ 15:51 2024 in Post-Transformer Architectures: State Space Models, RWKV [Latent Space LIVE! @ NeurIPS 2024]
Dec 25, 2024 negative
Assertion Not checkable as stated
Neubig: AI Models Are Poor At Pixel-Based Web Navigation
“The first way is this, the simplest way and the newest way, but it doesn't work very well, which is you take a screenshot of the website and then you click on a particular pixel value on the website and like models are not very good at that at the moment. Like…”
Graham Neubig Dec 25, 2024 ▶ 35:23 Best of 2024 in Agents (from #1 on SWE-Bench Full, Prof. Graham Neubig of OpenHands/AllHands)
Feb 1, 2025 bearish
Prediction Not checkable as stated
Nguyen: Website clicks will drop as internet access shifts to AI models
“In my opinion, like, people in, like, few years will click On, like, websites way less. I want to see the plot of, like, website clicks over time, but then my prediction is, like, it will go down and, like, people's access to the internet will be through the m…”
Karina Nguyen Feb 1, 2025 ▶ 56:35 The Agent Reasoning Interface: Claude, ChatGPT Canvas, Tasks, Operator — with Karina Nguyen, OpenAI
Mar 4, 2025 positive
Insight
Pokémon is ideal for testing AI agents because delays bring no penalty
“Pokemon's actually really nice because, like, if you don't do anything for five seconds, like, there's typically not a consequence by the nature of, like, doing inference on a model every, like, snapshot of time. It's actually a pretty good game to be able to …”
David Hershey Mar 4, 2025 ▶ 6:10 How Claude Plays Pokémon was made
May 7, 2025 positive
Assertion Not checkable as stated
Scandurra: Claude 3.7 uniquely gathers full context first before executing code edits
“Like, we've seen this pattern in Cloud 3.7, which is kind of different from the other models where, like, it really loves to gather up context first, and then it starts doing edits.”
Antonio Scandurra May 7, 2025 ▶ 17:39 Zed Agents — with Zed Cofounders Nathan Sobo & Antonio Scandurra
May 7, 2025 bullish
Prediction Not checkable as stated
Cherny: Foundation models will eventually subsume external memory and RAG architectures
“Everything is the model. Like that's the thing that wins in the end. And it just, as the model gets better, it's it subsumes everything else. So, you know, at some point the model will encode its own knowledge graph. It'll encode its own like KV story if you j…”
Boris Cherny May 7, 2025 ▶ 44:52 Claude Code: Anthropic's CLI Agent
May 9, 2025 bullish
Insight
Scaling autonomous task duration is a plausible path to AGI
“And that, if you can crack that scaling direction of, like, pushing the boundary of how long these models can go out and do these things for, that is a plausible path towards things that become marvelous.”
Will Brown May 9, 2025 ▶ 2:02 ⚡️Open Questions in Agentic RL — Will Brown (Prime Intellect)
May 29, 2025 bearish
Insight
External scaffolding provides higher leverage for coding agents than fine-tuning models
“But our take in general is that freezing the model at a specific quality level and freezing the model at a specific data set just feels like it's lower leverage than continuing to iterate on all these external systems.”
Eno Reyes May 29, 2025 ▶ 28:59 The AI Coding Factory
Jul 11, 2025 neutral
Assertion Not checkable as stated
Hsu: Only a few TTS models handle multilingual code-switching properly
“It's actually like only a few models are able to speak two languages in the same sentence and then pronounce them properly.”
Andrew Hsu Jul 11, 2025 ▶ 47:33 Personalized AI Language Education — with Andrew Hsu, Speak
Jul 16, 2025 bullish
Prediction Not checkable as stated
Rizwan: Developers will shift from tab autocomplete to natural language agents
“As the models get better, people are going to find themselves using natural language, working with an agent more and more and less being in the weeds and editing code and tab autocomplete.”
Saoud (Saud) Rizwan Jul 16, 2025 ▶ 11:33 Cline: The Collaborative AI Coder
Jul 31, 2025 positive
Insight
Lambert: RLVR on math does not degrade knowledge benchmark performance
“I think part of the intuition of RLVR is that the model is good at knowing which prompt area it is, which is why the models don't get worse on knowledge benchmarks if you're trading on like just math or precise instruction following. So the model just kind of …”
Nathan Lambert Jul 31, 2025 ▶ 59:33 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Aug 15, 2025 positive
Assertion Not checkable as stated
Brockman: LLMs consistently generalize to untrained preferences
“In order to get them to be able to operate according to different preferences and values, we just need to show that to them during training, and they are able to sort of generalize to different preferences and values that we didn't actually train against, and …”
Greg Brockman Aug 15, 2025 ▶ 38:56 Greg Brockman on OpenAI's Road to AGI
Aug 29, 2025 negative
Opinion
Morcos: Current AI models waste massive capacity memorizing unnecessary knowledge
“We're wasting a ton of capacity in these models on knowledge that is just totally unnecessary for them to have.”
Ari Morcos Aug 29, 2025 ▶ 1:06:42 Better Data is All You Need — Ari Morcos, Datology
Aug 29, 2025 bullish
Prediction Not checkable as stated
Morcos: Most AI models used in three years will be under 10B parameters
“Most of the models that the vast majority of people will be using in say three years will be single digit B or smaller.”
Ari Morcos Aug 29, 2025 ▶ 1:03:42 Better Data is All You Need — Ari Morcos, Datology
Sep 30, 2025 bearish
Opinion
Krieger: Current vision models lack the precision of skilled visual designers
“The models don't see as well as they could. They see, okay, you know, you ask them analyze a complex photo and they're able to do it, but I want them to be as persnickety as a like really good visual designer. Like, no, that looks, the baseline looks a little …”
Mike Krieger Sep 30, 2025 ▶ 10:07 ⚡️Claude Sonnet 4.5 and Anthropic's roadmap for Agents and Developers — Mike Krieger, Anthropic
Dec 26, 2025 positive
Disclosure
Fioca: OpenAI trains models to flexibly adapt across varied developer toolsets
“Initially, you know, our models are trained the way they were trained to use tools, and that kind of bakes in a habit, and so we've been getting the models better at using different types of tools.”
Brian Fioca Dec 26, 2025 ▶ 4:50 ⚡️GPT5-Codex-Max: Training Agents with Personality, Tools & Trust — Brian Fioca + Bill Chen, OpenAI
Dec 28, 2025 neutral
Insight
Underlying LLMs are technically completely uninvolved in the Model Context Protocol
“MCP is a protocol between the AI application and like servers, right? So the model is actually technically not involved in MCP.”
David Soria Parra Dec 28, 2025 ▶ 21:21 One Year of MCP — with David Soria Parria and AAIF leads from OpenAI, Goose, Linux Foundation
Dec 30, 2025 bearish
Insight
Nair: AI models lag orders of magnitude behind human one-shot error learning
“It seems like we're kind of, like, a few orders of magnitude of, like, kind of data efficiency, basically, away from, like, that kind of, like, you know, you do something once, or, like, you make a mistake like, you, yeah, you introduce, like, a bug in your co…”
Ashvin Nair Dec 30, 2025 ▶ 35:45 [State of RL/Reasoning] IMO/IOI Gold, OpenAI o3/GPT-5, and Cursor Composer — Ashvin Nair, Cursor
Feb 5, 2026 neutral
Opinion
Bissell: Subliminal learning typically only affects models sharing initial random seeds
“I think it only applies to models that were initialized from the same starting Z. Usually, yes.”
Mark Bissell Feb 5, 2026 ▶ 13:24 Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell
Mar 17, 2026 neutral
Opinion
Rieseberg: Unclear if agent hyper-optimizations remain relevant in next-gen models
“Will those gaps still exist in the next few generations of models? It's like a little unclear to me though. Because right now these like hyper optimizations we make, I'm not sure for how long they're still really relevant.”
Felix Rieseberg Mar 17, 2026 ▶ 21:13 Anthropic’s Felix Rieseberg on AI Coworkers, Local-First Agents, and the Future of Knowledge Work
Mar 17, 2026 bearish
Prediction Not checkable as stated
Rieseberg: Specialized AI wrapper apps won't survive as models generalize
“I think we're going to see a lot of like applications and companies that do very impressive things with AI that in the short term might seem very effective because they're very specialized to individual use cases. But I think once models get better at generali…”
Felix Rieseberg Mar 17, 2026 ▶ 25:29 Anthropic’s Felix Rieseberg on AI Coworkers, Local-First Agents, and the Future of Knowledge Work
May 5, 2026 bullish
Opinion
Frontier AI Models Can Solve Six-Month Graduate Physics Starter Problems
“And I think the issue is that many such problems now, I would say these models can probably crush. Yeah. These are problems that we usually take again, you know, timescale for a theoretical physics paper is six months to a year. That's pretty typical.”
Alex Lupsasca May 5, 2026 ▶ 56:31 🔬How GPT‑5 derived new results in theoretical physics and quantum gravity — Alex Lupsasca, OpenAI
Jun 1, 2026
Insight
Ethan He: External heuristic engineering gets absorbed into models
“From our experience, the heuristic engineering also have the models get absorbed into the models themselves.”
Ethan He Jun 1, 2026 ▶ 1:37:17 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Jun 4, 2026 bearish
Assertion Partly supported
Petersson: Models Score No Better Than Random on BlueprintBench Floorplans
“And it turns out the models are absolutely horrible at this. No one scores statistically better than random chance.”
Lukas Petersson Jun 4, 2026 ▶ 57:58 When AI Agents Run Businesses — Lukas Petersson and Axel Backlund of Andon Labs
Jun 22, 2026 negative
Insight
Fredrikson: Evaluation-aware AI models often execute harmful actions because it is a simulation
“If you make, if you're testing the model for robustness or safety, right? And it's aware that it's being tested because you've set things up in a very artificial way, right? Like the email addresses are at example.com. The webpage is clearly not a real webpage…”
Matt Fredrikson Jun 22, 2026 ▶ 23:32 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Jun 22, 2026 positive
Assertion Open · timeframe Jun 2026
Fredrikson: Top AI browser agents yielded only a handful of successful breaks
“There were a couple of models that seemed to be very, very robust, right? Like the red teamers found just a handful of successful breaks on them.”
Matt Fredrikson Jun 22, 2026 ▶ 22:29 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Jun 25, 2026 bullish
Disclosure
Chen: OpenAI's three-year goal is models conducting end-to-end research
“When we look at our kind of three-year roadmap, right the end goal that we want to reach is one where You know, the models are just doing end-to-end research, and I think a part of that problem is just being able to have the model come up with good taste.”
Mark Chen Jun 25, 2026 ▶ 34:15 Cooking with OpenAI’s Research Chief: AGI, o1, Evals, and Scaling Laws — Mark Chen
Jul 16, 2026 bullish
Insight
Beam: Cross-domain training reduces domain data requirements in science models
“And so again, the core bet that we're making is that is true for science. That if the model is trained on an increasingly broad set of data, the amount of data that you need in a given domain, that data requirement is reduced. In some cases will be reduced to …”
Andy Beam Jul 16, 2026 ▶ 30:20 🔬 RL with Verifiable Rewards, but the Verifier is a Lab — Lila Sciences
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.