O'Driscoll: Advanced AI models have had a massive impact on cyber risk
“I mean, I think to be fair, unlike some of the other PDOOM stuff, right, there's real evidence that the impact of these models on cyber risk has been massive.”
Sawhney: AI will produce exponentially more math, making it easier to absorb
“Along, I mean, of course models are going to help us produce exponentially more mathematics, but they also make it much easier to absorb it and right now, okay, it's still a bit of a challenge back and forth, but I think it's, for me at least much, much faster…”
AI foundation models are not fungible commodities
“I think for people who believe the models are commodities or totally fungible, you just haven't actually used the models.”
AI models differ by cognitive disposition rather than raw intelligence
“It's not that one is ahead of another one is more intelligent. It's rather One is, sort of, has a mind that's shaped in one direction, perhaps creativity and openness for Quen, and others that are shaped in other directions, like, you know, neuroticism and pre…”
Keep a low-stakes persistent project chassis to test new AI models
“And I think that if you don't have a chassis on which to like with which to use the models, it's really hard to come up with an idea from scratch every time. So I I'd say like work on something. It's actually better if it's not important with a capital I, and …”
Reyes: AI foundation model margins are worse than application margins
“The margin profile of the models is definitely worse than the applications.”
Reyes: Agent harnesses solve AI problems, not models or gateways
“People really want the problems to be solved, like sort of somewhere else, like in the model or in the gateway, but more and more we see it's the harness that solves these problems.”
Acharya: An AI agent is simply a model in a loop with tools and memory
“Agent is just a model in a loop with sort of tools and memory and a few other things.”
Movva: Models have trained on the whole internet and exhausted human web data
“Models have seen the entire internet many times over at this point, and there is not a whole lot more to be done on human data from the internet.”
Movva: AI models will excel at kernel engineering within six months
“I'm sure in six months time we'll have much better models on kernel engineering, and I'm sure the labs would tell you that they already do a lot of their kernel engineering in a fully automated way.”
Gupta: AI models should be treated as interpreters, not data analyzers
“And I feel we should look at models as interpreters rather than someone who's analyzing. Analysis is still going to be on the consumer side.”
Sharp: AI will commoditize software features across all vendors
“Building software is so productive, especially copying software. So if I, like, if you have a feature, I can implement a version of that feature. Super easy. You point the models at it, and they make you something that looks like that, feels like that, does th…”
Friedberg: China's power production advantage will dominate if model IP gap closes
“If they can eliminate the IP advantage and the knowledge advantage that sits in models, they have the advantage with power production in every which way.”
Schrader: LLMs lack embedded operational workflows despite massive training data
“The models are trained on so much data and they're so large and yet they actually don't really know how to do any of this work. Like they don't have the workflows encoded in any way.”
Altman: Frontier AI models still cannot learn continuously as they operate
“The model also, although brilliant, is still not learning continuously as it goes.
And that feels to me like, maybe not a hard requirement for AGI, but certainly something that I'd like.”
Cherny: AI evals saturate and must be discarded every few generations
“I think evals, they outlive the harness a little bit, but not quite that much. Like, an eval might live for maybe one, two, three model generations, but nowadays the, you know, we're on the exponential. The model is improving so quickly, very often we just sat…”
Mandia: Criminals will inevitably deploy ungated frontier AI models
“The models may have gates that prevent you from doing things, but even those folks that make them probably need to remove some of the gating factors because it's so important to make sure there will be models that can think, learn, and do incredible things tha…”
Beam: Cross-domain training reduces domain data requirements in science models
“And so again, the core bet that we're making is that is true for science. That if the model is trained on an increasingly broad set of data, the amount of data that you need in a given domain, that data requirement is reduced. In some cases will be reduced to …”
Field: AI model interfaces will evolve far beyond basic text prompting
“I just think that the interfaces for how you interact with models will evolve so much. And to the degree that models are part of software, we will see a ton of exploration there.”
Brown predicts AI will zero-shot his entire PhD thesis within one year
“And I wouldn't be surprised if, you know, six months or a year from now, the model is able to do zero shot an entire poker solver, basically my entire PhD thesis in one go.”
Chen: OpenAI's three-year goal is models conducting end-to-end research
“When we look at our kind of three-year roadmap, right the end goal that we want to reach is one where You know, the models are just doing end-to-end research, and I think a part of that problem is just being able to have the model come up with good taste.”
Fredrikson: Top AI browser agents yielded only a handful of successful breaks
“There were a couple of models that seemed to be very, very robust, right? Like the red teamers found just a handful of successful breaks on them.”
Fredrikson: Evaluation-aware AI models often execute harmful actions because it is a simulation
“If you make, if you're testing the model for robustness or safety, right? And it's aware that it's being tested because you've set things up in a very artificial way, right? Like the email addresses are at example.com. The webpage is clearly not a real webpage…”
AI coding improvements boost model performance across unrelated non-coding tasks
“When the models improve in coding, they have out of domain performance. They also improve in other tasks that have nothing to do with coding for reasons we don't fully understand.”
Franceschi: Organizing Model Context Is the Bottleneck for Most AI Applications
“I think a lot of the work to your point is the, how do you organize the context for the model? And you can use a model to help, but that is the bottleneck for most things.”
Ambati: AI foundation models will capture more than 10% of software
“We are starting to see more evidence that models are going to own that more than the 10%, ah, you know, whether it's the new plugins that, ah, you know, Claude has released that shook up the whole SecOps market.”
Petersson: Models Score No Better Than Random on BlueprintBench Floorplans
“And it turns out the models are absolutely horrible at this. No one scores statistically better than random chance.”
Ethan He: External heuristic engineering gets absorbed into models
“From our experience, the heuristic engineering also have the models get absorbed into the models themselves.”
Shipper: AI models fundamentally make yesterday's human competence cheap
“What a new model drop does, or what models do in general, is they make yesterday's human competence cheap.”
Shipper: AI models will structurally always trail behind creative human users
“Structurally, because of the way the models work, because of the financial incentives of model of model companies to like make them compliant and aligned structurally, they're always going to be trailing behind those people who are taking, taking the models an…”
Dave Morin: Google's new AI model is worse than top competitors
“The most reason, the model's not even competitive. What's not competitive about it? It's not competitive. Like, it's just less good than every conceivable benchmark. It's less good than five, five and, ah, four, seven. Like, meaningful. Like, you wouldn't use …”
Baker: Cerebras hardware can theoretically run any model size
“A Cerebrus machine can theoretically run any size model.”
Baker: AI profits accrue to infrastructure and models, not applications
“Today in that, in Jensen's five layer cake of AI. The profits. They're accruing to energy. They're accruing to data centers. They're accruing to chips. They're accruing to models. Not really accruing to the applications.”
Burazin: Current AI models do not learn from on-the-job execution
“Models actually don't learn right now. So you use a model and if you solve memory, it has memory of things. So it has context of these things. And so it can, oh, here's the context, and so it can have a better answer because it has the context, but it doesn't …”
Frontier AI Models Can Solve Six-Month Graduate Physics Starter Problems
“And I think the issue is that many such problems now, I would say these models can probably crush. Yeah. These are problems that we usually take again, you know, timescale for a theoretical physics paper is six months to a year. That's pretty typical.”
Nair: AI foundation models will commoditize just like CPUs did
“Developers are going to think, just like CPUs used to be like so Intel inside, and now nobody cares. I think models will become that.”
Ross Mike: Treat AI models like new employees, not all-knowing magic
“We should treat models and these agents like very new employees versus like these black magic boxes that like know everything. Right? They know everything because they've been trained on a lot of data, but they don't know your workflow, your steps, right?”
Sankar: Foundational AI models are rapidly commoditizing
“What I see happening empirically is the models are being commoditized and we're always under pressure.”
Jensen Huang: AI models are underlying technology, not standalone products
“Models is a technology, not a product. Models is a technology, not a service.”
Rieseberg: Unclear if agent hyper-optimizations remain relevant in next-gen models
“Will those gaps still exist in the next few generations of models? It's like a little unclear to me though. Because right now these like hyper optimizations we make, I'm not sure for how long they're still really relevant.”
Rieseberg: Specialized AI wrapper apps won't survive as models generalize
“I think we're going to see a lot of like applications and companies that do very impressive things with AI that in the short term might seem very effective because they're very specialized to individual use cases. But I think once models get better at generali…”
Evans: No AI lab has unique models because everyone builds identical tech
“What's happened so far is that people leapfrog each other every couple of weeks or every month or two, but because everyone is basically building the same stuff nobody has anything unique in the models.”
Cherny: Strong evidence shows LLMs perform reasoning beyond next-token prediction
“You know, like a long time ago, we weren't sure if the model was just predicting the next token or is doing something a little bit deeper. Now I think there's actually quite strong evidence that it is doing something a little bit deeper.”
Lonsdale: 8VC Application Startups Regularly Swap Underlying AI Models
“A lot of our level five companies, whether they're doing healthcare billing or work, logistics workflows, or gosh, there's a whole lot of lists of them we can go through. Like they're using whatever models are best at the time, but they're swapping between mod…”
Wu: AI models could execute multi-hour to day-long tasks in 12-18 months
“If you follow this trend, like, I think, like, in the next 12 to 18 months, we could see models that could do multi-hour long tasks very, very coherently. At some point, it might reach, like, you know, six hours a day long task.”
Steinberger: Perceived AI model degradation is actually just rising user expectations
“No, they didn't do anything, you just adapted to the new standard, and now your expectations went up, but the model is still the average.”
Bissell: Subliminal learning typically only affects models sharing initial random seeds
“I think it only applies to models that were initialized from the same starting Z. Usually, yes.”
Das: AI models are unreliable because they lack core structural understanding
“There are many cases where models Should be more deterministic if they understood the principle involved. But they're not. They're very undeterministic because they have no core structural understanding of things.”
Steinberger: Software tools must be designed for AI models, not human users
“And if you are smart, you build it in a way that just uses what the model already expects, you know, don't build it for humans, build it for models.”
Arnovitz: Multi-model peer reviews mitigate individual AI coding weaknesses
“So I think that using all these models and basically playing to their strengths and mitigating their weaknesses by using other models is, is a game changer for me. So I'll do peer review a bunch of times and I'll have other models review other models code and …”
Izmailov: Rogue AI science fiction in training data likely causes deceptive behavior
“I think at least part of it is probably The models seeing descriptions of AI, like in the science fiction literature going rogue and like, yeah, that probably affects how the models behave in similar scenarios.”
Reganti: 80% of AI engineering is understanding workflows, not complex models
“80% of so-called AI engineers, AI PMs spend their time actually understanding their workflows very well. They're not building the fanciest and the, you know, most cool models or workflows around it. They're actually in the weeds understanding their customers' …”
Nair: AI models lag orders of magnitude behind human one-shot error learning
“It seems like we're kind of, like, a few orders of magnitude of, like, kind of data efficiency, basically, away from, like, that kind of, like, you know, you do something once, or, like, you make a mistake like, you, yeah, you introduce, like, a bug in your co…”
Underlying LLMs are technically completely uninvolved in the Model Context Protocol
“MCP is a protocol between the AI application and like servers, right? So the model is actually technically not involved in MCP.”
Fioca: OpenAI trains models to flexibly adapt across varied developer toolsets
“Initially, you know, our models are trained the way they were trained to use tools, and that kind of bakes in a habit, and so we've been getting the models better at using different types of tools.”
Xanthos: AI model progress has not stalled and remains exponential
“I do believe we're still in an exponential improvement curve with AI. I think the models keep improving quite a bit because there was maybe a concern, maybe a few months ago, more like, you know, end of last year, I would say a year ago, whether let's say the …”
Sacerdote: Inference-time reasoning will be adopted by all AI models
“And now deep seek is basically solidifying this inference time reasoning is going to be adopted by all the models.”
Senra: Koenigsegg sold out entire production runs sight unseen
“Conaseg's entire run of some models would sell out sight unseen.”
Evans: AI models are commodities outside specialist use cases
“It's been very clear for like the last year that these models are commodities outside of specialist use cases and outside of very, people who are very, very deeply into using them all the time every day. Unless you're doing image, for image generation or codin…”
Randle: Financial models in VC only serve to test conviction
“And so I think beyond being like a yardstick to test your conviction, models aren't that useful, but for that, they're really, really good.”