Opinion
Enterprises demand full AI delegation, not 15% to 20% speed gains
“I think that the product experience of delegation is really, really immature right now. And most enterprises though, see that as the holy grail, not like going 15% or 20% faster.”
Prediction Not checkable as stated
Reyes: Inner-loop coding will soon be fully delegated to AI agents
“The outer loop of software development and what a software developer does, planning, talking with other human beings, interacting around what needs to get done, is something that's going to continue to be very human-driven, while the inner loop, the actual exe…”
Insight
Claude 3.7 Sonnet post-training heavily biases the model toward CLI tools
“So for example, Sonnet 3.7 clearly has it smells like cloud code, right? Same with codex. It very much impacted the way that those models want to write and edit code such that they seem to have a personality that wants to be in a CLI based tool.”
Insight
Reyes: Developers Should Not Need to Prompt Engineer AI Agents
“A lot of users we believe should not need to prompt engineer agents, right? If your time is being spent hyper optimizing every line and question that you pass to one of these systems, you're going to have a bad time.”
Insight
External scaffolding provides higher leverage for coding agents than fine-tuning models
“But our take in general is that freezing the model at a specific quality level and freezing the model at a specific data set just feels like it's lower leverage than continuing to iterate on all these external systems.”
Opinion
Amplitude and Statsig are closer to semantic AI observability than LLM tools
“I think Amplitude and Statsig and a lot of the feature flag companies actually are closer to this than the existing tools.”
Insight
$20/month local IDE pricing limits AI inference quality per task
“The cost when you are local first and your typical consumer is on a free plan or a Like, 20 dollar a month paid plan limits the amount of high quality inference you can do, and the scale or volume of inference you can do per, like, outcome.”
Insight
Reyes: High-quality enterprise codebases experience only 3% to 4% code churn
“In very high quality code bases, you'll see three percent, four percent code churn when they're at scale, right? This is like millions of lines of code. In poor code bases or poorly maintained code bases or early stage companies that are just changing a lot at…”
Insight
Reyes: Enterprise migration bottlenecks lie in planning and bureaucracy, not coding
“So a process that typically gets ultimately bottlenecked, not by like skilled humans writing lines of code, but by bureaucracy and technical complexity and understanding, Now gets condensed into basically how fast can a human being delegate the tasks appropria…”
Disclosure
Factory avoids parallel generation because the quality delta fails cost-benefit analysis
“Originally we had a lot of techniques that would generate a lot of stuff in parallel, and we still know how to do that. And we're very excited to bring that. But right now we don't do it because it's cost prohibitive and the quality Delta Is not enough to just…”
Prediction Not checkable as stated
Models post-trained for multi-hour agentic trajectories will be released very soon
“I mean, I'm thinking right off the bat, probably the biggest thing is models that have been post-trained on more general agentic trajectories over very long time spans. That feels like something that is There's an effort for that right now. But what I mean is …”
Insight
Reyes: AI agents can accurately imitate human-defined brand voice and design systems
“Being able to have someone set principles that are then consumable by our own agents, right? Design systems and consistency. I think it's pretty surprising the degree to which even like droids can actually imitate a brand voice and style that Cal created for u…”