Insight certainty 3/5 debate potential 3/5

Goyal: Systems betting on intrinsic LLM reasoning improvements are more durable

Ankur Goyal · Production AI Engineering starts with Evals · Oct 11, 2024 · at 1:45:53

Ankur Goyal is the founder and CEO of Braintrust. He describes architectural design principles for building software around evolving AI foundation models.

0:00 / 0:11exact quote · 11.6s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“If you build your system in a way that Kind of assumes LLMs will get better at reasoning and get better at sort of agentic tasks in the LLM itself. Then I think you will build a more durable system.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Ankur Goyal

Opinion
Goyal: Running LLM workloads at scale is impractical outside OpenAI
“It's just not practical outside of OpenAI to run use cases at scale in a lot of cases. Like, you can do it, but it requires quite a bit of work. And Because OpenAI is so good at making their models so available, I think they get a lot of credit for the science…”
Ankur Goyal Oct 11, 2024 ▶ 1:27:25 Production AI Engineering starts with Evals
Prediction Not checkable as stated
Goyal: Agent control flow and graph routing will move into models
“It feels very clear to me that this type of logic is going to be built into the model. Anytime there is control flow complexity or uncertainty complexity, I think the history of AI has been to push more and more into the model.”
Ankur Goyal Oct 11, 2024 ▶ 1:33:10 Production AI Engineering starts with Evals
Prediction Not checkable as stated
Goyal: OpenAI o1 will make agentic frameworks obsolete
“And I think O-one is going to do that to agentic frameworks as well. Hey, I think To me, it seems very unlikely that the, you know, you and me sort of like sipping an espresso and thinking about how, like, different personified roles of people should interact …”
Ankur Goyal Oct 11, 2024 ▶ 1:34:00 Production AI Engineering starts with Evals
Opinion
Goyal: Open-source inference providers are far less reliable than OpenAI
“They are nowhere near as reliable as, I mean, every single time I use any of those products and run a benchmark, I find a bug, text the CEO, and they fix something. It's nowhere near where OpenAI is.”
Ankur Goyal Oct 11, 2024 ▶ 1:37:24 Production AI Engineering starts with Evals
Insight
Goyal: Publishing Public Benchmarks Is Marketing, Not Product Improvement
“It's just that the value proposition of publishing an eval is completely orthogonal to the value proposition of building evals in service of building a good product. I think the purpose of publishing benchmarks is marketing, and it's good marketing.”
Ankur Goyal Dec 7, 2025 ▶ 7:53 The Great Evals Debate — Ankur Goyal & Malte Ubl
Insight
Goyal: Providing eval criteria and examples is more effective than writing specs
“In many ways coming to the table of product building with representative examples and criteria that articulate what good versus bad is for a use case is just a more precise and usable form of product management than writing a spec.”
Ankur Goyal Dec 7, 2025 ▶ 18:36 The Great Evals Debate — Ankur Goyal & Malte Ubl
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.