Insight certainty 4/5 debate potential 2/5

Agarwal: Rigorous eval suites are core IP for leading AI startups

Anish Agarwal · ⚡️Traversal: Causal ML and Reinforcement Learning · Oct 5, 2025 · at 35:31

Anish Agarwal (co-founder of Traversal) discusses why AI companies must invest heavily in proprietary benchmarking and evaluation pipelines.

0:00 / 0:30exact quote · 30.5s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“The best AI companies will always have to be the edge of what the models can do, right? I think you always want to be threading the line. If everything works all the time, then you're not really pushing the limit and you're not innovating, right? So I think you always kind of want to be pushing. And so if you're always pushing this, sometimes your system will work, sometimes it won't work. And so then evaluation I would say it's become sometimes a big bottleneck. You're spending the most engineering hours just evaluating. And so the, I think the best companies will be investing a lot of time in really good eval. And I think that's core IP.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Anish Agarwal

Opinion
Agarwal: No AI incident troubleshooting competitor genuinely works in production
“And I don't think we've seen any other company in our space having something actually work at production.”
Anish Agarwal Oct 5, 2025 ▶ 25:49 ⚡️Traversal: Causal ML and Reinforcement Learning
Prediction Not checkable as stated
Agarwal: AI coding assistants will cause uninterpretable outages and endless firefighting
“And so it's pretty clear to us that, you know, and we were starting to use it ourselves and sometimes we didn't understand what the code was doing, but you know, we shipped it. And so it's like, well, if this is clearly, this is going to happen a lot more. And…”
Anish Agarwal Oct 5, 2025 ▶ 7:40 ⚡️Traversal: Causal ML and Reinforcement Learning
Insight
Agarwal: Incident troubleshooting requires adaptive search over LLM context dumping
“You cannot just Put all of it into context of an LLM and hope something great happens. You have to search the data sequentially and adaptively, right? And that's what these agentic systems are fundamentally about.”
Anish Agarwal Oct 5, 2025 ▶ 7:07 ⚡️Traversal: Causal ML and Reinforcement Learning
Insight
Agarwal: LLMs are really bad at processing time series data
“Because most of the data you're looking at is like time series data. And these LLMs are really bad at processing time series data, right? And that's really where like good statistics comes in.”
Anish Agarwal Oct 5, 2025 ▶ 17:58 ⚡️Traversal: Causal ML and Reinforcement Learning
Insight
Agarwal: Observability incumbents lack incentive to analyze competitor telemetry data
“The typically the way they work is, is they price based on the amount of data they're storing, right? And so, you know, they have very little incentive for company A To provide you any insight on data being stored on, on company, observability company B, right…”
Anish Agarwal Oct 5, 2025 ▶ 23:15 ⚡️Traversal: Causal ML and Reinforcement Learning
Disclosure
Agarwal: Sequoia met 20 AI SRE startups before backing Traversal
“Our first VC backer was Sequoia, and I think they had met, I think, like, I think like 20 VC, 20, ah, companies before this, but unbeknownst to us, and that was the first question, like Bogomol, who's a board member from there asked, like, I've heard this pitc…”
Anish Agarwal Oct 5, 2025 ▶ 8:34 ⚡️Traversal: Causal ML and Reinforcement Learning
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.