Insight certainty 4/5 debate potential 2/5

Shankar: Evals are necessary to train AI reasoning models

Shreya Shankar · [Lightning Pod] Evals: How to Improve AI Consistently — with Hamel Husain and Shreya Shankar · Mar 13, 2025 · at 27:13

ML researcher Shreya Shankar notes the dependency between evaluation frameworks and training reasoning models during a discussion on verifiers with Swix.

0:00 / 0:02exact quote · 2.3s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“You need evals to train your reasoning models.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Shreya Shankar

Insight
Shankar: AI evals differ from MLOps due to data scarcity
“The other thing is I think that AI engineering Evaluation or evals here is actually different from MLOps or ML evaluation for traditional ML models. We were in a much more, you know, data rich setting in MLOps. So we were taught to come up with loss metrics or…”
Shreya Shankar Mar 13, 2025 ▶ 3:20 [Lightning Pod] Evals: How to Improve AI Consistently — with Hamel Husain and Shreya Shankar
Insight
Shankar: Decompose LLM pipelines into unit tasks with standalone intermediate assertions
“The idea is to have each node in your graph kind of be a standalone, do a standalone thing that you can have standalone assertions for. And if you think about, you know, infinitely many inputs flowing through your pipeline, there's going to be some fraction of…”
Shreya Shankar Sep 28, 2024 ▶ 57:35 [Paper Club] Who Validates the Validators? Aligning LLM-Judges with Humans (w/ Eugene Yan)
Insight
Shankar: LLM failure modes and evaluation techniques have largely stabilized
“I think techniques have stabilized. I think the kinds of failure modes of LLMs, I mean, they're still there, but it's not like changing every single day. We know that LLMs are bad at certain things. We know a little bit more about say limitations of the transf…”
Shreya Shankar Mar 13, 2025 ▶ 11:21 [Lightning Pod] Evals: How to Improve AI Consistently — with Hamel Husain and Shreya Shankar
Insight
Shankar: Grounded synthetic data generation beats slow human annotation
“People don't know how to do anything other than Plan A, which is to go out and try to collect as much real world data as possible and take months because we're going to employ, like, human annotator teams to do this. Or Plan B, which is I'm going to ask an LLM…”
Shreya Shankar Mar 13, 2025 ▶ 13:48 [Lightning Pod] Evals: How to Improve AI Consistently — with Hamel Husain and Shreya Shankar
Insight
Shreya Shankar: Chunking efficacy is task-specific, requiring automated pipeline optimization
“Sometimes it's beneficial to chunk and sometimes you should not chunk. And we have observed this in a number of workloads and the insight that we've gained is that we will never know what it's all task specific and data specific. And we are so further convince…”
Shreya Shankar Nov 29, 2024 ▶ 19:31 [Paper Club] DocETL: Agentic Query Rewriting + Eval for Complex Document Processing w Shreya Shankar
Insight
Shreya Shankar: DocETL builds semantic unstructured layers, not point-lookup RAG systems
“This is very different from traditional rag or Q&A or document processing for a chatbot. Like, the kinds of queries that people are, people want to use .etl for can be expressed as etl style sweep and harvest, kind of, I want to look at my entire dataset. I wo…”
Shreya Shankar Nov 29, 2024 ▶ 47:13 [Paper Club] DocETL: Agentic Query Rewriting + Eval for Complex Document Processing w Shreya Shankar
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.