Shreya Shankar

Assistant Professor, Carnegie Mellon University · 3 appearances on the record.

computed by AI from the episodes · how this works → · full disclaimer →

academicscientistengineerauthor@sh_reya ↗LinkedIn ↗sh-reya.com ↗

Shankar created DocETL, an open-source declarative system that utilizes agentic query rewriting and language models to analyze complex unstructured data at scale. Alongside Hamel Husain, she co-instructs a course on AI evals and co-authored the book Evals for AI Engineers.

10statements → 0claims → 0claims resolved → 3.8/5average certainty → 1.9/5average debate potential → 1said about them ↓

8 insights · 1 disclosure · 1 what if · every statement was checked. none of them is a claim the record can settle: opinions, insights and disclosures never carry an assessment.

Everything Shreya Shankar said on Latent Space that made the record, most notable first. Filter by type, assessment or year in the ledger →

Insight
Shankar: AI evals differ from MLOps due to data scarcity
“The other thing is I think that AI engineering Evaluation or evals here is actually different from MLOps or ML evaluation for traditional ML models. We were in a much more, you know, data rich setting in MLOps. So we were taught to come up with loss metrics or…”
Shreya Shankar Mar 13, 2025 ▶ 3:20 [Lightning Pod] Evals: How to Improve AI Consistently — with Hamel Husain and Shreya Shankar
Insight
Shankar: Decompose LLM pipelines into unit tasks with standalone intermediate assertions
“The idea is to have each node in your graph kind of be a standalone, do a standalone thing that you can have standalone assertions for. And if you think about, you know, infinitely many inputs flowing through your pipeline, there's going to be some fraction of…”
Shreya Shankar Sep 28, 2024 ▶ 57:35 [Paper Club] Who Validates the Validators? Aligning LLM-Judges with Humans (w/ Eugene Yan)
Insight
Shankar: LLM failure modes and evaluation techniques have largely stabilized
“I think techniques have stabilized. I think the kinds of failure modes of LLMs, I mean, they're still there, but it's not like changing every single day. We know that LLMs are bad at certain things. We know a little bit more about say limitations of the transf…”
Shreya Shankar Mar 13, 2025 ▶ 11:21 [Lightning Pod] Evals: How to Improve AI Consistently — with Hamel Husain and Shreya Shankar
Insight
Shankar: Grounded synthetic data generation beats slow human annotation
“People don't know how to do anything other than Plan A, which is to go out and try to collect as much real world data as possible and take months because we're going to employ, like, human annotator teams to do this. Or Plan B, which is I'm going to ask an LLM…”
Shreya Shankar Mar 13, 2025 ▶ 13:48 [Lightning Pod] Evals: How to Improve AI Consistently — with Hamel Husain and Shreya Shankar
Insight
Shankar: Evals are necessary to train AI reasoning models
“You need evals to train your reasoning models.”
Shreya Shankar Mar 13, 2025 ▶ 27:13 [Lightning Pod] Evals: How to Improve AI Consistently — with Hamel Husain and Shreya Shankar
Insight
Shreya Shankar: Chunking efficacy is task-specific, requiring automated pipeline optimization
“Sometimes it's beneficial to chunk and sometimes you should not chunk. And we have observed this in a number of workloads and the insight that we've gained is that we will never know what it's all task specific and data specific. And we are so further convince…”
Shreya Shankar Nov 29, 2024 ▶ 19:31 [Paper Club] DocETL: Agentic Query Rewriting + Eval for Complex Document Processing w Shreya Shankar
Insight
Shreya Shankar: DocETL builds semantic unstructured layers, not point-lookup RAG systems
“This is very different from traditional rag or Q&A or document processing for a chatbot. Like, the kinds of queries that people are, people want to use .etl for can be expressed as etl style sweep and harvest, kind of, I want to look at my entire dataset. I wo…”
Shreya Shankar Nov 29, 2024 ▶ 47:13 [Paper Club] DocETL: Agentic Query Rewriting + Eval for Complex Document Processing w Shreya Shankar
Insight
Shreya Shankar: Practical DocETL users never have pre-annotated ground truth datasets
“So we're redoing our evaluation to be on data sets where we actually have ground truth from human annotators, but that's just not a practical setting. Like, nobody's coming to doc ETL with the ground truth.”
Shreya Shankar Nov 29, 2024 ▶ 49:36 [Paper Club] DocETL: Agentic Query Rewriting + Eval for Complex Document Processing w Shreya Shankar
What-if
Shreya Shankar: GPT-4o Mini reduces DocETL optimization cost by 90%
“The reason it was a hundred dollars, if I ran the optimizer with GPT-Foro mini as the LLMs, it would be 10 dollars. But we use GPT four. Oh, just because I think we did this at a time where many hadn't come out yet.”
Shreya Shankar Nov 29, 2024 ▶ 45:58 [Paper Club] DocETL: Agentic Query Rewriting + Eval for Complex Document Processing w Shreya Shankar
Disclosure
Shankar: New LLM evaluation framework to be open-sourced in ChainForge
“Not yet. It's not out yet, but we will. So the conference is in three weeks. We have to have it out by then, but it'll be implemented in chainforge.ai, which is an open source LLM pipeline building tool.”
Shreya Shankar Sep 28, 2024 ▶ 56:49 [Paper Club] Who Validates the Validators? Aligning LLM-Judges with Humans (w/ Eugene Yan)

The other half of the tape: Shreya Shankar's own voice is left out of every number here. 1 statement on the record names them. every mention, with the transcript →

Statements about Shreya Shankar, by other people (1)

Prediction Not checkable as stated
Husain: HCI and workflow evaluations will enter AI tools by 2027
“If I were to fast forward one or two years, I would expect to see those and all the tools. It just hasn't arrived yet.”
Hamel Husain Mar 13, 2025 ▶ 26:15 [Lightning Pod] Evals: How to Improve AI Consistently — with Hamel Husain and Shreya Shankar

Appearances (3)

EpisodeDateSpeaking time
[Lightning Pod] Evals: How to Improve AI Consistently — with Hamel Husain and Shreya Shank Mar 13, 2025 5m
[Paper Club] DocETL: Agentic Query Rewriting + Eval for Complex Document Processing w Shre Nov 29, 2024 10m
[Paper Club] Who Validates the Validators? Aligning LLM-Judges with Humans (w/ Eugene Yan) Sep 28, 2024 1m
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.