Jun 10, 2025 · 33m · a16z

Giving New Life to Unstructured Data with LLMs and Agents

Anant Bhardwaj · 24m spoken Guido Appenzeller · 6m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this a16z podcast episode, Instabase Founder and CEO Anant Bhardwaj joins Partner Guido Appenzeller to discuss how compound AI systems, compile-time agentic workflows, and federated execution frameworks are solving the challenge of unstructured data and transforming enterprise automation beyond legacy RPA.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The host as informed peer 3.6 Guest teaching 4.0 Guest disagreement 1.1 The host pushing back 0.6
05100:0010:0020:0030:000:39–7:27 · The host as informed peer 3/10 Title Card: LLMs and Agents for Unstructured Data Anant details the evolution of unstructured data processing from MIT research to Instabase, BERT, and InstaLM. Guido demonstrates tech familiarity by mentioning two-dimensional rotary encoding and the bitter lesson.7:27–13:06 · The host as informed peer 3/10 Enterprise Use Cases and System Reliability Beyond LLMs Guido shares a personal anecdote about sorting scanned PDFs before Anant explains why standard RAG and LLMs fail on complex financial documents due to precision versus completeness trade-offs.13:06–15:52 · The host as informed peer 4/10 Predictability vs. Perfection and User Experience Shifts Guido asks how enterprise compliance acceptance criteria are shifting from absolute perfection to human error benchmarks. Anant gently reframes the issue, explaining that predictability of errors matters more than raw accuracy.15:52–21:24 · The host as informed peer 3/10 Conversational Lending and Barriers to Enterprise AI Adoption Anant shares examples of conversational lending over WhatsApp and notes enterprise slowness. Guido offers a mild counterpoint that enterprises are moving faster with AI than in past tech cycles.21:24–25:21 · The host as informed peer 4/10 AI Agents in the Enterprise: Compile-Time vs. Run-Time Anant introduces the distinction between build-time compile-time agentic generation and deterministic runtime execution. Guido synthesizes this with broader industry debates on full autonomy versus workflow freezing.25:21–28:05 · The host as informed peer 2/10 Future Vision: Federated AI Execution Frameworks Anant shares his long-term vision of federated AI execution frameworks replacing traditional robotic process automation (RPA) in enterprises.28:05–32:35 · The host as informed peer 5/10 End-to-End Workflow Automation and Model Context Protocol Anant outlines how Model Context Protocol and identity pass-through enable end-to-end workflow automation. Guido pushes back on full capability pass-through, arguing agent privileges should be capped like an intern.32:35–33:55 · The host as informed peer 5/10 Conclusion and Strategic Lessons for Enterprise Leaders Guido synthesizes strategic lessons for executives, comparing AI adoption to the dot-com era where delay risked obsolescence. Anant summarizes three core business outcomes.0:39–7:27 · Guest teaching 5/10 Title Card: LLMs and Agents for Unstructured Data Anant details the evolution of unstructured data processing from MIT research to Instabase, BERT, and InstaLM. Guido demonstrates tech familiarity by mentioning two-dimensional rotary encoding and the bitter lesson.7:27–13:06 · Guest teaching 5/10 Enterprise Use Cases and System Reliability Beyond LLMs Guido shares a personal anecdote about sorting scanned PDFs before Anant explains why standard RAG and LLMs fail on complex financial documents due to precision versus completeness trade-offs.13:06–15:52 · Guest teaching 4/10 Predictability vs. Perfection and User Experience Shifts Guido asks how enterprise compliance acceptance criteria are shifting from absolute perfection to human error benchmarks. Anant gently reframes the issue, explaining that predictability of errors matters more than raw accuracy.15:52–21:24 · Guest teaching 4/10 Conversational Lending and Barriers to Enterprise AI Adoption Anant shares examples of conversational lending over WhatsApp and notes enterprise slowness. Guido offers a mild counterpoint that enterprises are moving faster with AI than in past tech cycles.21:24–25:21 · Guest teaching 4/10 AI Agents in the Enterprise: Compile-Time vs. Run-Time Anant introduces the distinction between build-time compile-time agentic generation and deterministic runtime execution. Guido synthesizes this with broader industry debates on full autonomy versus workflow freezing.25:21–28:05 · Guest teaching 5/10 Future Vision: Federated AI Execution Frameworks Anant shares his long-term vision of federated AI execution frameworks replacing traditional robotic process automation (RPA) in enterprises.28:05–32:35 · Guest teaching 4/10 End-to-End Workflow Automation and Model Context Protocol Anant outlines how Model Context Protocol and identity pass-through enable end-to-end workflow automation. Guido pushes back on full capability pass-through, arguing agent privileges should be capped like an intern.32:35–33:55 · Guest teaching 1/10 Conclusion and Strategic Lessons for Enterprise Leaders Guido synthesizes strategic lessons for executives, comparing AI adoption to the dot-com era where delay risked obsolescence. Anant summarizes three core business outcomes.0:39–7:27 · Guest disagreement 1/10 Title Card: LLMs and Agents for Unstructured Data Anant details the evolution of unstructured data processing from MIT research to Instabase, BERT, and InstaLM. Guido demonstrates tech familiarity by mentioning two-dimensional rotary encoding and the bitter lesson.7:27–13:06 · Guest disagreement 1/10 Enterprise Use Cases and System Reliability Beyond LLMs Guido shares a personal anecdote about sorting scanned PDFs before Anant explains why standard RAG and LLMs fail on complex financial documents due to precision versus completeness trade-offs.13:06–15:52 · Guest disagreement 2/10 Predictability vs. Perfection and User Experience Shifts Guido asks how enterprise compliance acceptance criteria are shifting from absolute perfection to human error benchmarks. Anant gently reframes the issue, explaining that predictability of errors matters more than raw accuracy.15:52–21:24 · Guest disagreement 1/10 Conversational Lending and Barriers to Enterprise AI Adoption Anant shares examples of conversational lending over WhatsApp and notes enterprise slowness. Guido offers a mild counterpoint that enterprises are moving faster with AI than in past tech cycles.21:24–25:21 · Guest disagreement 1/10 AI Agents in the Enterprise: Compile-Time vs. Run-Time Anant introduces the distinction between build-time compile-time agentic generation and deterministic runtime execution. Guido synthesizes this with broader industry debates on full autonomy versus workflow freezing.25:21–28:05 · Guest disagreement 1/10 Future Vision: Federated AI Execution Frameworks Anant shares his long-term vision of federated AI execution frameworks replacing traditional robotic process automation (RPA) in enterprises.28:05–32:35 · Guest disagreement 2/10 End-to-End Workflow Automation and Model Context Protocol Anant outlines how Model Context Protocol and identity pass-through enable end-to-end workflow automation. Guido pushes back on full capability pass-through, arguing agent privileges should be capped like an intern.32:35–33:55 · Guest disagreement 0/10 Conclusion and Strategic Lessons for Enterprise Leaders Guido synthesizes strategic lessons for executives, comparing AI adoption to the dot-com era where delay risked obsolescence. Anant summarizes three core business outcomes.0:39–7:27 · The host pushing back 0/10 Title Card: LLMs and Agents for Unstructured Data Anant details the evolution of unstructured data processing from MIT research to Instabase, BERT, and InstaLM. Guido demonstrates tech familiarity by mentioning two-dimensional rotary encoding and the bitter lesson.7:27–13:06 · The host pushing back 0/10 Enterprise Use Cases and System Reliability Beyond LLMs Guido shares a personal anecdote about sorting scanned PDFs before Anant explains why standard RAG and LLMs fail on complex financial documents due to precision versus completeness trade-offs.13:06–15:52 · The host pushing back 1/10 Predictability vs. Perfection and User Experience Shifts Guido asks how enterprise compliance acceptance criteria are shifting from absolute perfection to human error benchmarks. Anant gently reframes the issue, explaining that predictability of errors matters more than raw accuracy.15:52–21:24 · The host pushing back 1/10 Conversational Lending and Barriers to Enterprise AI Adoption Anant shares examples of conversational lending over WhatsApp and notes enterprise slowness. Guido offers a mild counterpoint that enterprises are moving faster with AI than in past tech cycles.21:24–25:21 · The host pushing back 0/10 AI Agents in the Enterprise: Compile-Time vs. Run-Time Anant introduces the distinction between build-time compile-time agentic generation and deterministic runtime execution. Guido synthesizes this with broader industry debates on full autonomy versus workflow freezing.25:21–28:05 · The host pushing back 0/10 Future Vision: Federated AI Execution Frameworks Anant shares his long-term vision of federated AI execution frameworks replacing traditional robotic process automation (RPA) in enterprises.28:05–32:35 · The host pushing back 3/10 End-to-End Workflow Automation and Model Context Protocol Anant outlines how Model Context Protocol and identity pass-through enable end-to-end workflow automation. Guido pushes back on full capability pass-through, arguing agent privileges should be capped like an intern.32:35–33:55 · The host pushing back 0/10 Conclusion and Strategic Lessons for Enterprise Leaders Guido synthesizes strategic lessons for executives, comparing AI adoption to the dot-com era where delay risked obsolescence. Anant summarizes three core business outcomes.

speaking balance: gold is the host, purple is the guest (3 minute bins)

0:00 · the host 0% · guest 100%0:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%18:00 · the host 0% · guest 100%18:00 · the host 0% · guest 100%21:00 · the host 0% · guest 100%21:00 · the host 0% · guest 100%24:00 · the host 0% · guest 100%24:00 · the host 0% · guest 100%27:00 · the host 0% · guest 100%27:00 · the host 0% · guest 100%30:00 · the host 0% · guest 100%30:00 · the host 0% · guest 100%33:00 · the host 0% · guest 100%33:00 · the host 0% · guest 100%
Sharpest disagreement ▶ 13:49 Guest reframes accuracy vs predictability

Anant directly reframes Guido's question about error rate benchmarks, stating that enterprises do not care about raw accuracy as much as error predictability.

Hardest push from the host ▶ 31:34 Host challenges full identity pass-through for agents

Guido refuses the premise that agents should hold identical access rights to human users, using an intern spending limit analogy to advocate for strict privilege boundaries.

Biggest teaching moment ▶ 8:34 Guest breaks down failure modes of naive RAG on enterprise documents

Anant educates on why context windows and vector retrieval fail on complex financial packets due to missed table cells and lack of completeness guarantees.

The host holds their own ▶ 32:35 Host frames executive strategy using historical tech cycles

Guido takes control of the segment to deliver strategic guidance to enterprise leaders, drawing historical parallels to the dot-com revolution and Barnes & Noble.

the scores for every segment, with the reasoning behind each
ChapterTopicThe host as informed peerGuest teachingGuest disagreementThe host pushing backWhy
Title Card: LLMs and Agents for Unstructured Data 3510 Anant details the evolution of unstructured data processing from MIT research to Instabase, BERT, and InstaLM. Guido demonstrates tech familiarity by mentioning two-dimensional rotary encoding and the bitter lesson.
Enterprise Use Cases and System Reliability Beyond LLMs 3510 Guido shares a personal anecdote about sorting scanned PDFs before Anant explains why standard RAG and LLMs fail on complex financial documents due to precision versus completeness trade-offs.
Predictability vs. Perfection and User Experience Shifts 4421 Guido asks how enterprise compliance acceptance criteria are shifting from absolute perfection to human error benchmarks. Anant gently reframes the issue, explaining that predictability of errors matters more than raw accuracy.
Conversational Lending and Barriers to Enterprise AI Adoption 3411 Anant shares examples of conversational lending over WhatsApp and notes enterprise slowness. Guido offers a mild counterpoint that enterprises are moving faster with AI than in past tech cycles.
AI Agents in the Enterprise: Compile-Time vs. Run-Time 4410 Anant introduces the distinction between build-time compile-time agentic generation and deterministic runtime execution. Guido synthesizes this with broader industry debates on full autonomy versus workflow freezing.
Future Vision: Federated AI Execution Frameworks 2510 Anant shares his long-term vision of federated AI execution frameworks replacing traditional robotic process automation (RPA) in enterprises.
End-to-End Workflow Automation and Model Context Protocol 5423 Anant outlines how Model Context Protocol and identity pass-through enable end-to-end workflow automation. Guido pushes back on full capability pass-through, arguing agent privileges should be capped like an intern.
Conclusion and Strategic Lessons for Enterprise Leaders 5100 Guido synthesizes strategic lessons for executives, comparing AI adoption to the dot-com era where delay risked obsolescence. Anant summarizes three core business outcomes.

Statements from this episode (13)

Insight
Bhardwaj defines unstructured data as anything non-queryable via SQL
“Anything that cannot be put into nice database tables where you can run SQL, anything that is not that is unstructured data.”
Anant Bhardwaj Jun 10, 2025 ▶ 1:00
Assertion Not checkable as stated
Instabase tripled its revenue between 2021 and 2022
“We tripled our revenue that year, twenty-twenty-one to twenty-twenty-two.”
Anant Bhardwaj Jun 10, 2025 ▶ 6:39
Assertion Not checkable as stated
Bhardwaj: Automated systems process bank lending applications in under five seconds
“So now you can do lending in like less than five seconds. Rather than earlier, that would have taken several, several weeks.”
Anant Bhardwaj Jun 10, 2025 ▶ 11:02
Insight
Bhardwaj: RAG alone is insufficient for high-accuracy enterprise data processing
“While drag is good for casual search, you need a complex workflow under the herd that is explainable, that is auditable, that is guaranteed to be accurate and correct, is important for solving many of these enterprise problems.”
Anant Bhardwaj Jun 10, 2025 ▶ 12:00
Prediction Not checkable as stated
Bhardwaj: Major capital will flow into building reliability systems around LLMs
“And that is going to be a lot of investment that you will see across the board, which is how do we build the right systems around AI and LLMs that solves the problem.”
Anant Bhardwaj Jun 10, 2025 ▶ 12:57
Insight
Bhardwaj: Enterprise AI clients prioritize predictability over baseline accuracy
“So I think in general what we have seen is enterprises are fine using AIs as long as we show them predictability. They don't care about, you know, 99% accuracy. You can be 90% accurate or even 80% accurate, But just tell us which 20% need to be reviewed or whi…”
Anant Bhardwaj Jun 10, 2025 ▶ 14:37
Prediction Not checkable as stated
Bhardwaj: AI will automate filtering unstructured documents down to key reviewable items
“Whenever unstructured data like documents come in, humans will still see some kind of dashboard with like whatever stuff is, and only the thing of interest they will go and double click on. And AI will do a lot of things to minimize their time to get to that t…”
Anant Bhardwaj Jun 10, 2025 ▶ 15:19
Assertion Not checkable as stated
Bhardwaj: Indian bank conducts entire lending process conversationally over WhatsApp
“I was working with a bank in India, and now given that AI is, has become reasonably reliable, they are offering entire lending over WhatsApp.”
Anant Bhardwaj Jun 10, 2025 ▶ 16:20
Insight
Bhardwaj: Enterprise AI adoption hinges on data security and auditability
“Two key things that they care about is, how do you guarantee that my data is safe and secure? So that's number one. And second is, how do you give me auditability and predictability? That's the two more, like if you boil down to all their questions, they event…”
Anant Bhardwaj Jun 10, 2025 ▶ 20:15
Insight
Bhardwaj: Enterprises reject black-box AI decisions without clear step-by-step explainability
“Nobody wants, like, AI making a decision, even if it is correct, if they cannot Explain here are the set of steps that it took, because if something wrong happened, they have to explain, because in human world you can explain.”
Anant Bhardwaj Jun 10, 2025 ▶ 20:34
Prediction Not checkable as stated
Bhardwaj: Enterprise AI agents will operate at compile-time, not runtime
“So I do not believe that autonomous agent would be a runtime phenomena. However, there would be a build time or compile time phenomena.”
Anant Bhardwaj Jun 10, 2025 ▶ 23:16
Prediction Not checkable as stated
Anant Bhardwaj: AI automation will fully replace legacy RPA
“The bet that we are taking is that AI will drive automation in a significant way. RPA would be fully eaten by AI automation, and the future is likely going to be more of decentralized, federated execution.”
Anant Bhardwaj Jun 10, 2025 ▶ 27:42
Insight
Bhardwaj: RPA cannot effectively process unstructured data due to variability
“You can't do robotic process for unstructured data because it's not it's not fixed. They change it. So anything will be very, very brittle.”
Anant Bhardwaj Jun 10, 2025 ▶ 29:36
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.