Apr 24, 2024 · 1h 2m · latent-space

High Agency Pydantic over VC Backed Frameworks — with Jason Liu of Instructor

Jason Liu · 37m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of Latent Space, Jason Liu, creator of Instructor, joins Alessio and Swix to discuss the mechanics of structured LLM outputs, the architectural advantages of pragmatic engineering over VC-backed frameworks, and the emerging discipline of AI engineering.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 5.6 Guest teaching 4.8 Guest disagreement 3.2 The hosts pushing back 1.9
05100:0015:0030:0045:001:00:001:20–5:04 · The hosts as informed peer 6/10 Stitch Fix Experience and Early LLM Skepticism Swix connects RAG to classical recommendation systems (RecSys) and probes why Jason built the Flight recommendation system from scratch at Stitch Fix. Jason explains the operational necessity of standardizing bespoke data science codebases to enable observability and performance tuning.5:05–7:16 · The hosts as informed peer 6/10 Similarity Search, Multimodal Embeddings, and Real-World Impact Swix asks technical clarifying questions about whether off-the-shelf GPT-3 embeddings worked for multi-modal clothing retrieval. Jason clarifies that fine-tuning ResNets and joint text-image representations was necessary to beat generic models at scale.7:16–10:02 · The hosts as informed peer 3/10 Fashion, Human Narrative, and Recommender Limitations Swix jokes about tech fashion and asks whether tech workers should use Stitch Fix. Jason reframes the premise by contrasting utility-driven clothing delivery with human narrative and experiential shopping.10:02–13:24 · The hosts as informed peer 6/10 The Inception and Architectural Philosophy of Instructor Swix and Jason discuss library architecture, contrasting broad frameworks like Django with lightweight standard utilities like Requests. Jason details how function calling separated schemas from prompts, driving the inception of Instructor.13:24–16:35 · The hosts as informed peer 7/10 Typed Outputs, JSON Mode, and Function Calling Mechanics Alessio asks technical questions comparing JSON mode to function calling and inquires about SDK-level handling of parallel tool invocation. Jason explains why typed schemas with validation outperform raw JSON output modes.16:36–20:40 · The hosts as informed peer 8/10 Tool Selection, Context Scaling, and Ranking Architectures Alessio pushes back on relying solely on vector similarity for tool selection, sharing firsthand incubation data on 780 API endpoints where LLMs select tools better than semantic search. Jason defends ranking via XGBoost over stuffing prompts with tool definitions.20:41–23:23 · The hosts as informed peer 6/10 Production Benchmarks and Anthropic vs OpenAI Trade-offs Alessio notes Jason's contrarian evaluation of Anthropic function calling benchmarks. Jason shares real production telemetry, revealing edge-case parsing bugs despite compelling pricing advantages.23:23–26:41 · The hosts as informed peer 5/10 Mapping the Use Cases and Surface Area of Instructor Swix invites Jason to delineate the capabilities of Instructor across graphs, sanitation, and query understanding. Jason schools the audience on why vector embeddings fail on temporal aggregations and relational queries.26:42–28:55 · The hosts as informed peer 6/10 Workflows and Deterministic DAGs versus ReAct Agent Loops Swix asks how Instructor compares with LangChain's self-querying and prompts Jason on workflow design. Jason advocates for deterministic DAGs and fine-tuning compilers over non-deterministic ReAct prompting loops.28:56–31:00 · The hosts as informed peer 6/10 Autonomous Agents, Execution Horizons, and Plan Churn Swix discusses evaluating long-horizon coding agents like Devin and cautions against rewarding runtime duration. Jason mocks early agent failure modes and argues for monotonic progress metrics over loop thrashing.31:00–33:40 · The hosts as informed peer 6/10 AI Engineering Stack and Pragmatic Infrastructure Swix inquires about Jason's preferred AI engineering stack. Jason dismisses the dozens of VC-backed LLM observability startups, explaining that a single Postgres table and standard APM tools like Datadog and Sentry are more than sufficient.33:41–37:10 · The hosts as informed peer 5/10 Choosing Independent Consulting over Venture Capital Backing Swix asks why Jason declined venture backing for Instructor. Jason explains his philosophy of building a standard utility like Requests while running a high-margin, flexible solo consulting practice.37:10–41:42 · The hosts as informed peer 6/10 Startup Realities, Solo Founding, and Entrepreneurial Courage Alessio and Swix discuss the realities of seed funding versus high-paying corporate employment. Jason highlights founder breakup rates and emphasizes having the courage to build with minimal overhead.41:42–43:59 · The hosts as informed peer 4/10 Overcoming Career Setbacks and Starting from Zero Alessio asks about overcoming early rejection after Jason was turned down by top AI labs in early 2023. Jason uses a Jiu-Jitsu analogy to urge engineers to start immediately rather than waiting for ideal conditions.43:59–49:10 · The hosts as informed peer 6/10 High Agency, Process Metrics, and the Clay Metaphor Jason defines agency as focusing on process volume rather than final outcome metrics using a pottery clay metaphor. Swix directly pushes back, asking whether this equates to flawed developer metrics like counting lines of code.49:10–54:40 · The hosts as informed peer 5/10 Automating Personal Workflows and Future AI Capabilities Alessio and Jason discuss personal workflow automation. Jason explains how hiring human assistants helped him uncover precise, modular heuristic prompts for calendar management and client summaries.54:40–57:22 · The hosts as informed peer 6/10 DSPy, Prompt Optimization, and Business-Level Metrics Swix raises DSPy and the thesis that AI is a better prompt optimizer than humans. Jason argues that DSPy works well on narrow programmatic heuristics but breaks down when evaluating long-term business outcomes like user churn.57:22–1:01:36 · The hosts as informed peer 6/10 AI Engineers vs ML Engineers and Organizational Talent Swix asks about the distinction between MLEs and AI engineers. Jason delivers a strong take advising application startups against hiring traditional MLEs, who quickly churn when faced with TypeScript and missing training pipelines.1:01:36–1:02:33 · The hosts as informed peer 4/10 AI Engineer World's Fair Preview and Episode Conclusion Swix previews the AI Engineer World's Fair. Jason jokes about his talk title while reinforcing the core philosophy behind Pydantic-based structured outputs.1:20–5:04 · Guest teaching 5/10 Stitch Fix Experience and Early LLM Skepticism Swix connects RAG to classical recommendation systems (RecSys) and probes why Jason built the Flight recommendation system from scratch at Stitch Fix. Jason explains the operational necessity of standardizing bespoke data science codebases to enable observability and performance tuning.5:05–7:16 · Guest teaching 5/10 Similarity Search, Multimodal Embeddings, and Real-World Impact Swix asks technical clarifying questions about whether off-the-shelf GPT-3 embeddings worked for multi-modal clothing retrieval. Jason clarifies that fine-tuning ResNets and joint text-image representations was necessary to beat generic models at scale.7:16–10:02 · Guest teaching 6/10 Fashion, Human Narrative, and Recommender Limitations Swix jokes about tech fashion and asks whether tech workers should use Stitch Fix. Jason reframes the premise by contrasting utility-driven clothing delivery with human narrative and experiential shopping.10:02–13:24 · Guest teaching 6/10 The Inception and Architectural Philosophy of Instructor Swix and Jason discuss library architecture, contrasting broad frameworks like Django with lightweight standard utilities like Requests. Jason details how function calling separated schemas from prompts, driving the inception of Instructor.13:24–16:35 · Guest teaching 5/10 Typed Outputs, JSON Mode, and Function Calling Mechanics Alessio asks technical questions comparing JSON mode to function calling and inquires about SDK-level handling of parallel tool invocation. Jason explains why typed schemas with validation outperform raw JSON output modes.16:36–20:40 · Guest teaching 4/10 Tool Selection, Context Scaling, and Ranking Architectures Alessio pushes back on relying solely on vector similarity for tool selection, sharing firsthand incubation data on 780 API endpoints where LLMs select tools better than semantic search. Jason defends ranking via XGBoost over stuffing prompts with tool definitions.20:41–23:23 · Guest teaching 5/10 Production Benchmarks and Anthropic vs OpenAI Trade-offs Alessio notes Jason's contrarian evaluation of Anthropic function calling benchmarks. Jason shares real production telemetry, revealing edge-case parsing bugs despite compelling pricing advantages.23:23–26:41 · Guest teaching 8/10 Mapping the Use Cases and Surface Area of Instructor Swix invites Jason to delineate the capabilities of Instructor across graphs, sanitation, and query understanding. Jason schools the audience on why vector embeddings fail on temporal aggregations and relational queries.26:42–28:55 · Guest teaching 5/10 Workflows and Deterministic DAGs versus ReAct Agent Loops Swix asks how Instructor compares with LangChain's self-querying and prompts Jason on workflow design. Jason advocates for deterministic DAGs and fine-tuning compilers over non-deterministic ReAct prompting loops.28:56–31:00 · Guest teaching 4/10 Autonomous Agents, Execution Horizons, and Plan Churn Swix discusses evaluating long-horizon coding agents like Devin and cautions against rewarding runtime duration. Jason mocks early agent failure modes and argues for monotonic progress metrics over loop thrashing.31:00–33:40 · Guest teaching 4/10 AI Engineering Stack and Pragmatic Infrastructure Swix inquires about Jason's preferred AI engineering stack. Jason dismisses the dozens of VC-backed LLM observability startups, explaining that a single Postgres table and standard APM tools like Datadog and Sentry are more than sufficient.33:41–37:10 · Guest teaching 4/10 Choosing Independent Consulting over Venture Capital Backing Swix asks why Jason declined venture backing for Instructor. Jason explains his philosophy of building a standard utility like Requests while running a high-margin, flexible solo consulting practice.37:10–41:42 · Guest teaching 4/10 Startup Realities, Solo Founding, and Entrepreneurial Courage Alessio and Swix discuss the realities of seed funding versus high-paying corporate employment. Jason highlights founder breakup rates and emphasizes having the courage to build with minimal overhead.41:42–43:59 · Guest teaching 4/10 Overcoming Career Setbacks and Starting from Zero Alessio asks about overcoming early rejection after Jason was turned down by top AI labs in early 2023. Jason uses a Jiu-Jitsu analogy to urge engineers to start immediately rather than waiting for ideal conditions.43:59–49:10 · Guest teaching 6/10 High Agency, Process Metrics, and the Clay Metaphor Jason defines agency as focusing on process volume rather than final outcome metrics using a pottery clay metaphor. Swix directly pushes back, asking whether this equates to flawed developer metrics like counting lines of code.49:10–54:40 · Guest teaching 4/10 Automating Personal Workflows and Future AI Capabilities Alessio and Jason discuss personal workflow automation. Jason explains how hiring human assistants helped him uncover precise, modular heuristic prompts for calendar management and client summaries.54:40–57:22 · Guest teaching 6/10 DSPy, Prompt Optimization, and Business-Level Metrics Swix raises DSPy and the thesis that AI is a better prompt optimizer than humans. Jason argues that DSPy works well on narrow programmatic heuristics but breaks down when evaluating long-term business outcomes like user churn.57:22–1:01:36 · Guest teaching 5/10 AI Engineers vs ML Engineers and Organizational Talent Swix asks about the distinction between MLEs and AI engineers. Jason delivers a strong take advising application startups against hiring traditional MLEs, who quickly churn when faced with TypeScript and missing training pipelines.1:01:36–1:02:33 · Guest teaching 2/10 AI Engineer World's Fair Preview and Episode Conclusion Swix previews the AI Engineer World's Fair. Jason jokes about his talk title while reinforcing the core philosophy behind Pydantic-based structured outputs.1:20–5:04 · Guest disagreement 2/10 Stitch Fix Experience and Early LLM Skepticism Swix connects RAG to classical recommendation systems (RecSys) and probes why Jason built the Flight recommendation system from scratch at Stitch Fix. Jason explains the operational necessity of standardizing bespoke data science codebases to enable observability and performance tuning.5:05–7:16 · Guest disagreement 1/10 Similarity Search, Multimodal Embeddings, and Real-World Impact Swix asks technical clarifying questions about whether off-the-shelf GPT-3 embeddings worked for multi-modal clothing retrieval. Jason clarifies that fine-tuning ResNets and joint text-image representations was necessary to beat generic models at scale.7:16–10:02 · Guest disagreement 3/10 Fashion, Human Narrative, and Recommender Limitations Swix jokes about tech fashion and asks whether tech workers should use Stitch Fix. Jason reframes the premise by contrasting utility-driven clothing delivery with human narrative and experiential shopping.10:02–13:24 · Guest disagreement 3/10 The Inception and Architectural Philosophy of Instructor Swix and Jason discuss library architecture, contrasting broad frameworks like Django with lightweight standard utilities like Requests. Jason details how function calling separated schemas from prompts, driving the inception of Instructor.13:24–16:35 · Guest disagreement 3/10 Typed Outputs, JSON Mode, and Function Calling Mechanics Alessio asks technical questions comparing JSON mode to function calling and inquires about SDK-level handling of parallel tool invocation. Jason explains why typed schemas with validation outperform raw JSON output modes.16:36–20:40 · Guest disagreement 3/10 Tool Selection, Context Scaling, and Ranking Architectures Alessio pushes back on relying solely on vector similarity for tool selection, sharing firsthand incubation data on 780 API endpoints where LLMs select tools better than semantic search. Jason defends ranking via XGBoost over stuffing prompts with tool definitions.20:41–23:23 · Guest disagreement 4/10 Production Benchmarks and Anthropic vs OpenAI Trade-offs Alessio notes Jason's contrarian evaluation of Anthropic function calling benchmarks. Jason shares real production telemetry, revealing edge-case parsing bugs despite compelling pricing advantages.23:23–26:41 · Guest disagreement 4/10 Mapping the Use Cases and Surface Area of Instructor Swix invites Jason to delineate the capabilities of Instructor across graphs, sanitation, and query understanding. Jason schools the audience on why vector embeddings fail on temporal aggregations and relational queries.26:42–28:55 · Guest disagreement 4/10 Workflows and Deterministic DAGs versus ReAct Agent Loops Swix asks how Instructor compares with LangChain's self-querying and prompts Jason on workflow design. Jason advocates for deterministic DAGs and fine-tuning compilers over non-deterministic ReAct prompting loops.28:56–31:00 · Guest disagreement 4/10 Autonomous Agents, Execution Horizons, and Plan Churn Swix discusses evaluating long-horizon coding agents like Devin and cautions against rewarding runtime duration. Jason mocks early agent failure modes and argues for monotonic progress metrics over loop thrashing.31:00–33:40 · Guest disagreement 5/10 AI Engineering Stack and Pragmatic Infrastructure Swix inquires about Jason's preferred AI engineering stack. Jason dismisses the dozens of VC-backed LLM observability startups, explaining that a single Postgres table and standard APM tools like Datadog and Sentry are more than sufficient.33:41–37:10 · Guest disagreement 4/10 Choosing Independent Consulting over Venture Capital Backing Swix asks why Jason declined venture backing for Instructor. Jason explains his philosophy of building a standard utility like Requests while running a high-margin, flexible solo consulting practice.37:10–41:42 · Guest disagreement 3/10 Startup Realities, Solo Founding, and Entrepreneurial Courage Alessio and Swix discuss the realities of seed funding versus high-paying corporate employment. Jason highlights founder breakup rates and emphasizes having the courage to build with minimal overhead.41:42–43:59 · Guest disagreement 2/10 Overcoming Career Setbacks and Starting from Zero Alessio asks about overcoming early rejection after Jason was turned down by top AI labs in early 2023. Jason uses a Jiu-Jitsu analogy to urge engineers to start immediately rather than waiting for ideal conditions.43:59–49:10 · Guest disagreement 4/10 High Agency, Process Metrics, and the Clay Metaphor Jason defines agency as focusing on process volume rather than final outcome metrics using a pottery clay metaphor. Swix directly pushes back, asking whether this equates to flawed developer metrics like counting lines of code.49:10–54:40 · Guest disagreement 2/10 Automating Personal Workflows and Future AI Capabilities Alessio and Jason discuss personal workflow automation. Jason explains how hiring human assistants helped him uncover precise, modular heuristic prompts for calendar management and client summaries.54:40–57:22 · Guest disagreement 4/10 DSPy, Prompt Optimization, and Business-Level Metrics Swix raises DSPy and the thesis that AI is a better prompt optimizer than humans. Jason argues that DSPy works well on narrow programmatic heuristics but breaks down when evaluating long-term business outcomes like user churn.57:22–1:01:36 · Guest disagreement 5/10 AI Engineers vs ML Engineers and Organizational Talent Swix asks about the distinction between MLEs and AI engineers. Jason delivers a strong take advising application startups against hiring traditional MLEs, who quickly churn when faced with TypeScript and missing training pipelines.1:01:36–1:02:33 · Guest disagreement 1/10 AI Engineer World's Fair Preview and Episode Conclusion Swix previews the AI Engineer World's Fair. Jason jokes about his talk title while reinforcing the core philosophy behind Pydantic-based structured outputs.1:20–5:04 · The hosts pushing back 2/10 Stitch Fix Experience and Early LLM Skepticism Swix connects RAG to classical recommendation systems (RecSys) and probes why Jason built the Flight recommendation system from scratch at Stitch Fix. Jason explains the operational necessity of standardizing bespoke data science codebases to enable observability and performance tuning.5:05–7:16 · The hosts pushing back 2/10 Similarity Search, Multimodal Embeddings, and Real-World Impact Swix asks technical clarifying questions about whether off-the-shelf GPT-3 embeddings worked for multi-modal clothing retrieval. Jason clarifies that fine-tuning ResNets and joint text-image representations was necessary to beat generic models at scale.7:16–10:02 · The hosts pushing back 1/10 Fashion, Human Narrative, and Recommender Limitations Swix jokes about tech fashion and asks whether tech workers should use Stitch Fix. Jason reframes the premise by contrasting utility-driven clothing delivery with human narrative and experiential shopping.10:02–13:24 · The hosts pushing back 1/10 The Inception and Architectural Philosophy of Instructor Swix and Jason discuss library architecture, contrasting broad frameworks like Django with lightweight standard utilities like Requests. Jason details how function calling separated schemas from prompts, driving the inception of Instructor.13:24–16:35 · The hosts pushing back 2/10 Typed Outputs, JSON Mode, and Function Calling Mechanics Alessio asks technical questions comparing JSON mode to function calling and inquires about SDK-level handling of parallel tool invocation. Jason explains why typed schemas with validation outperform raw JSON output modes.16:36–20:40 · The hosts pushing back 5/10 Tool Selection, Context Scaling, and Ranking Architectures Alessio pushes back on relying solely on vector similarity for tool selection, sharing firsthand incubation data on 780 API endpoints where LLMs select tools better than semantic search. Jason defends ranking via XGBoost over stuffing prompts with tool definitions.20:41–23:23 · The hosts pushing back 1/10 Production Benchmarks and Anthropic vs OpenAI Trade-offs Alessio notes Jason's contrarian evaluation of Anthropic function calling benchmarks. Jason shares real production telemetry, revealing edge-case parsing bugs despite compelling pricing advantages.23:23–26:41 · The hosts pushing back 1/10 Mapping the Use Cases and Surface Area of Instructor Swix invites Jason to delineate the capabilities of Instructor across graphs, sanitation, and query understanding. Jason schools the audience on why vector embeddings fail on temporal aggregations and relational queries.26:42–28:55 · The hosts pushing back 2/10 Workflows and Deterministic DAGs versus ReAct Agent Loops Swix asks how Instructor compares with LangChain's self-querying and prompts Jason on workflow design. Jason advocates for deterministic DAGs and fine-tuning compilers over non-deterministic ReAct prompting loops.28:56–31:00 · The hosts pushing back 3/10 Autonomous Agents, Execution Horizons, and Plan Churn Swix discusses evaluating long-horizon coding agents like Devin and cautions against rewarding runtime duration. Jason mocks early agent failure modes and argues for monotonic progress metrics over loop thrashing.31:00–33:40 · The hosts pushing back 2/10 AI Engineering Stack and Pragmatic Infrastructure Swix inquires about Jason's preferred AI engineering stack. Jason dismisses the dozens of VC-backed LLM observability startups, explaining that a single Postgres table and standard APM tools like Datadog and Sentry are more than sufficient.33:41–37:10 · The hosts pushing back 1/10 Choosing Independent Consulting over Venture Capital Backing Swix asks why Jason declined venture backing for Instructor. Jason explains his philosophy of building a standard utility like Requests while running a high-margin, flexible solo consulting practice.37:10–41:42 · The hosts pushing back 1/10 Startup Realities, Solo Founding, and Entrepreneurial Courage Alessio and Swix discuss the realities of seed funding versus high-paying corporate employment. Jason highlights founder breakup rates and emphasizes having the courage to build with minimal overhead.41:42–43:59 · The hosts pushing back 1/10 Overcoming Career Setbacks and Starting from Zero Alessio asks about overcoming early rejection after Jason was turned down by top AI labs in early 2023. Jason uses a Jiu-Jitsu analogy to urge engineers to start immediately rather than waiting for ideal conditions.43:59–49:10 · The hosts pushing back 6/10 High Agency, Process Metrics, and the Clay Metaphor Jason defines agency as focusing on process volume rather than final outcome metrics using a pottery clay metaphor. Swix directly pushes back, asking whether this equates to flawed developer metrics like counting lines of code.49:10–54:40 · The hosts pushing back 1/10 Automating Personal Workflows and Future AI Capabilities Alessio and Jason discuss personal workflow automation. Jason explains how hiring human assistants helped him uncover precise, modular heuristic prompts for calendar management and client summaries.54:40–57:22 · The hosts pushing back 2/10 DSPy, Prompt Optimization, and Business-Level Metrics Swix raises DSPy and the thesis that AI is a better prompt optimizer than humans. Jason argues that DSPy works well on narrow programmatic heuristics but breaks down when evaluating long-term business outcomes like user churn.57:22–1:01:36 · The hosts pushing back 1/10 AI Engineers vs ML Engineers and Organizational Talent Swix asks about the distinction between MLEs and AI engineers. Jason delivers a strong take advising application startups against hiring traditional MLEs, who quickly churn when faced with TypeScript and missing training pipelines.1:01:36–1:02:33 · The hosts pushing back 1/10 AI Engineer World's Fair Preview and Episode Conclusion Swix previews the AI Engineer World's Fair. Jason jokes about his talk title while reinforcing the core philosophy behind Pydantic-based structured outputs.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 58:41 Startups should stop hiring PyTorch ML engineers

Jason forcefully dismisses the common startup hiring practice of recruiting PyTorch MLEs for application-layer products, pointing out that they inevitably churn when forced to fix TypeScript errors.

Hardest push from the hosts ▶ 45:34 Pushback on measuring process volume over outcomes

Swix directly confronts Jason's pottery clay metaphor by arguing that measuring process volume in software engineering is the equivalent of counting lines of code.

Biggest teaching moment ▶ 25:37 Embeddings fail on temporal and aggregation queries

Jason walks through explicit counterexamples showing why semantic vector search fails on date-relative or group-by queries, demonstrating the necessity of structured data extraction.

The host holds their own ▶ 19:08 Real-world incubation data on 780 API endpoints

Alessio counters theoretical tool retrieval suggestions by sharing concrete empirical data from a 780-endpoint integration startup where full-context LLM selection outperformed vector ranking.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Stitch Fix Experience and Early LLM Skepticism 6522 Swix connects RAG to classical recommendation systems (RecSys) and probes why Jason built the Flight recommendation system from scratch at Stitch Fix. Jason explains the operational necessity of standardizing bespoke data science codebases to enable observability and performance tuning.
Similarity Search, Multimodal Embeddings, and Real-World Impact 6512 Swix asks technical clarifying questions about whether off-the-shelf GPT-3 embeddings worked for multi-modal clothing retrieval. Jason clarifies that fine-tuning ResNets and joint text-image representations was necessary to beat generic models at scale.
Fashion, Human Narrative, and Recommender Limitations 3631 Swix jokes about tech fashion and asks whether tech workers should use Stitch Fix. Jason reframes the premise by contrasting utility-driven clothing delivery with human narrative and experiential shopping.
The Inception and Architectural Philosophy of Instructor 6631 Swix and Jason discuss library architecture, contrasting broad frameworks like Django with lightweight standard utilities like Requests. Jason details how function calling separated schemas from prompts, driving the inception of Instructor.
Typed Outputs, JSON Mode, and Function Calling Mechanics 7532 Alessio asks technical questions comparing JSON mode to function calling and inquires about SDK-level handling of parallel tool invocation. Jason explains why typed schemas with validation outperform raw JSON output modes.
Tool Selection, Context Scaling, and Ranking Architectures 8435 Alessio pushes back on relying solely on vector similarity for tool selection, sharing firsthand incubation data on 780 API endpoints where LLMs select tools better than semantic search. Jason defends ranking via XGBoost over stuffing prompts with tool definitions.
Production Benchmarks and Anthropic vs OpenAI Trade-offs 6541 Alessio notes Jason's contrarian evaluation of Anthropic function calling benchmarks. Jason shares real production telemetry, revealing edge-case parsing bugs despite compelling pricing advantages.
Mapping the Use Cases and Surface Area of Instructor 5841 Swix invites Jason to delineate the capabilities of Instructor across graphs, sanitation, and query understanding. Jason schools the audience on why vector embeddings fail on temporal aggregations and relational queries.
Workflows and Deterministic DAGs versus ReAct Agent Loops 6542 Swix asks how Instructor compares with LangChain's self-querying and prompts Jason on workflow design. Jason advocates for deterministic DAGs and fine-tuning compilers over non-deterministic ReAct prompting loops.
Autonomous Agents, Execution Horizons, and Plan Churn 6443 Swix discusses evaluating long-horizon coding agents like Devin and cautions against rewarding runtime duration. Jason mocks early agent failure modes and argues for monotonic progress metrics over loop thrashing.
AI Engineering Stack and Pragmatic Infrastructure 6452 Swix inquires about Jason's preferred AI engineering stack. Jason dismisses the dozens of VC-backed LLM observability startups, explaining that a single Postgres table and standard APM tools like Datadog and Sentry are more than sufficient.
Choosing Independent Consulting over Venture Capital Backing 5441 Swix asks why Jason declined venture backing for Instructor. Jason explains his philosophy of building a standard utility like Requests while running a high-margin, flexible solo consulting practice.
Startup Realities, Solo Founding, and Entrepreneurial Courage 6431 Alessio and Swix discuss the realities of seed funding versus high-paying corporate employment. Jason highlights founder breakup rates and emphasizes having the courage to build with minimal overhead.
Overcoming Career Setbacks and Starting from Zero 4421 Alessio asks about overcoming early rejection after Jason was turned down by top AI labs in early 2023. Jason uses a Jiu-Jitsu analogy to urge engineers to start immediately rather than waiting for ideal conditions.
High Agency, Process Metrics, and the Clay Metaphor 6646 Jason defines agency as focusing on process volume rather than final outcome metrics using a pottery clay metaphor. Swix directly pushes back, asking whether this equates to flawed developer metrics like counting lines of code.
Automating Personal Workflows and Future AI Capabilities 5421 Alessio and Jason discuss personal workflow automation. Jason explains how hiring human assistants helped him uncover precise, modular heuristic prompts for calendar management and client summaries.
DSPy, Prompt Optimization, and Business-Level Metrics 6642 Swix raises DSPy and the thesis that AI is a better prompt optimizer than humans. Jason argues that DSPy works well on narrow programmatic heuristics but breaks down when evaluating long-term business outcomes like user churn.
AI Engineers vs ML Engineers and Organizational Talent 6551 Swix asks about the distinction between MLEs and AI engineers. Jason delivers a strong take advising application startups against hiring traditional MLEs, who quickly churn when faced with TypeScript and missing training pipelines.
AI Engineer World's Fair Preview and Episode Conclusion 4211 Swix previews the AI Engineer World's Fair. Jason jokes about his talk title while reinforcing the core philosophy behind Pydantic-based structured outputs.

Statements from this episode (27)

Assertion Supported
Liu: Stitch Fix used Transformer models for recommendations before GPT-3
“We actually were using Transformers at Stitch Fix, like, before the GPT-III model, so we were just using Transformers for recommendation systems.”
Jason Liu Apr 24, 2024 ▶ 1:42
Insight
Liu: Fine-tuning on proprietary data at scale always beats off-the-shelf models
“We, because, I mean, at this point, we would have, like, you know, three million pieces of inventory, over, like, a billion interactions between users and clothes. Any kind of fine-taining would definitely outperform like, some off-the-shove model.”
Jason Liu Apr 24, 2024 ▶ 6:23
Disclosure
Liu: Admitted being bearish on LLMs for four years before recent breakthroughs
“I mean, the biggest one really was the fact that, like, I think for just four years I was so bearish on language models. And just NLP in general, I was just like, ah, like, none of this really works. Like, why would I spend time focusing on this? I gotta go do…”
Jason Liu Apr 24, 2024 ▶ 6:47
Insight
Liu: AI recommenders can sell $20 shirts but struggle with luxury narratives
“Narrative matters a lot to human beings. And I think the recommendation system, that's really hard to capture. Like, it's easy to sell, it's easy to use AI to sell, like, a 20 dollar shirt, but it's really hard for AI to sell, like, a 500 dollar shirt.”
Jason Liu Apr 24, 2024 ▶ 8:00
Insight
Liu: Function calling decouples schemas from prompt instructions
“Function calling lets you define the schema separate from the data and the instructions. And what this meant was you can kind of have a lot more complex schemas and just map them in Pydantic, and then you can just keep those very separate.”
Jason Liu Apr 24, 2024 ▶ 10:49
Disclosure
Liu: Instructor was designed as a minimal wrapper akin to Requests
“And so I just said, let me write, like, the most simple SDK around the OpenAI SDK, sorry, simple wrapper on the SDK, just handle the response model a bit, and kind of think of myself more like requests than an actual framework that people can use.”
Jason Liu Apr 24, 2024 ▶ 11:30
Insight
Liu: Developers love building custom frameworks but hate writing JSON parsers
“People want to build their own frameworks, but people don't want to build, like, JSON parsing.”
Jason Liu Apr 24, 2024 ▶ 11:54
Insight
Liu: LLM JSON mode is worse than function calling for structured outputs
“In terms of whether or not, like, JSON mode is better, I usually think it's almost worse unless you want to spend less money on, like, the prompt tokens that the function call represents. Primarily because with JSON mode you don't actually specify the schema.”
Jason Liu Apr 24, 2024 ▶ 14:08
Insight
Liu: Single schema extraction beats parallel function calling for relationship modeling
“In terms of an extraction workflow, I definitely think it's probably more helpful to have everything be a single schema. Just because you can sort of specify relationships between these entities, right, that you can't do in parallel function calling, you can h…”
Jason Liu Apr 24, 2024 ▶ 15:36
Insight
Liu: Rank and retrieve tool definitions instead of passing dozens to LLMs
“If you're running into issues where you have, like, 20 or 50 or 60 function calls, I think you're much better having those specifications saved in a vector database, and then have them be retrieved. So if there are 30 tools, like, you should basically be, like…”
Jason Liu Apr 24, 2024 ▶ 16:51
Prediction Not checkable as stated
Liu: XGBoost rankers will beat LLMs for large-scale tool selection
“Yeah, my money is on the rankers because you can do those so easily, right? You could just say, well, given the embeddings of my search query and the embeddings of the description, I can just train XGBoost and just make sure that I have very high, like, MRR, w…”
Jason Liu Apr 24, 2024 ▶ 19:54
Opinion
Liu: Claude 3 Haiku outperforms OpenAI models at function calling
“Overall, I'm like super happy with the anthropic models compared to the OpenAM models. Like, Sonnet is very cost effective. Haiku is, in function calling, it's actually better.”
Jason Liu Apr 24, 2024 ▶ 21:57
Insight
Liu: LLMs excel at identifying nodes and edges for knowledge graphs
“One of the things we found out about these language models is that not only can you define nodes, it's really good at figuring out what are nodes and what are edges.”
Jason Liu Apr 24, 2024 ▶ 24:32
Insight
Liu: Structured LLM outputs unlock traditional computer science reasoning algorithms
“Embeddings really is kind of like the lowest hanging fruit, and using something like Instructor can really help produce a data structure, and then you can just use your computer science to reason about this data structure.”
Jason Liu Apr 24, 2024 ▶ 26:02
Insight
Liu: AI agents should use explicit DAG workflows over ReAct loops
“Instead of doing like a react type reasoning loop, I think my belief is that we should be using like workflows, right? If we do this, then we always have a request and a complete workflow. We can fine tune a model that has a better workflow. Whereas it's hard …”
Jason Liu Apr 24, 2024 ▶ 27:10
Prediction Not checkable as stated
Liu: Prefect and Zapier are well-positioned to build AI workflow UIs
“I think, you know, people like Prefect and Zapier have a pretty good shot at doing a good job.”
Jason Liu Apr 24, 2024 ▶ 31:50
Insight
Liu: LLM observability startups ignore full systems; just use Postgres
“The issue really is the fact that these observability companies isn't actually doing observability for the system, it's just doing the LLM thing. Like I still end up using like Datadog, right? Or like, you know, Sentry to do, like, latency. And so I just have …”
Jason Liu Apr 24, 2024 ▶ 32:35
Opinion
Liu: Instructor is not a billion-dollar startup opportunity
“But to back to the instructor thing, I just don't think it's a billion dollar company.”
Jason Liu Apr 24, 2024 ▶ 36:29
Insight
Liu: High-agency individuals focus on process metrics over outcome metrics
“I think the higher agency person is more focused on like process metrics versus outcome metrics. Right? Like, from pottery, like, one thing I learned was, if you want to be good at pottery, you shouldn't count, like, the number of cups or bowls you make. You s…”
Jason Liu Apr 24, 2024 ▶ 44:49
Insight
Liu: Agency drives machine learning experiment volume; experience filters wasteful trials
“So, agency lets you sort of capture the volume of experiments, and, like, experience lets you figure out, like, oh, that other half, it's not worth doing.”
Jason Liu Apr 24, 2024 ▶ 46:53
Insight
Liu: Define explicit conditions for revisiting negative machine learning experiment results
“Like what you should write down is like, here are the conditions. This is the inputs and the outputs we tried the experiment on. And then one thing that's really valuable is basically writing down under what conditions would I revisit these experiments?”
Jason Liu Apr 24, 2024 ▶ 48:04
Opinion
Liu: GPT-4 and Claude 3 Opus still write poor quality essays
“Or those are two sort of systems that I wish you before or Opus was actually good enough to just write me an essay, but most of the essays are still pretty bad.”
Jason Liu Apr 24, 2024 ▶ 49:50
Insight
Liu: Companies abandon LLM frameworks to regain control over prompts
“So much of it is changing that if you give control of these systems away too early, you end up ultimately wanting them back. Like many companies I know that I reach out or ones were like, oh, we're going off of the frameworks because now that we know what the …”
Jason Liu Apr 24, 2024 ▶ 52:08
Opinion
Liu: LangChain and LlamaIndex can hit $100M revenue, but billions uncertain
“I think the bigger challenge is like, okay, a hundred million dollars, probably pretty, pretty easy. It's just time and effort. And they have both like the manpower and the money to sort of solve those problems. I think it's just like, again, if you go the VC …”
Jason Liu Apr 24, 2024 ▶ 53:43
Insight
Liu: DSPy excels at micro-tasks but fails at end-to-end business workflows
“I think something like DSPy can work because there are like very short term metrics to measure success, right? It is like, did you find the PII or like, did you write the multi-hop question the correct way? But in these like workflows that I've been managing, …”
Jason Liu Apr 24, 2024 ▶ 55:15
Opinion
Liu: App-layer AI startups should avoid hiring traditional machine learning engineers
“I think a lot of these app layer startups should not be hiring MLEs because they end up churning.”
Jason Liu Apr 24, 2024 ▶ 58:42
Prediction Not checkable as stated
Liu: Data science skill sets will outvalue traditional machine learning engineering
“I think a lot more data science is going to come in versus machine learning engineering, because a lot of it now is just quantifying, like, what does the business actually want as an outcome, right?”
Jason Liu Apr 24, 2024 ▶ 1:01:10
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.