Apr 24, 2024 · 1h 2m · latent-space
High Agency Pydantic over VC Backed Frameworks — with Jason Liu of Instructor
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this episode of Latent Space, Jason Liu, creator of Instructor, joins Alessio and Swix to discuss the mechanics of structured LLM outputs, the architectural advantages of pragmatic engineering over VC-backed frameworks, and the emerging discipline of AI engineering.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Jason forcefully dismisses the common startup hiring practice of recruiting PyTorch MLEs for application-layer products, pointing out that they inevitably churn when forced to fix TypeScript errors.
Hardest push from the hosts ▶ 45:34 Pushback on measuring process volume over outcomesSwix directly confronts Jason's pottery clay metaphor by arguing that measuring process volume in software engineering is the equivalent of counting lines of code.
Biggest teaching moment ▶ 25:37 Embeddings fail on temporal and aggregation queriesJason walks through explicit counterexamples showing why semantic vector search fails on date-relative or group-by queries, demonstrating the necessity of structured data extraction.
The host holds their own ▶ 19:08 Real-world incubation data on 780 API endpointsAlessio counters theoretical tool retrieval suggestions by sharing concrete empirical data from a 780-endpoint integration startup where full-context LLM selection outperformed vector ranking.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Stitch Fix Experience and Early LLM Skepticism | 6 | 5 | 2 | 2 | Swix connects RAG to classical recommendation systems (RecSys) and probes why Jason built the Flight recommendation system from scratch at Stitch Fix. Jason explains the operational necessity of standardizing bespoke data science codebases to enable observability and performance tuning. | |
| Similarity Search, Multimodal Embeddings, and Real-World Impact | 6 | 5 | 1 | 2 | Swix asks technical clarifying questions about whether off-the-shelf GPT-3 embeddings worked for multi-modal clothing retrieval. Jason clarifies that fine-tuning ResNets and joint text-image representations was necessary to beat generic models at scale. | |
| Fashion, Human Narrative, and Recommender Limitations | 3 | 6 | 3 | 1 | Swix jokes about tech fashion and asks whether tech workers should use Stitch Fix. Jason reframes the premise by contrasting utility-driven clothing delivery with human narrative and experiential shopping. | |
| The Inception and Architectural Philosophy of Instructor | 6 | 6 | 3 | 1 | Swix and Jason discuss library architecture, contrasting broad frameworks like Django with lightweight standard utilities like Requests. Jason details how function calling separated schemas from prompts, driving the inception of Instructor. | |
| Typed Outputs, JSON Mode, and Function Calling Mechanics | 7 | 5 | 3 | 2 | Alessio asks technical questions comparing JSON mode to function calling and inquires about SDK-level handling of parallel tool invocation. Jason explains why typed schemas with validation outperform raw JSON output modes. | |
| Tool Selection, Context Scaling, and Ranking Architectures | 8 | 4 | 3 | 5 | Alessio pushes back on relying solely on vector similarity for tool selection, sharing firsthand incubation data on 780 API endpoints where LLMs select tools better than semantic search. Jason defends ranking via XGBoost over stuffing prompts with tool definitions. | |
| Production Benchmarks and Anthropic vs OpenAI Trade-offs | 6 | 5 | 4 | 1 | Alessio notes Jason's contrarian evaluation of Anthropic function calling benchmarks. Jason shares real production telemetry, revealing edge-case parsing bugs despite compelling pricing advantages. | |
| Mapping the Use Cases and Surface Area of Instructor | 5 | 8 | 4 | 1 | Swix invites Jason to delineate the capabilities of Instructor across graphs, sanitation, and query understanding. Jason schools the audience on why vector embeddings fail on temporal aggregations and relational queries. | |
| Workflows and Deterministic DAGs versus ReAct Agent Loops | 6 | 5 | 4 | 2 | Swix asks how Instructor compares with LangChain's self-querying and prompts Jason on workflow design. Jason advocates for deterministic DAGs and fine-tuning compilers over non-deterministic ReAct prompting loops. | |
| Autonomous Agents, Execution Horizons, and Plan Churn | 6 | 4 | 4 | 3 | Swix discusses evaluating long-horizon coding agents like Devin and cautions against rewarding runtime duration. Jason mocks early agent failure modes and argues for monotonic progress metrics over loop thrashing. | |
| AI Engineering Stack and Pragmatic Infrastructure | 6 | 4 | 5 | 2 | Swix inquires about Jason's preferred AI engineering stack. Jason dismisses the dozens of VC-backed LLM observability startups, explaining that a single Postgres table and standard APM tools like Datadog and Sentry are more than sufficient. | |
| Choosing Independent Consulting over Venture Capital Backing | 5 | 4 | 4 | 1 | Swix asks why Jason declined venture backing for Instructor. Jason explains his philosophy of building a standard utility like Requests while running a high-margin, flexible solo consulting practice. | |
| Startup Realities, Solo Founding, and Entrepreneurial Courage | 6 | 4 | 3 | 1 | Alessio and Swix discuss the realities of seed funding versus high-paying corporate employment. Jason highlights founder breakup rates and emphasizes having the courage to build with minimal overhead. | |
| Overcoming Career Setbacks and Starting from Zero | 4 | 4 | 2 | 1 | Alessio asks about overcoming early rejection after Jason was turned down by top AI labs in early 2023. Jason uses a Jiu-Jitsu analogy to urge engineers to start immediately rather than waiting for ideal conditions. | |
| High Agency, Process Metrics, and the Clay Metaphor | 6 | 6 | 4 | 6 | Jason defines agency as focusing on process volume rather than final outcome metrics using a pottery clay metaphor. Swix directly pushes back, asking whether this equates to flawed developer metrics like counting lines of code. | |
| Automating Personal Workflows and Future AI Capabilities | 5 | 4 | 2 | 1 | Alessio and Jason discuss personal workflow automation. Jason explains how hiring human assistants helped him uncover precise, modular heuristic prompts for calendar management and client summaries. | |
| DSPy, Prompt Optimization, and Business-Level Metrics | 6 | 6 | 4 | 2 | Swix raises DSPy and the thesis that AI is a better prompt optimizer than humans. Jason argues that DSPy works well on narrow programmatic heuristics but breaks down when evaluating long-term business outcomes like user churn. | |
| AI Engineers vs ML Engineers and Organizational Talent | 6 | 5 | 5 | 1 | Swix asks about the distinction between MLEs and AI engineers. Jason delivers a strong take advising application startups against hiring traditional MLEs, who quickly churn when faced with TypeScript and missing training pipelines. | |
| AI Engineer World's Fair Preview and Episode Conclusion | 4 | 2 | 1 | 1 | Swix previews the AI Engineer World's Fair. Jason jokes about his talk title while reinforcing the core philosophy behind Pydantic-based structured outputs. |