Mar 19, 2025 · 31m · latent-space

Fullstack-Bench: The Eval for Coding Agents — with Sujay Jayakar, Chief Scientist, Convex

Sujay Jayakar · 21m spoken Shawn Wang · 5m spoken Alessio Fanelli · 1m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of Latent Space, Convex Co-founder and Chief Scientist Sujay Jayakar introduces Fullstack-Bench, a comprehensive benchmark evaluating AI coding agents across different backend architectures. He discusses model failure modes, the evolution from Developer Experience to Agent Experience, and the infrastructure required to support autonomous software development.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 4.2% of the talking time here. How this is scored →

The hosts as informed peer 5.0 Guest teaching 3.9 Guest disagreement 1.1 The hosts pushing back 1.7
05100:0010:0020:0030:001:52–8:00 · The hosts as informed peer 6/10 Fullstack-Bench Experimental Setup and Benchmark Design Shawn Wang brings up existing benchmark comparisons like 7GUIs and Thinkster RealWorld app suites to contextualize Fullstack-Bench. Sujay readily agrees with the limitations regarding templates and scaffolds.8:00–12:36 · The hosts as informed peer 5/10 Benchmark Results and Human-in-the-Loop Interventions Alessio queries how human intervention and hint-giving works when models run off-track for long periods. Sujay explains the practical difficulties of manual eval administration and debugging loops.12:36–18:09 · The hosts as informed peer 6/10 Model Discrepancies, Cursor Rules, and SWE-bench Comparison Alessio shares practical experience using Cursor rules to stop models defaulting to deprecated RSpec matchers, while Sujay discusses Claude 3.7 regressing compared to 3.5 on internal Convex evals.18:09–20:43 · The hosts as informed peer 3/10 Future Benchmarking: Concurrency, Verification, and Throughput Sujay outlines advanced verification plans using Jepsen model checking for serializability and TPC-style throughput benchmarks, with the hosts purely listening.20:43–25:55 · The hosts as informed peer 6/10 Developer Experience vs. Agent Experience and API Calcification Shawn challenges the assumption that good DX maps directly to agent experience (AX) and posits that pre-training cutoffs calcify API design. Sujay agrees that models struggle with subtle differences from Firebase.25:55–28:57 · The hosts as informed peer 5/10 Agent Architectures, Durable Execution, and AI Town Shawn asks why the industry saw a lack of practical follow-through after the viral AI Town demo. Sujay explains the distinction between traditional apps built via AI and fundamentally new agent backend architectures.28:57–30:58 · The hosts as informed peer 4/10 Key Takeaways for AI Tool Builders and Podcast Conclusion Sujay wraps up with key lessons from the benchmark, contrasting procedural debugging with declarative pitfalls like Supabase RLS infinite recursion before concluding the episode.1:52–8:00 · Guest teaching 3/10 Fullstack-Bench Experimental Setup and Benchmark Design Shawn Wang brings up existing benchmark comparisons like 7GUIs and Thinkster RealWorld app suites to contextualize Fullstack-Bench. Sujay readily agrees with the limitations regarding templates and scaffolds.8:00–12:36 · Guest teaching 4/10 Benchmark Results and Human-in-the-Loop Interventions Alessio queries how human intervention and hint-giving works when models run off-track for long periods. Sujay explains the practical difficulties of manual eval administration and debugging loops.12:36–18:09 · Guest teaching 4/10 Model Discrepancies, Cursor Rules, and SWE-bench Comparison Alessio shares practical experience using Cursor rules to stop models defaulting to deprecated RSpec matchers, while Sujay discusses Claude 3.7 regressing compared to 3.5 on internal Convex evals.18:09–20:43 · Guest teaching 6/10 Future Benchmarking: Concurrency, Verification, and Throughput Sujay outlines advanced verification plans using Jepsen model checking for serializability and TPC-style throughput benchmarks, with the hosts purely listening.20:43–25:55 · Guest teaching 3/10 Developer Experience vs. Agent Experience and API Calcification Shawn challenges the assumption that good DX maps directly to agent experience (AX) and posits that pre-training cutoffs calcify API design. Sujay agrees that models struggle with subtle differences from Firebase.25:55–28:57 · Guest teaching 3/10 Agent Architectures, Durable Execution, and AI Town Shawn asks why the industry saw a lack of practical follow-through after the viral AI Town demo. Sujay explains the distinction between traditional apps built via AI and fundamentally new agent backend architectures.28:57–30:58 · Guest teaching 4/10 Key Takeaways for AI Tool Builders and Podcast Conclusion Sujay wraps up with key lessons from the benchmark, contrasting procedural debugging with declarative pitfalls like Supabase RLS infinite recursion before concluding the episode.1:52–8:00 · Guest disagreement 1/10 Fullstack-Bench Experimental Setup and Benchmark Design Shawn Wang brings up existing benchmark comparisons like 7GUIs and Thinkster RealWorld app suites to contextualize Fullstack-Bench. Sujay readily agrees with the limitations regarding templates and scaffolds.8:00–12:36 · Guest disagreement 1/10 Benchmark Results and Human-in-the-Loop Interventions Alessio queries how human intervention and hint-giving works when models run off-track for long periods. Sujay explains the practical difficulties of manual eval administration and debugging loops.12:36–18:09 · Guest disagreement 2/10 Model Discrepancies, Cursor Rules, and SWE-bench Comparison Alessio shares practical experience using Cursor rules to stop models defaulting to deprecated RSpec matchers, while Sujay discusses Claude 3.7 regressing compared to 3.5 on internal Convex evals.18:09–20:43 · Guest disagreement 0/10 Future Benchmarking: Concurrency, Verification, and Throughput Sujay outlines advanced verification plans using Jepsen model checking for serializability and TPC-style throughput benchmarks, with the hosts purely listening.20:43–25:55 · Guest disagreement 2/10 Developer Experience vs. Agent Experience and API Calcification Shawn challenges the assumption that good DX maps directly to agent experience (AX) and posits that pre-training cutoffs calcify API design. Sujay agrees that models struggle with subtle differences from Firebase.25:55–28:57 · Guest disagreement 1/10 Agent Architectures, Durable Execution, and AI Town Shawn asks why the industry saw a lack of practical follow-through after the viral AI Town demo. Sujay explains the distinction between traditional apps built via AI and fundamentally new agent backend architectures.28:57–30:58 · Guest disagreement 1/10 Key Takeaways for AI Tool Builders and Podcast Conclusion Sujay wraps up with key lessons from the benchmark, contrasting procedural debugging with declarative pitfalls like Supabase RLS infinite recursion before concluding the episode.1:52–8:00 · The hosts pushing back 2/10 Fullstack-Bench Experimental Setup and Benchmark Design Shawn Wang brings up existing benchmark comparisons like 7GUIs and Thinkster RealWorld app suites to contextualize Fullstack-Bench. Sujay readily agrees with the limitations regarding templates and scaffolds.8:00–12:36 · The hosts pushing back 2/10 Benchmark Results and Human-in-the-Loop Interventions Alessio queries how human intervention and hint-giving works when models run off-track for long periods. Sujay explains the practical difficulties of manual eval administration and debugging loops.12:36–18:09 · The hosts pushing back 2/10 Model Discrepancies, Cursor Rules, and SWE-bench Comparison Alessio shares practical experience using Cursor rules to stop models defaulting to deprecated RSpec matchers, while Sujay discusses Claude 3.7 regressing compared to 3.5 on internal Convex evals.18:09–20:43 · The hosts pushing back 0/10 Future Benchmarking: Concurrency, Verification, and Throughput Sujay outlines advanced verification plans using Jepsen model checking for serializability and TPC-style throughput benchmarks, with the hosts purely listening.20:43–25:55 · The hosts pushing back 4/10 Developer Experience vs. Agent Experience and API Calcification Shawn challenges the assumption that good DX maps directly to agent experience (AX) and posits that pre-training cutoffs calcify API design. Sujay agrees that models struggle with subtle differences from Firebase.25:55–28:57 · The hosts pushing back 2/10 Agent Architectures, Durable Execution, and AI Town Shawn asks why the industry saw a lack of practical follow-through after the viral AI Town demo. Sujay explains the distinction between traditional apps built via AI and fundamentally new agent backend architectures.28:57–30:58 · The hosts pushing back 0/10 Key Takeaways for AI Tool Builders and Podcast Conclusion Sujay wraps up with key lessons from the benchmark, contrasting procedural debugging with declarative pitfalls like Supabase RLS infinite recursion before concluding the episode.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 5% · guest 95%0:00 · the hosts 5% · guest 95%3:00 · the hosts 11.4% · guest 88.6%3:00 · the hosts 11.4% · guest 88.6%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 9.4% · guest 90.6%9:00 · the hosts 9.4% · guest 90.6%12:00 · the hosts 14.3% · guest 85.7%12:00 · the hosts 14.3% · guest 85.7%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 3.3% · guest 96.7%18:00 · the hosts 3.3% · guest 96.7%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 23:20 Pushing past the assumption that DX equals AX

Sujay rejects the conventional wisdom that developer ergonomics naturally translate to LLM usability, noting models get tripped up by subtle deviations from legacy frameworks like Firebase.

Hardest push from the hosts ▶ 24:06 swyx argues AI locks in legacy APIs

Shawn pushes back on API flexibility, arguing that LLM training data effectively calcifies API design and makes switching syntax nearly impossible.

Biggest teaching moment ▶ 18:30 Sujay details formal concurrency verification for agents

Sujay educates the audience and hosts on applying distributed systems tooling like Jepsen and Elle to automatically grade agent-generated database code under high concurrency.

The host holds their own ▶ 5:11 swyx contextualizes benchmark with 7GUIs and RealWorld

Shawn demonstrates deep familiarity with UI/backend evaluation literature by connecting Fullstack-Bench to historical projects like 7GUIs and Thinkster's RealWorld spec.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Fullstack-Bench Experimental Setup and Benchmark Design 6312 Shawn Wang brings up existing benchmark comparisons like 7GUIs and Thinkster RealWorld app suites to contextualize Fullstack-Bench. Sujay readily agrees with the limitations regarding templates and scaffolds.
Benchmark Results and Human-in-the-Loop Interventions 5412 Alessio queries how human intervention and hint-giving works when models run off-track for long periods. Sujay explains the practical difficulties of manual eval administration and debugging loops.
Model Discrepancies, Cursor Rules, and SWE-bench Comparison 6422 Alessio shares practical experience using Cursor rules to stop models defaulting to deprecated RSpec matchers, while Sujay discusses Claude 3.7 regressing compared to 3.5 on internal Convex evals.
Future Benchmarking: Concurrency, Verification, and Throughput 3600 Sujay outlines advanced verification plans using Jepsen model checking for serializability and TPC-style throughput benchmarks, with the hosts purely listening.
Developer Experience vs. Agent Experience and API Calcification 6324 Shawn challenges the assumption that good DX maps directly to agent experience (AX) and posits that pre-training cutoffs calcify API design. Sujay agrees that models struggle with subtle differences from Firebase.
Agent Architectures, Durable Execution, and AI Town 5312 Shawn asks why the industry saw a lack of practical follow-through after the viral AI Town demo. Sujay explains the distinction between traditional apps built via AI and fundamentally new agent backend architectures.
Key Takeaways for AI Tool Builders and Podcast Conclusion 4410 Sujay wraps up with key lessons from the benchmark, contrasting procedural debugging with declarative pitfalls like Supabase RLS infinite recursion before concluding the episode.

Statements from this episode (11)

Assertion Supported
Claude 3.5 in Cursor autonomously codes for 10 plus minutes
“When it has the right feedback in cursor composer, and this was even on cloud three, five, it can just autonomously code for a 10 plus minutes and it can fix its own bugs. It can get to the point where it's like pretty much a fully working app with just an ini…”
Sujay Jayakar Mar 19, 2025 ▶ 3:21
Assertion Supported
Cursor Composer solves Convex benchmarks but fails on alternative backends
“We did notice that I mean, with convex, it pretty much autonomously just solves the first two tasks. It has a few round trips on like some errors that are only show up and playing with the front end. And then it's able to complete this files task and kind of g…”
Sujay Jayakar Mar 19, 2025 ▶ 9:36
Assertion Not checkable as stated
Jayakar: AI Models Stalled on Convex's Subtle Distinction Between Null and Undefined
“Convex has like a pretty subtle distinction between null and undefined, like just similar to JavaScript. And we noticed that like, because this is a subtle thing that's unexpected, it was The model got stuck on it and couldn't even figure it out. And so this i…”
Sujay Jayakar Mar 19, 2025 ▶ 13:01
Assertion Not checkable as stated
Claude 3.7 performed worse than Claude 3.5 on Convex evals
“For example, we just tried clod three seven and it performs worse than clod three five on convex evals with the same prompting.”
Sujay Jayakar Mar 19, 2025 ▶ 17:23
Assertion Not checkable as stated
OpenAI's o3 outperforms GPT-4o on Convex evals by a small margin
“You know, oh, three does do better than four. Oh, I mean, we use brain trust for tracking all this quantitatively, but I can't remember off the top of my head, but it's not like a slam dunk.”
Sujay Jayakar Mar 19, 2025 ▶ 17:56
Disclosure
Jayakar prototypes AI coding benchmark using Jepsen's Elle model checker
“I think like, you know, I've wired up a version of this, you know, using L, which is the like model checker from Jepson and, you know, can a model write highly correct, highly like code that executes under very high concurrency, even when there's”
Sujay Jayakar Mar 19, 2025 ▶ 18:57
Assertion Not checkable as stated
AI hallucinates Convex code due to API similarities with Firebase
“I mean, I think one thing that's kind of interesting given these, you know, models inherent internal structure is that we noticed that a lot of hallucinations come from parts of our API that are like very close to Firebase, but not exactly Firebase.”
Sujay Jayakar Mar 19, 2025 ▶ 22:56
Opinion
Swyx: AI pre-training datasets calcify software APIs against future changes
“Like one tricky thing, I feel like it's almost like AI calcifies your API. The cost of switching APIs is so high because it's, you don't know what the hell is in those, the data set. You're not going to reach out to every model lab and like ask them to update …”
Shawn Wang Mar 19, 2025 ▶ 24:08
Insight
End-to-end type safety and tight feedback loops improve AI agent performance
“It's like, you know, having just really tight feedback loops and having strong guardrails, like in convex, that's like end to end type safety, but that's like, you know, could take it in many, many different forms, right? Like making it so code is very cheap a…”
Sujay Jayakar Mar 19, 2025 ▶ 29:12
Assertion Supported
AI models struggle debugging Supabase RLS recursion compared to procedural code
“The particular example was like RLS rules and Supabase where debugging like an infinite loop for infinite recursion for the RLS rules was something that the models just really struggled with in a way that we didn't see for procedural code.”
Sujay Jayakar Mar 19, 2025 ▶ 29:57
Insight
Strong library abstractions prevent AI models from breaking backend wiring
“Models, you know, when given the full flexibility of fast API and doing SSE and wiring everything from scratch. It just couldn't help, but messed it up. So having like kind of strong life, like choosing good libraries that have strong abstractions is like, and…”
Sujay Jayakar Mar 19, 2025 ▶ 30:14
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.