Dec 17, 2023 · 1h 33m · latent-space
The "Normsky" architecture for AI coding agents — with Beyang Liu + Steve Yegge of SourceGraph
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
Sourcegraph's Beyang Liu and Steve Yegge discuss the General Availability of Cody and explain the 'Normsky' architecture, which synthesizes deterministic code graph indexing with statistical LLMs to master enterprise-scale code intelligence.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 20.3% of the talking time here. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Beyang forcefully reframes Swyx's scaling argument by demanding proof that pure transformer scaling can beat tree search algorithms in chess or coding.
Hardest push from the hosts ▶ 59:36 Swyx pushing back on agent skepticism with compute scaling lawsSwyx refuses the guests' dismissive stance on autonomous agents, detailing compute scaling calculations up to 2030 to challenge their near-term limits.
Biggest teaching moment ▶ 49:55 Beyang Liu detailing LSP versus SCIP and Kythe trade-offsBeyang delivers a comprehensive breakdown of why LSP's range-based abstraction fails to capture symbolic code relationships and why custom knowledge graphs are essential.
The host holds their own ▶ 59:36 Swyx calculating human-year compute scaling projectionsSwyx demonstrates deep domain expertise by citing George Hotz's petaflop calculations and Sam Altman's exponential model generation projections.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Introductions and Origins of Sourcegraph and Grok | 4 | 2 | 1 | 1 | The conversation starts with warm, collegial rapport. Swyx demonstrates strong familiarity with Steve's famous engineering essays, Google/Amazon tenure, and Grab's engineering culture. | |
| Introducing Cody and the Power of RAG Context | 6 | 3 | 1 | 1 | Alessio demonstrates sharp domain knowledge by quoting specific benchmark discrepancies between Copilot and Cody regarding package.json start commands. Beyang and Steve elaborate on how RAG serves as an expert consultant compared to fine-tuning. | |
| Code Search as a Recommendation Engine and Context Ranking | 6 | 2 | 1 | 1 | Swyx connects RAG ranking to classical recommendation systems and highlights context attention curves like lost-in-the-middle phenomena. Steve and Beyang validate this architectural view. | |
| Inline UX and Skepticism Toward Autonomous LLM Agents | 4 | 4 | 3 | 2 | Beyang and Steve strongly reject the AI hype around fully autonomous transformer agents, calling multi-hop LLM agents a boondoggle. Beyang articulates why deterministic search algorithms like A-star are necessary. | |
| The 'Normsky' Architecture: Merging Norvig and Chomsky | 3 | 5 | 2 | 1 | Beyang and Steve educate the hosts on the historic Chomsky versus Norvig AI debates and unveil Sourcegraph's hybrid 'NORMSKI' architectural framing. Steve shares firsthand perspective from 1990s Google on scaling Chomskyan systems. | |
| Data Pre-processing Moats and the Future of DSLs | 5 | 3 | 1 | 1 | Alessio asks an informed question regarding whether LLM code generation will eliminate the need for human-oriented DSLs. Steve and Beyang clarify that data pre-processing and compiler AST parsing represent their true competitive moat. | |
| Graph Protocols: LSP, Kythe, SCIP, and the BFG Indexer | 4 | 6 | 2 | 1 | Beyang and Steve deliver an in-depth schooling on graph indexing protocols, contrasting LSP's range-based limitations with Kythe's and SCIP's semantic graphs. They explain how their BFG engine eliminates type errors. | |
| Transformer Scaling Limits versus Algorithmic Search | 7 | 4 | 4 | 6 | Swyx directly challenges the guests' bearishness on agents by presenting compute scaling curves extrapolated to 2030. Beyang forcefully pushes back with a chess algorithm comparison, questioning whether pure transformer scaling ever replaces tree search. | |
| Sourcegraph's Production AI Stack and Database Architecture | 5 | 3 | 1 | 1 | Swyx probes into Sourcegraph's production stack, and Steve reveals they use Postgres and flat files rather than a dedicated graph database. Swyx demonstrates industry familiarity with Fireworks AI leadership. | |
| AI Tooling Wishlists, Synthetic Data, and OpenAI Dynamics | 5 | 2 | 1 | 1 | Swyx analyzes Replit's bounty data moat and compares it to OpenAI dependencies. The group discusses contingency plans during the OpenAI CEO firing weekend. | |
| Managing Codebase Complexity and Enterprise Productivity | 5 | 4 | 1 | 1 | Beyang explains how AI code generators exacerbate codebase complexity and why engineering managers need codebase-level understanding over raw code generation. Swyx shares a personal user story navigating a complex Twitter scraping repo via Cody Web. | |
| Lightning Round and Future Outlook | 4 | 3 | 2 | 1 | In the lightning round, the group explores multimodal coding workflows. Steve forcefully warns engineers who refuse to use coding assistants that they need to start planning another career. |