Aug 31, 2023 · 36m · news
Noam Shazeer: How We Spent $2M to Train a Single AI Model and Grew Character.ai to 20M Users | E1055 · 20VC with Harry Stebbings
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this episode of TwentyVC, host Harry Stebbings interviews Noam Shazeer, co-founder and CEO of Character.ai, about his journey from Google veteran to leading a massive consumer AI startup. Shazeer shares his philosophies on building full-stack general-purpose systems, the scaling laws of model training, and the positive societal potential of AI companionship.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Harry holds 24.2% of the talking time here. How this is scored →
speaking balance: gold is Harry, purple is the guest (3 minute bins)
Noam forcefully rejects traditional product strategy advice, stating he explicitly refuses to hire product managers who advocate narrowing the product into specialized verticals.
Hardest push from Harry ▶ 11:23 Challenging Horizontal vs Vertical Model QualityHarry directly challenges Noam's belief in horizontal generalist models, refusing his premise and asking how quality can match specialized vertical tools.
Biggest teaching moment ▶ 34:00 Dense Hardware vs Sparse Memory MechanicsNoam provides a technical breakdown explaining why early deep learning sparse computation failed due to hardware mechanics optimized for dense matrix multiplication.
Harry holds his own ▶ 11:23 Applying Tech Strategy Frameworks to Model GeneralizationHarry demonstrates strong understanding of product strategy by applying classic software specialization trade-offs to challenge Noam's model architecture decisions.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | Harry as informed peer | Guest teaching | Guest disagreement | Harry pushing back | Why |
|---|---|---|---|---|---|---|
| Noam's Google Origin Story: Rebuilding Google's Spelling Corrector | 2 | 3 | 1 | 1 | Harry asks conversational background questions about Noam's 20-year career at Google and the spelling corrector project. Noam explains how Google's early 50,000-word third-party dictionary was ill-suited for web search queries compared to word processors. | |
| The Full-Stack Direct-to-Consumer Strategy of Character.ai | 2 | 4 | 2 | 2 | Harry prompts Noam on key takeaways from Google and how they shaped Character.ai. Noam explains his preference for full-stack B2C systems over B2B foundational model stacks, gently contrasting his vision with standard VC consensus. | |
| Character.ai's Mission: Humility and User Agency | 3 | 4 | 2 | 4 | Harry presses Noam on what specific mission statement he uses to motivate his team rather than general corporate platitudes. Noam reframes company leadership around user agency and unexpected emergent use cases like therapeutic video game bots. | |
| Three Core Elements Driving Character.ai's Growth | 4 | 3 | 3 | 6 | Harry directly challenges Noam's optimism regarding AI companionship, asking whether talking to AI bots replaces real human interaction or creates unhealthy habits. Noam maintains that AI conversations serve as valuable practice for individuals with social anxiety. | |
| The Product Challenge: Rejecting Vertical Specialization for Generality | 5 | 7 | 4 | 6 | Noam dismisses standard product management advice about narrowing focus to specific verticals, stating he refuses to hire PMs who advocate that strategy. Harry pushes back firmly, questioning how a generalist model can compete on quality against specialized vertical solutions. | |
| The Evolution of Language Modeling Since 2015 | 4 | 6 | 2 | 3 | Harry asks well-informed questions about market hype cycles and the trade-offs between model size and data volume. Noam clarifies that total compute operations, rather than raw model parameters or dataset sizes, represent the primary bottleneck in model capabilities. | |
| Training the $2 Million AI Model | 4 | 5 | 2 | 3 | Harry inquires about model constraints and proprietary data moats across different domains. Noam uses an analogy of human general education versus job-specific training to explain pretraining versus task fine-tuning. | |
| Why Startups Outpace Incumbents in AI Innovation | 5 | 4 | 3 | 5 | Harry presses Noam to take a stance on whether startups or incumbents will dominate AI innovation, citing arguments from industry figures like Yann LeCun. Noam avoids taking a binary side, explaining the economic trade-offs between batch serving efficiency and startup agility. | |
| The Electricity Moment: Public Perception and Hallucinations | 4 | 5 | 2 | 3 | Harry asks about public misconceptions surrounding AI risk and whether hallucinations represent a flaw or a feature. Noam explains that for consumer, creative, and emotional support use cases, hallucinations act as a feature enabling imagination. | |
| Noam's Role as CEO: Prioritizing Utility Over Fun | 3 | 5 | 3 | 6 | Noam asserts that he prioritizes utility over personal fun in his role as CEO, prompting Harry to push back and share his own experience of finding fun and utility tightly linked. Noam shares his personal perspective on fatherhood, responsibility, and maturity. | |
| Quickfire Round: The Wright Brothers Moment of AI | 5 | 6 | 2 | 3 | Harry asks about identifying signal amidst noise in AI research, referencing notable AI pioneers like Yann LeCun and Yoshua Bengio. Noam compares current machine learning research to alchemy, noting that only empirical positive results cut through the noise. | |
| Sparse Computation vs. Dense Hardware: Noam's Biggest Mistake | 3 | 7 | 1 | 2 | Harry asks Noam about past misconceptions he held. Noam details his early attempts at sparse computation before realizing that modern hardware acceleration relies on dense matrix multiplication, which led to his groundbreaking work on Mixture of Experts. |