Sep 30, 2025 · 1h 4m · y-combinator
Anthropic Head of Pretraining on Scaling Laws, Compute, and the Future of AI · Y Combinator
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this Y Combinator podcast interview, host Ankit Gupta speaks with Nick Joseph, Head of Pre-Training at Anthropic, about the fundamental principles, hardware infrastructure, organizational strategies, and alignment challenges of building frontier AI models. Nick shares detailed insights on scaling laws, cluster engineering, web data dynamics, and the key technical skills needed to build superintelligent systems.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →
speaking balance: gold is the partners, purple is the guest (3 minute bins)
Nick candidly rejects the culture of academic research labs like FAIR, asserting that external critics missed the big picture and made nonsensical arguments against obvious empirical scaling laws.
Hardest push from the partners ▶ 32:42 Ankit challenging data scarcity with PageRank filteringAnkit pushes back against the premise that finding useful internet data is unsolvable by proposing that standard PageRank thresholding should cleanly separate useful training data from web noise.
Biggest teaching moment ▶ 32:55 Nick explaining why web graph metrics fail for AI trainingNick educates the host on why PageRank and click-based graph metrics fail to capture training utility, explaining that the true frontier lies in the unlinked long tail of internet text.
The partners hold their own ▶ 55:49 Ankit challenging compute scaling with architectural optimizationsAnkit demonstrates deep industry awareness by citing specific architectural modifications from open-source Chinese labs, such as attention caching and custom attention kernels, to challenge whether scaling alone suffices.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The partners as informed peer | Guest teaching | Guest disagreement | The partners pushing back | Why |
|---|---|---|---|---|---|---|
| Nick Joseph's Background: Vicarious, OpenAI, and Early AI Safety | 3 | 3 | 1 | 0 | Ankit Gupta opens the interview exploring Nick Joseph's background across GiveWell, Vicarious, and OpenAI. Nick explains his transition from economics and charity evaluation to AI safety. | |
| Defining Pre-Training, Next-Word Prediction, and Scaling Laws | 4 | 4 | 1 | 1 | Ankit asks why autoregressive next-token prediction won out over masked language modeling like BERT. Nick emphasizes that autoregressive training allows direct sampling for products and provides an empirical scaling objective. | |
| Scaling Laws, Hyperparameters, and Small-Scale Testing | 4 | 4 | 1 | 1 | Ankit explores hyperparameter selection and neural architecture search. Nick explains how empirical scaling laws reveal power-law curves and how scaling compute dominates hyperparameter fine-tuning. | |
| Early Infrastructure, Cloud Providers, and Cluster Topologies | 4 | 5 | 2 | 1 | Ankit asks about Anthropic's early infrastructure when it was a small team. Nick reveals that they wrote custom all-reduce operations because they aimed to scale beyond what existing frameworks like PyTorch supported at Facebook. | |
| Scaling Conviction vs. Academic Research Lab Culture | 4 | 5 | 3 | 1 | Nick expresses surprise at external skepticism toward scaling laws, criticizing academic research culture at places like FAIR for ignoring empirical scaling. He then details hardware profiling and Model FLOPs Utilization (MFU). | |
| Pair Programming, Internal Knowledge, and Debugging Tools | 4 | 4 | 1 | 1 | Nick describes learning profiling and debugging through pair programming with Tom Brown and Sam McCandlish. He contrasts specialist and generalist staffing strategies in pre-training teams. | |
| Distributed Failure Domains and Data Center Scale Challenges | 4 | 5 | 1 | 1 | Ankit asks about unexpected challenges at scale. Nick highlights single failure domains in distributed training runs and how hardware components in massive clusters frequently break. | |
| TPU vs. GPU Trade-offs in Pre-Training and Inference | 4 | 5 | 1 | 1 | Nick breaks down hardware bottlenecks, explaining why inference is high-bandwidth memory bound while pre-training is compute-bound. He discusses building minimal reproducers to collaborate with cloud providers on low-level bugs. | |
| The Evolution of Post-Training, RL, and Organizational Dynamics | 4 | 4 | 1 | 1 | Ankit inquires whether the rise of reinforcement learning shifts focus away from pre-training. Nick explains that RL introduces its own scaling laws and stresses the importance of avoiding organizational competition between teams. | |
| Internet Data Limits, PageRank, and Finding Quality Data | 4 | 6 | 3 | 2 | Ankit suggests PageRank could filter high-quality internet data. Nick pushes back, explaining that web popularity metrics do not correlate with what an AI model needs to learn long-tail knowledge. | |
| Synthetic Data, Distillation Risks, and Web Data Contamination | 4 | 4 | 1 | 1 | Ankit and Nick discuss synthetic data distillation and the risks of recursive training on LLM-generated web content, illustrated by how false data distributions compound. | |
| Evaluating Frontier Models, Loss Metrics, and Domain Benchmarks | 4 | 4 | 1 | 1 | Ankit discusses difficult evaluation domains like medical AI conversations. Nick argues that driving down next-token loss on high-quality doctor-patient transcripts remains an underappreciated and effective benchmark. | |
| Defining AI Alignment, Steering Models, and Democratic Values | 3 | 4 | 1 | 1 | Nick defines AI alignment both in terms of long-term AGI goal-sharing and near-term persona steering via Constitutional AI, using the analogy of building a steering wheel before choosing where to drive. | |
| Integrating Alignment into Pre-Training vs. Fast Post-Training Loops | 3 | 4 | 1 | 1 | Nick explains why alignment is predominantly handled in post-training due to fast iteration cycles, while noting that fundamental robustness may eventually require embedding alignment signals directly into pre-training. | |
| Future AI Paradigm Shifts and the Risks of Unnoticed Machine Learning Bugs | 3 | 4 | 1 | 1 | Nick discusses future risks, noting that unnoticed low-level machine learning bugs are what scare him most because they can derail entire multi-month training runs. | |
| Full-Stack Debugging: From High-Level ML Dynamics to Network Packets | 4 | 5 | 1 | 1 | Ankit describes classic architecture wiring bugs. Nick expands on the challenge of full-stack debugging, from high-level learning dynamics down to low-level kernel precision casting and network packet protocols. | |
| Hiring Strategy: Systems Engineers, Physicists, and Deep Technical Problem Solvers | 3 | 4 | 2 | 1 | Nick dispels the idea that pre-training teams are primarily ML research PhDs, emphasizing that frontier labs critically need systems engineers and problem-solving physicists capable of deep technical debugging. | |
| Alternative Architectures vs. Reliable Scaling of Transformers | 4 | 5 | 2 | 1 | Ankit brings up alternative architectures and optimizations from open-source labs. Nick explains that reliable transformer scaling consistently outperforms novel architectures, and highlights how pre-training co-designs models around inference constraints. | |
| Scaling Compute and the Challenges of Infinite Resources | 4 | 4 | 1 | 1 | Ankit and Nick discuss operating under 100x compute scaling. Nick highlights startup opportunities in automated cluster chip verification and cautions against building heavy scaffolding around current model limitations. | |
| Career Advice for Entering the AI Workforce | 2 | 3 | 1 | 0 | Nick offers career advice for entering the AI industry, recommending that newcomers focus on deep systems engineering and preparing for post-AGI governance rather than classical ML math. |