Sep 25, 2023 · 20m · a16z

Improving AI with Anthropic's Dario Amodei

Dario Amodei · 14m spoken Anjney Midha · 4m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this a16z podcast interview, Anthropic CEO Dario Amodei discusses the empirical power of AI scaling laws, Anthropic's physics-driven hiring philosophy, and their frameworks for constitutional safety and enterprise model deployment.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The host as informed peer 2.7 Guest teaching 5.2 Guest disagreement 1.8 The host pushing back 1.3
05100:0010:0020:000:40–3:59 · The host as informed peer 2/10 Dario Amodei's Background and Founding Anthropic The host opens with an anecdotal memory about funding and prompts Dario on his career trajectory. Dario explains his physics and neuroscience background and details how early bad GPT-2 translations signaled underlying architectural potential.3:59–7:54 · The host as informed peer 2/10 GPT-3 Breakthroughs and Emergent Code Reasoning The host prompts Dario on GPT-3 breakthroughs and scaling bottlenecks over the coming years. Dario explains how minimal Python training data yielded emergent code reasoning and details projected model training costs scaling to billions.7:54–10:09 · The host as informed peer 3/10 Inference Economics and Why Physicists Excel in Frontier AI The host asks about inference cost economics and the team's strong physics bias. Dario reframes the talent dynamic, explaining how fast-moving fields favor talented generalists over traditional experts where accumulated domain knowledge can be a hindrance.10:09–14:50 · The host as informed peer 3/10 Maintaining Talent Density While Scaling Anthropic The host questions how Dario maintains talent density while scaling, and directly challenges whether Anthropic imposes its own values via Constitutional AI. Dario explains RLHF limitations and argues they rely on universal principles like the UN Declaration of Human Rights.14:50–18:09 · The host as informed peer 3/10 Addressing the Safety Paradox and Safe Scaling Checkpoints The host identifies the central paradox of Anthropic scaling models rapidly while advocating for safety. Dario defends this by dismissing purely theoretical safety research as ineffective, arguing that safety solutions require increasingly capable AI systems.18:09–20:27 · The host as informed peer 3/10 Enterprise Context Windows and Developer Ecosystem Outlook The host asks practical ecosystem questions regarding context windows and presses on when infinite context windows will arrive. Dario dismisses the idea of infinite context windows by pointing out hard compute cost constraints.0:40–3:59 · Guest teaching 4/10 Dario Amodei's Background and Founding Anthropic The host opens with an anecdotal memory about funding and prompts Dario on his career trajectory. Dario explains his physics and neuroscience background and details how early bad GPT-2 translations signaled underlying architectural potential.3:59–7:54 · Guest teaching 5/10 GPT-3 Breakthroughs and Emergent Code Reasoning The host prompts Dario on GPT-3 breakthroughs and scaling bottlenecks over the coming years. Dario explains how minimal Python training data yielded emergent code reasoning and details projected model training costs scaling to billions.7:54–10:09 · Guest teaching 6/10 Inference Economics and Why Physicists Excel in Frontier AI The host asks about inference cost economics and the team's strong physics bias. Dario reframes the talent dynamic, explaining how fast-moving fields favor talented generalists over traditional experts where accumulated domain knowledge can be a hindrance.10:09–14:50 · Guest teaching 5/10 Maintaining Talent Density While Scaling Anthropic The host questions how Dario maintains talent density while scaling, and directly challenges whether Anthropic imposes its own values via Constitutional AI. Dario explains RLHF limitations and argues they rely on universal principles like the UN Declaration of Human Rights.14:50–18:09 · Guest teaching 6/10 Addressing the Safety Paradox and Safe Scaling Checkpoints The host identifies the central paradox of Anthropic scaling models rapidly while advocating for safety. Dario defends this by dismissing purely theoretical safety research as ineffective, arguing that safety solutions require increasingly capable AI systems.18:09–20:27 · Guest teaching 5/10 Enterprise Context Windows and Developer Ecosystem Outlook The host asks practical ecosystem questions regarding context windows and presses on when infinite context windows will arrive. Dario dismisses the idea of infinite context windows by pointing out hard compute cost constraints.0:40–3:59 · Guest disagreement 1/10 Dario Amodei's Background and Founding Anthropic The host opens with an anecdotal memory about funding and prompts Dario on his career trajectory. Dario explains his physics and neuroscience background and details how early bad GPT-2 translations signaled underlying architectural potential.3:59–7:54 · Guest disagreement 1/10 GPT-3 Breakthroughs and Emergent Code Reasoning The host prompts Dario on GPT-3 breakthroughs and scaling bottlenecks over the coming years. Dario explains how minimal Python training data yielded emergent code reasoning and details projected model training costs scaling to billions.7:54–10:09 · Guest disagreement 2/10 Inference Economics and Why Physicists Excel in Frontier AI The host asks about inference cost economics and the team's strong physics bias. Dario reframes the talent dynamic, explaining how fast-moving fields favor talented generalists over traditional experts where accumulated domain knowledge can be a hindrance.10:09–14:50 · Guest disagreement 2/10 Maintaining Talent Density While Scaling Anthropic The host questions how Dario maintains talent density while scaling, and directly challenges whether Anthropic imposes its own values via Constitutional AI. Dario explains RLHF limitations and argues they rely on universal principles like the UN Declaration of Human Rights.14:50–18:09 · Guest disagreement 3/10 Addressing the Safety Paradox and Safe Scaling Checkpoints The host identifies the central paradox of Anthropic scaling models rapidly while advocating for safety. Dario defends this by dismissing purely theoretical safety research as ineffective, arguing that safety solutions require increasingly capable AI systems.18:09–20:27 · Guest disagreement 2/10 Enterprise Context Windows and Developer Ecosystem Outlook The host asks practical ecosystem questions regarding context windows and presses on when infinite context windows will arrive. Dario dismisses the idea of infinite context windows by pointing out hard compute cost constraints.0:40–3:59 · The host pushing back 0/10 Dario Amodei's Background and Founding Anthropic The host opens with an anecdotal memory about funding and prompts Dario on his career trajectory. Dario explains his physics and neuroscience background and details how early bad GPT-2 translations signaled underlying architectural potential.3:59–7:54 · The host pushing back 0/10 GPT-3 Breakthroughs and Emergent Code Reasoning The host prompts Dario on GPT-3 breakthroughs and scaling bottlenecks over the coming years. Dario explains how minimal Python training data yielded emergent code reasoning and details projected model training costs scaling to billions.7:54–10:09 · The host pushing back 1/10 Inference Economics and Why Physicists Excel in Frontier AI The host asks about inference cost economics and the team's strong physics bias. Dario reframes the talent dynamic, explaining how fast-moving fields favor talented generalists over traditional experts where accumulated domain knowledge can be a hindrance.10:09–14:50 · The host pushing back 3/10 Maintaining Talent Density While Scaling Anthropic The host questions how Dario maintains talent density while scaling, and directly challenges whether Anthropic imposes its own values via Constitutional AI. Dario explains RLHF limitations and argues they rely on universal principles like the UN Declaration of Human Rights.14:50–18:09 · The host pushing back 2/10 Addressing the Safety Paradox and Safe Scaling Checkpoints The host identifies the central paradox of Anthropic scaling models rapidly while advocating for safety. Dario defends this by dismissing purely theoretical safety research as ineffective, arguing that safety solutions require increasingly capable AI systems.18:09–20:27 · The host pushing back 2/10 Enterprise Context Windows and Developer Ecosystem Outlook The host asks practical ecosystem questions regarding context windows and presses on when infinite context windows will arrive. Dario dismisses the idea of infinite context windows by pointing out hard compute cost constraints.

speaking balance: gold is the host, purple is the guest (3 minute bins)

0:00 · the host 0% · guest 100%0:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%18:00 · the host 0% · guest 100%18:00 · the host 0% · guest 100%
Sharpest disagreement ▶ 15:11 Dismissal of external theoretical safety approaches

Dario forcefully criticizes purely theoretical safety researchers, stating their efforts separate from capabilities development have been largely unsuccessful.

Hardest push from the host ▶ 13:21 Challenging Constitutional AI moral imposition

Anjney Midha directly pushes back on Dario, asking how Anthropic grapples with the critique that they are imposing their own personal values on AI systems.

Biggest teaching moment ▶ 8:58 Reframing domain expertise vs generalist talent

Dario educates the host on field dynamics, demonstrating why deep domain expertise in legacy ML can actually be a disadvantage compared to raw physics background.

The host holds their own ▶ 14:50 Articulating the Anthropic safety paradox

The host sharply formulates the core contradiction of Anthropic accelerating scaling while simultaneously advocating for safety caution.

the scores for every segment, with the reasoning behind each
ChapterTopicThe host as informed peerGuest teachingGuest disagreementThe host pushing backWhy
Dario Amodei's Background and Founding Anthropic 2410 The host opens with an anecdotal memory about funding and prompts Dario on his career trajectory. Dario explains his physics and neuroscience background and details how early bad GPT-2 translations signaled underlying architectural potential.
GPT-3 Breakthroughs and Emergent Code Reasoning 2510 The host prompts Dario on GPT-3 breakthroughs and scaling bottlenecks over the coming years. Dario explains how minimal Python training data yielded emergent code reasoning and details projected model training costs scaling to billions.
Inference Economics and Why Physicists Excel in Frontier AI 3621 The host asks about inference cost economics and the team's strong physics bias. Dario reframes the talent dynamic, explaining how fast-moving fields favor talented generalists over traditional experts where accumulated domain knowledge can be a hindrance.
Maintaining Talent Density While Scaling Anthropic 3523 The host questions how Dario maintains talent density while scaling, and directly challenges whether Anthropic imposes its own values via Constitutional AI. Dario explains RLHF limitations and argues they rely on universal principles like the UN Declaration of Human Rights.
Addressing the Safety Paradox and Safe Scaling Checkpoints 3632 The host identifies the central paradox of Anthropic scaling models rapidly while advocating for safety. Dario defends this by dismissing purely theoretical safety research as ineffective, arguing that safety solutions require increasingly capable AI systems.
Enterprise Context Windows and Developer Ecosystem Outlook 3522 The host asks practical ecosystem questions regarding context windows and presses on when infinite context windows will arrive. Dario dismisses the idea of infinite context windows by pointing out hard compute cost constraints.

Statements from this episode (15)

Prediction Not checkable as stated
Amodei: AI scaling laws will continue driving improvements without algorithmic breakthroughs
“I think we are on track, even if there were no algorithmic improvements from here, even if we just scaled up what we had so far, ah, I think the scaling laws are going to continue, and I think that's going to lead to amazing improvements.”
Dario Amodei Sep 25, 2023 ▶ 0:00
Insight
Amodei: Next-word prediction scaling has no limits and will continue
“Our view was that, look, this is the beginning of something amazing because there's no limit and you can continue to scale it up. And there's no reason why the patterns we've seen before won't continue to hold. The objective of predicting the next word is so r…”
Dario Amodei Sep 25, 2023 ▶ 3:28
Disclosure
Amodei: Python made up only 0.1% to 1% of GPT-3's training data
“When we looked through it was like, you know, it's hard to estimate, but it was something like, you know, maybe .1% to one percent of the data that we scraped was Python data.”
Dario Amodei Sep 25, 2023 ▶ 5:21
Prediction Not checkable as stated
Amodei: Scaling compute and coding data will inevitably advance AI reasoning
“We're getting more compute. We can scale up the models more and we can greatly increase the amount of data that, that is programming. So, so, so we have so many ways that we can amplify this. And so Of course it's gonna work. It's just a matter of time.”
Dario Amodei Sep 25, 2023 ▶ 5:40
Assertion Supported
Amodei: Most expensive AI models in 2023 cost around $100M to train
“The most expensive models made today cost about a hundred million dollars, say plus or minus a factor of two.”
Dario Amodei Sep 25, 2023 ▶ 6:54
Prediction Didn’t hold up
Amodei: AI training costs will hit $1B in 2024 and up to $10B in 2025
“I think that next year we're probably going to see from multiple players models on the order of one billion dollars, and in 2025 we're going to see models on the order of Several billion, I don't know, perhaps even ten billion dollars.”
Dario Amodei Sep 25, 2023 ▶ 7:02
Prediction Not checkable as stated
Amodei: AI inference will remain affordable over next 3-4 years
“I think my basic view is that inference will not get that much more expensive. The simple, you know, the basic logic of the scaling laws is that if you increase compute by a factor of n, you need to increase data by a factor of square root of n. And size the m…”
Dario Amodei Sep 25, 2023 ▶ 7:55
Insight
Amodei: Generalists outperform long-time domain experts in fast-moving AI
“AI was and still is to some extent very young and is definitely moving very fast. And so when that's the case, really talented generalists can often outperform those who have been in the field for a long time because things are being shaken up so much. If anyt…”
Dario Amodei Sep 25, 2023 ▶ 9:29
Insight
Dario Amodei: Talent density beats talent mass every time
“Our general view is, you know, talent density beats talent mass every time.”
Dario Amodei Sep 25, 2023 ▶ 10:53
Assertion Supported
Dario Amodei says he co-invented RLHF while working at OpenAI
“I was one of the, like, co-inventors of that at OpenAI, but since then it's been, you know, improved to power ChatGPT”
Dario Amodei Sep 25, 2023 ▶ 12:06
Disclosure
Amodei: Anthropic's core AI safety constitution is five pages long
“Our set of principles is in our constitution. It's very short. It's five pages.”
Dario Amodei Sep 25, 2023 ▶ 12:35
Prediction Held up
Anthropic plans to formalize a safe scaling framework with capability checkpoints
“One broad way that we've been thinking about it from the beginning and are probably going to work more on formalizing it over the coming months and years is this idea of sort of safe scaling or checkpointing. So there could be this kind of alternating step whe…”
Dario Amodei Sep 25, 2023 ▶ 16:49
Insight
Amodei: Bureaucratic AI regulation will cause the West to lose to adversaries
“If this is like, you have to fill out a thousand pages of paperwork and get 15 different licenses from different bodies to, you know, to make an AI system, that's never gonna work, that's gonna slow things down, you know, other adversaries You know, authoritar…”
Dario Amodei Sep 25, 2023 ▶ 17:22
Insight
Amodei: Long context windows and search retrieval remain heavily underappreciated
“One thing that I think people are starting to realize, but I think is, is still underappreciated is the longer context and things that come along with that, that we're working on, you know, things in the direction of, you know, retrieval or search really open …”
Dario Amodei Sep 25, 2023 ▶ 18:39
Prediction Not checkable as stated
Amodei: AI models will never have infinite context windows due to compute
“Really the main thing holding back infinite context windows is just, you know, as you make the context window longer and longer, of course, the majority of the compute starts to be in the context window. So at some point it just becomes too, too expensive in t…”
Dario Amodei Sep 25, 2023 ▶ 19:59
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.