Sep 25, 2023 · 17m · a16z

Universally Accessible Intelligence with Character.ai's Noam Shazeer

Noam Shazeer · 11m spoken Sarah Wang · 3m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In an a16z stage interview, Character.AI Co-founder and CEO Noam Shazeer discusses how consumer-driven conversational AI accelerates the path to Artificial General Intelligence (AGI). Alongside a live interactive comparison with an AI chatbot trained on his persona, Shazeer explores compute scaling, full-stack model architecture, and the transition into an era of universally accessible intelligence.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The host as informed peer 3.5 Guest teaching 4.3 Guest disagreement 1.8 The host pushing back 1.8
05100:0010:000:10–3:52 · The host as informed peer 2/10 Title Sequence and Important Disclosures Sarah introduces the guest and asks standard friendly icebreaker questions regarding Duke and leaving Google. Noam explains early LLM developments at Google, internal corporate renaming, and leaving due to big company brand risk.3:52–7:04 · The host as informed peer 3/10 Question 3: Scaling Laws and AGI Existential Risk Sarah connects Noam's scaling perspectives with other industry leaders like Mira Murati and Dario Amodei. Noam dismisses safety pause doom-mongering with a joke about needing four months to acquire more H100s and outlines progress in training costs.7:04–10:07 · The host as informed peer 4/10 Consumer Engagement and Parasocial AI Relationships Sarah presents precise platform engagement metrics including 20 billion messages and 2-hour daily usage averages. Noam reframes the entertainment industry as parasocial imaginary friends and argues model hallucination is a feature rather than a flaw for companions.10:07–12:38 · The host as informed peer 5/10 Generalist vs Specialised Domain AI Models Sarah uses a16z startup insights to question Character's generalist model approach compared to specialized edtech or mental health startups. Noam explains why specialized rule-tuning fails to generalize and defends full-stack model development.12:38–16:03 · The host as informed peer 4/10 Theory of Mind and Global Compute Scaling Sarah quotes research on Theory of Mind and references Noam's quote on universally accessible intelligence. Noam performs back-of-the-envelope calculations on global Nvidia H100 output to demonstrate compute accessibility.16:03–17:12 · The host as informed peer 3/10 Fundamental Breakthroughs and Compute Value Sarah pushes on whether current Transformer scaling is sufficient or if new breakthroughs are needed to reach AGI. Noam highlights that operational compute costs are vastly lower than human time value.0:10–3:52 · Guest teaching 3/10 Title Sequence and Important Disclosures Sarah introduces the guest and asks standard friendly icebreaker questions regarding Duke and leaving Google. Noam explains early LLM developments at Google, internal corporate renaming, and leaving due to big company brand risk.3:52–7:04 · Guest teaching 4/10 Question 3: Scaling Laws and AGI Existential Risk Sarah connects Noam's scaling perspectives with other industry leaders like Mira Murati and Dario Amodei. Noam dismisses safety pause doom-mongering with a joke about needing four months to acquire more H100s and outlines progress in training costs.7:04–10:07 · Guest teaching 5/10 Consumer Engagement and Parasocial AI Relationships Sarah presents precise platform engagement metrics including 20 billion messages and 2-hour daily usage averages. Noam reframes the entertainment industry as parasocial imaginary friends and argues model hallucination is a feature rather than a flaw for companions.10:07–12:38 · Guest teaching 4/10 Generalist vs Specialised Domain AI Models Sarah uses a16z startup insights to question Character's generalist model approach compared to specialized edtech or mental health startups. Noam explains why specialized rule-tuning fails to generalize and defends full-stack model development.12:38–16:03 · Guest teaching 6/10 Theory of Mind and Global Compute Scaling Sarah quotes research on Theory of Mind and references Noam's quote on universally accessible intelligence. Noam performs back-of-the-envelope calculations on global Nvidia H100 output to demonstrate compute accessibility.16:03–17:12 · Guest teaching 4/10 Fundamental Breakthroughs and Compute Value Sarah pushes on whether current Transformer scaling is sufficient or if new breakthroughs are needed to reach AGI. Noam highlights that operational compute costs are vastly lower than human time value.0:10–3:52 · Guest disagreement 2/10 Title Sequence and Important Disclosures Sarah introduces the guest and asks standard friendly icebreaker questions regarding Duke and leaving Google. Noam explains early LLM developments at Google, internal corporate renaming, and leaving due to big company brand risk.3:52–7:04 · Guest disagreement 2/10 Question 3: Scaling Laws and AGI Existential Risk Sarah connects Noam's scaling perspectives with other industry leaders like Mira Murati and Dario Amodei. Noam dismisses safety pause doom-mongering with a joke about needing four months to acquire more H100s and outlines progress in training costs.7:04–10:07 · Guest disagreement 3/10 Consumer Engagement and Parasocial AI Relationships Sarah presents precise platform engagement metrics including 20 billion messages and 2-hour daily usage averages. Noam reframes the entertainment industry as parasocial imaginary friends and argues model hallucination is a feature rather than a flaw for companions.10:07–12:38 · Guest disagreement 2/10 Generalist vs Specialised Domain AI Models Sarah uses a16z startup insights to question Character's generalist model approach compared to specialized edtech or mental health startups. Noam explains why specialized rule-tuning fails to generalize and defends full-stack model development.12:38–16:03 · Guest disagreement 1/10 Theory of Mind and Global Compute Scaling Sarah quotes research on Theory of Mind and references Noam's quote on universally accessible intelligence. Noam performs back-of-the-envelope calculations on global Nvidia H100 output to demonstrate compute accessibility.16:03–17:12 · Guest disagreement 1/10 Fundamental Breakthroughs and Compute Value Sarah pushes on whether current Transformer scaling is sufficient or if new breakthroughs are needed to reach AGI. Noam highlights that operational compute costs are vastly lower than human time value.0:10–3:52 · The host pushing back 1/10 Title Sequence and Important Disclosures Sarah introduces the guest and asks standard friendly icebreaker questions regarding Duke and leaving Google. Noam explains early LLM developments at Google, internal corporate renaming, and leaving due to big company brand risk.3:52–7:04 · The host pushing back 2/10 Question 3: Scaling Laws and AGI Existential Risk Sarah connects Noam's scaling perspectives with other industry leaders like Mira Murati and Dario Amodei. Noam dismisses safety pause doom-mongering with a joke about needing four months to acquire more H100s and outlines progress in training costs.7:04–10:07 · The host pushing back 2/10 Consumer Engagement and Parasocial AI Relationships Sarah presents precise platform engagement metrics including 20 billion messages and 2-hour daily usage averages. Noam reframes the entertainment industry as parasocial imaginary friends and argues model hallucination is a feature rather than a flaw for companions.10:07–12:38 · The host pushing back 3/10 Generalist vs Specialised Domain AI Models Sarah uses a16z startup insights to question Character's generalist model approach compared to specialized edtech or mental health startups. Noam explains why specialized rule-tuning fails to generalize and defends full-stack model development.12:38–16:03 · The host pushing back 1/10 Theory of Mind and Global Compute Scaling Sarah quotes research on Theory of Mind and references Noam's quote on universally accessible intelligence. Noam performs back-of-the-envelope calculations on global Nvidia H100 output to demonstrate compute accessibility.16:03–17:12 · The host pushing back 2/10 Fundamental Breakthroughs and Compute Value Sarah pushes on whether current Transformer scaling is sufficient or if new breakthroughs are needed to reach AGI. Noam highlights that operational compute costs are vastly lower than human time value.

speaking balance: gold is the host, purple is the guest (3 minute bins)

0:00 · the host 0% · guest 100%0:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%
Sharpest disagreement ▶ 4:16 Rejection of the global AGI safety pause premise

Noam cheekily deflects the existential risk question by proposing a pause only long enough to bring more H100 GPUs online.

Hardest push from the host ▶ 10:07 Questioning generalist models versus niche domain models

Sarah uses VC market data from a16z to challenge Noam on why a single model beats domain-specific startups in areas like mental health.

Biggest teaching moment ▶ 13:15 Fermi estimation of global compute capacity

Noam breaks down exact hardware counts and floating-point operations per second to educate the host on global compute math per human.

The host holds their own ▶ 3:53 Contextualizing scaling laws across industry peers

Sarah demonstrates domain expertise by framing Noam's take on scaling laws alongside identical positions from Mira Murati at OpenAI and Dario Amodei at Anthropic.

the scores for every segment, with the reasoning behind each
ChapterTopicThe host as informed peerGuest teachingGuest disagreementThe host pushing backWhy
Title Sequence and Important Disclosures 2321 Sarah introduces the guest and asks standard friendly icebreaker questions regarding Duke and leaving Google. Noam explains early LLM developments at Google, internal corporate renaming, and leaving due to big company brand risk.
Question 3: Scaling Laws and AGI Existential Risk 3422 Sarah connects Noam's scaling perspectives with other industry leaders like Mira Murati and Dario Amodei. Noam dismisses safety pause doom-mongering with a joke about needing four months to acquire more H100s and outlines progress in training costs.
Consumer Engagement and Parasocial AI Relationships 4532 Sarah presents precise platform engagement metrics including 20 billion messages and 2-hour daily usage averages. Noam reframes the entertainment industry as parasocial imaginary friends and argues model hallucination is a feature rather than a flaw for companions.
Generalist vs Specialised Domain AI Models 5423 Sarah uses a16z startup insights to question Character's generalist model approach compared to specialized edtech or mental health startups. Noam explains why specialized rule-tuning fails to generalize and defends full-stack model development.
Theory of Mind and Global Compute Scaling 4611 Sarah quotes research on Theory of Mind and references Noam's quote on universally accessible intelligence. Noam performs back-of-the-envelope calculations on global Nvidia H100 output to demonstrate compute accessibility.
Fundamental Breakthroughs and Compute Value 3412 Sarah pushes on whether current Transformer scaling is sufficient or if new breakthroughs are needed to reach AGI. Noam highlights that operational compute costs are vastly lower than human time value.

Statements from this episode (16)

Assertion Not checkable as stated
Shazeer: Dialogue is the world's number one pastime
“Dialogue, like it's world's number one pastime”
Noam Shazeer Sep 25, 2023 ▶ 2:43
Opinion
Shazeer is not currently afraid of AGI destroying humanity
“No. Not yet. Not yet. I think there's a lot of possibility of you know, a lot of potential benefits and yeah, we're going to work on it as the technology improves.”
Noam Shazeer Sep 25, 2023 ▶ 4:31
Disclosure
Character.ai's model cost $2M to train, now $500k
“The model we're serving now, we, you know, cost us about, like, two million dollars worth of compute cycles to train last year, and could probably repeat it for, like, half a million now.”
Noam Shazeer Sep 25, 2023 ▶ 5:43
Prediction Not checkable as stated
Character.ai will launch a significantly smarter model by late 2023
“We're going to launch something, tens of IQ points smarter hopefully by the by the end of the year.”
Noam Shazeer Sep 25, 2023 ▶ 5:53
Insight
Entertainment is a $2T industry built on imaginary friends
“Entertainment is like this two trillion dollar a year industry. And like the dirty secret is that entertainment Is imaginary friends that don't know you exist.”
Noam Shazeer Sep 25, 2023 ▶ 7:34
Insight
LLM hallucination is a feature, not a bug, for AI companions
“Like you want to launch something that's a doctor, It's going to be a lot slower because you want to be really, really, really careful about not providing like false information, but friend, you can do like really fast. Like it's just entertainment. It makes t…”
Noam Shazeer Sep 25, 2023 ▶ 8:21
Prediction Not checkable as stated
AI companion entertainment is ready for immediate explosive adoption
“If I, like, I want to push this technology ahead fast, like, that's what I want to go with, because like, A, it's, you know, it's ready for an explosion, like right now, not like in five years when we solve all the problems, but like, now.”
Noam Shazeer Sep 25, 2023 ▶ 9:11
Insight
Narrow AI use cases tempt builders into non-generalizable rules
“The more you get to like mission critical, you know, particular use case, the more you get tempted into like writing particular rules and like doing things that will not generalize well.”
Noam Shazeer Sep 25, 2023 ▶ 10:44
Prediction Not checkable as stated
Character.AI aims for AGI by building consumer products at scale
“Our goal is to be like an AGI company and a product first company. And the way to do that is by picking the right product that forces us to work on the right things, things that generalize, make the model smarter, make it, like, do what people, you know, what …”
Noam Shazeer Sep 25, 2023 ▶ 11:00
Disclosure
Frustration over inability to launch AI products drove Google exit
“You know, some people are motivated by publishing, you know, like, I was frustrated I couldn't launch at Google.”
Noam Shazeer Sep 25, 2023 ▶ 12:29
Prediction Not checkable as stated
Theory of mind in AI is an emergent property of scale
“Yeah, just make the thing smarter. It's gonna have a better theory of mind, and I think that's definitely something massively, massively important. It seems like one of these emergent properties that just Is gonna, going to come with scale.”
Noam Shazeer Sep 25, 2023 ▶ 13:03
Prediction Not checkable as stated
AI will improve massively from scaling without new breakthroughs
“And without any breakthroughs, like it's going to get like massively better as everyone just Kind of scales up to use it, and there will be more breakthroughs because now, you know, like all the scientists in the world are like working on like making this stuf…”
Noam Shazeer Sep 25, 2023 ▶ 14:59
Prediction Not checkable as stated
Top-tier AI capabilities will reach academic labs in a few years
“Like, you know, we're going to see like a huge amount of innovation and, you know, what, what's possible in the largest companies now can be possible in, you know, in somebody's academic lab or garage in a few years.”
Noam Shazeer Sep 25, 2023 ▶ 15:16
Prediction Not checkable as stated
AI capable of curing cancer is only a few years away
“I'd love to get to the point where you can just ask it how to cure cancer or something, you know. I mean, it's, you know, it seems a few years away for now, but, you know, like.”
Noam Shazeer Sep 25, 2023 ▶ 15:43
Assertion Not checkable as stated
AI scaling laws show no signs of stopping
“I don't think anyone's seen like these scaling laws, you know, stop. I think as far as anybody has experimented, stuff just gets, keeps, keeps getting smarter.”
Noam Shazeer Sep 25, 2023 ▶ 16:05
Assertion Supported
Individual compute operations currently cost around $10^-18
“Like operations cost like 10 to the -18 dollars these days”
Noam Shazeer Sep 25, 2023 ▶ 16:47
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.