May 22, 2024 · 39m · no-priors

No Priors Ep. 65 | With Scale AI CEO Alexandr Wang

Alexandr Wang · 29m spoken Sarah Guo · 5m spoken Elad Gil · 2m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of No Priors, Scale AI founder and CEO Alexandr Wang joins Sarah Guo and Elad Gil to discuss the evolution of data infrastructure, the transition to expert-driven frontier data, rigorous AI benchmarking, and the iterative path toward AGI.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 20.9% of the talking time here. How this is scored →

The hosts as informed peer 4.2 Guest teaching 2.3 Guest disagreement 0.9 The hosts pushing back 1.0
05100:0010:0020:0030:002:32–7:45 · The hosts as informed peer 3/10 The Evolution of Scale AI: From Autonomous Vehicles to Generative AI Sarah sets the context recalling Scale's early days and asks Alex to describe the evolution of the business from autonomous vehicles to defense and LLM RLHF. Alex provides an expansive historical overview in an agreeable, narrative style.7:46–13:08 · The hosts as informed peer 4/10 Navigating Data Scarcity and Frontier Data Production Elad asks about emerging enterprise and sovereign AI demands and draws a parallel to Google's early mission to digitize books. Alex details the transition from scraping easy web data to producing high-signal frontier expert data.13:08–15:33 · The hosts as informed peer 4/10 Enterprise Proprietary Data and Hybrid Human-AI Synthetic Pipelines Sarah shares investor observations that incumbent enterprise software lacks the structured data required for AI models. Alex agrees, highlighting proprietary enterprise volume and introducing hybrid human-AI synthetic data pipelines.15:33–20:00 · The hosts as informed peer 6/10 Centaur Intelligence and the Durability of Human Expertise The hosts press Alex on the limits of centaur intelligence. Elad cites Google's Med-PaLM 2 outperforming general practitioners to ask when human expertise becomes obsolete, while Sarah probes architectural breakthroughs in planning.20:01–22:13 · The hosts as informed peer 2/10 Scale AI's $1B Fundraise and Strategic Industry Role Sarah congratulates Alex on Scale's $1B fundraise at a $14B valuation and inquires about strategic investors. Alex outlines Scale's role as neutral infrastructure across the entire AI stack.22:14–26:00 · The hosts as informed peer 4/10 Evaluating Frontier AI Systems and the GSM-1K Benchmark Sarah asks what makes evaluating frontier models difficult. Alex explains benchmark contamination and introduces Scale's GSM-1K study, which demonstrated that multiple models overfitted existing benchmarks.26:00–32:00 · The hosts as informed peer 5/10 Application Layer Dynamics and Scale's Product Announcements Alex discusses the application layer hype cycle around GPT-4 and Scale's enterprise/defense tooling. Elad references Meta research proving smaller, higher-quality datasets produce superior models.32:01–34:47 · The hosts as informed peer 5/10 Multimodality, Lab Convergence, and the Need for Smarter Models The hosts and guest discuss the rapid convergence between Google Astra and OpenAI GPT-4o. Sarah offers dual explanations of natural technical convergence versus competitive intelligence, and Alex laments the lack of fundamentally smarter reasoning models.34:47–38:39 · The hosts as informed peer 5/10 The Path to AGI and Organizational Agility Alex presents his contrarian view that AGI progress mirrors curing individual cancers rather than a single vaccine. Sarah pushes back on video world models as a general foundation, which Alex rejects as narrative lacking scientific evidence.2:32–7:45 · Guest teaching 2/10 The Evolution of Scale AI: From Autonomous Vehicles to Generative AI Sarah sets the context recalling Scale's early days and asks Alex to describe the evolution of the business from autonomous vehicles to defense and LLM RLHF. Alex provides an expansive historical overview in an agreeable, narrative style.7:46–13:08 · Guest teaching 3/10 Navigating Data Scarcity and Frontier Data Production Elad asks about emerging enterprise and sovereign AI demands and draws a parallel to Google's early mission to digitize books. Alex details the transition from scraping easy web data to producing high-signal frontier expert data.13:08–15:33 · Guest teaching 2/10 Enterprise Proprietary Data and Hybrid Human-AI Synthetic Pipelines Sarah shares investor observations that incumbent enterprise software lacks the structured data required for AI models. Alex agrees, highlighting proprietary enterprise volume and introducing hybrid human-AI synthetic data pipelines.15:33–20:00 · Guest teaching 3/10 Centaur Intelligence and the Durability of Human Expertise The hosts press Alex on the limits of centaur intelligence. Elad cites Google's Med-PaLM 2 outperforming general practitioners to ask when human expertise becomes obsolete, while Sarah probes architectural breakthroughs in planning.20:01–22:13 · Guest teaching 1/10 Scale AI's $1B Fundraise and Strategic Industry Role Sarah congratulates Alex on Scale's $1B fundraise at a $14B valuation and inquires about strategic investors. Alex outlines Scale's role as neutral infrastructure across the entire AI stack.22:14–26:00 · Guest teaching 4/10 Evaluating Frontier AI Systems and the GSM-1K Benchmark Sarah asks what makes evaluating frontier models difficult. Alex explains benchmark contamination and introduces Scale's GSM-1K study, which demonstrated that multiple models overfitted existing benchmarks.26:00–32:00 · Guest teaching 2/10 Application Layer Dynamics and Scale's Product Announcements Alex discusses the application layer hype cycle around GPT-4 and Scale's enterprise/defense tooling. Elad references Meta research proving smaller, higher-quality datasets produce superior models.32:01–34:47 · Guest teaching 1/10 Multimodality, Lab Convergence, and the Need for Smarter Models The hosts and guest discuss the rapid convergence between Google Astra and OpenAI GPT-4o. Sarah offers dual explanations of natural technical convergence versus competitive intelligence, and Alex laments the lack of fundamentally smarter reasoning models.34:47–38:39 · Guest teaching 3/10 The Path to AGI and Organizational Agility Alex presents his contrarian view that AGI progress mirrors curing individual cancers rather than a single vaccine. Sarah pushes back on video world models as a general foundation, which Alex rejects as narrative lacking scientific evidence.2:32–7:45 · Guest disagreement 0/10 The Evolution of Scale AI: From Autonomous Vehicles to Generative AI Sarah sets the context recalling Scale's early days and asks Alex to describe the evolution of the business from autonomous vehicles to defense and LLM RLHF. Alex provides an expansive historical overview in an agreeable, narrative style.7:46–13:08 · Guest disagreement 0/10 Navigating Data Scarcity and Frontier Data Production Elad asks about emerging enterprise and sovereign AI demands and draws a parallel to Google's early mission to digitize books. Alex details the transition from scraping easy web data to producing high-signal frontier expert data.13:08–15:33 · Guest disagreement 0/10 Enterprise Proprietary Data and Hybrid Human-AI Synthetic Pipelines Sarah shares investor observations that incumbent enterprise software lacks the structured data required for AI models. Alex agrees, highlighting proprietary enterprise volume and introducing hybrid human-AI synthetic data pipelines.15:33–20:00 · Guest disagreement 3/10 Centaur Intelligence and the Durability of Human Expertise The hosts press Alex on the limits of centaur intelligence. Elad cites Google's Med-PaLM 2 outperforming general practitioners to ask when human expertise becomes obsolete, while Sarah probes architectural breakthroughs in planning.20:01–22:13 · Guest disagreement 0/10 Scale AI's $1B Fundraise and Strategic Industry Role Sarah congratulates Alex on Scale's $1B fundraise at a $14B valuation and inquires about strategic investors. Alex outlines Scale's role as neutral infrastructure across the entire AI stack.22:14–26:00 · Guest disagreement 1/10 Evaluating Frontier AI Systems and the GSM-1K Benchmark Sarah asks what makes evaluating frontier models difficult. Alex explains benchmark contamination and introduces Scale's GSM-1K study, which demonstrated that multiple models overfitted existing benchmarks.26:00–32:00 · Guest disagreement 0/10 Application Layer Dynamics and Scale's Product Announcements Alex discusses the application layer hype cycle around GPT-4 and Scale's enterprise/defense tooling. Elad references Meta research proving smaller, higher-quality datasets produce superior models.32:01–34:47 · Guest disagreement 1/10 Multimodality, Lab Convergence, and the Need for Smarter Models The hosts and guest discuss the rapid convergence between Google Astra and OpenAI GPT-4o. Sarah offers dual explanations of natural technical convergence versus competitive intelligence, and Alex laments the lack of fundamentally smarter reasoning models.34:47–38:39 · Guest disagreement 3/10 The Path to AGI and Organizational Agility Alex presents his contrarian view that AGI progress mirrors curing individual cancers rather than a single vaccine. Sarah pushes back on video world models as a general foundation, which Alex rejects as narrative lacking scientific evidence.2:32–7:45 · The hosts pushing back 0/10 The Evolution of Scale AI: From Autonomous Vehicles to Generative AI Sarah sets the context recalling Scale's early days and asks Alex to describe the evolution of the business from autonomous vehicles to defense and LLM RLHF. Alex provides an expansive historical overview in an agreeable, narrative style.7:46–13:08 · The hosts pushing back 0/10 Navigating Data Scarcity and Frontier Data Production Elad asks about emerging enterprise and sovereign AI demands and draws a parallel to Google's early mission to digitize books. Alex details the transition from scraping easy web data to producing high-signal frontier expert data.13:08–15:33 · The hosts pushing back 0/10 Enterprise Proprietary Data and Hybrid Human-AI Synthetic Pipelines Sarah shares investor observations that incumbent enterprise software lacks the structured data required for AI models. Alex agrees, highlighting proprietary enterprise volume and introducing hybrid human-AI synthetic data pipelines.15:33–20:00 · The hosts pushing back 4/10 Centaur Intelligence and the Durability of Human Expertise The hosts press Alex on the limits of centaur intelligence. Elad cites Google's Med-PaLM 2 outperforming general practitioners to ask when human expertise becomes obsolete, while Sarah probes architectural breakthroughs in planning.20:01–22:13 · The hosts pushing back 0/10 Scale AI's $1B Fundraise and Strategic Industry Role Sarah congratulates Alex on Scale's $1B fundraise at a $14B valuation and inquires about strategic investors. Alex outlines Scale's role as neutral infrastructure across the entire AI stack.22:14–26:00 · The hosts pushing back 0/10 Evaluating Frontier AI Systems and the GSM-1K Benchmark Sarah asks what makes evaluating frontier models difficult. Alex explains benchmark contamination and introduces Scale's GSM-1K study, which demonstrated that multiple models overfitted existing benchmarks.26:00–32:00 · The hosts pushing back 1/10 Application Layer Dynamics and Scale's Product Announcements Alex discusses the application layer hype cycle around GPT-4 and Scale's enterprise/defense tooling. Elad references Meta research proving smaller, higher-quality datasets produce superior models.32:01–34:47 · The hosts pushing back 1/10 Multimodality, Lab Convergence, and the Need for Smarter Models The hosts and guest discuss the rapid convergence between Google Astra and OpenAI GPT-4o. Sarah offers dual explanations of natural technical convergence versus competitive intelligence, and Alex laments the lack of fundamentally smarter reasoning models.34:47–38:39 · The hosts pushing back 3/10 The Path to AGI and Organizational Agility Alex presents his contrarian view that AGI progress mirrors curing individual cancers rather than a single vaccine. Sarah pushes back on video world models as a general foundation, which Alex rejects as narrative lacking scientific evidence.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 43.5% · guest 56.5%0:00 · the hosts 43.5% · guest 56.5%3:00 · the hosts 3.9% · guest 96.1%3:00 · the hosts 3.9% · guest 96.1%6:00 · the hosts 13.6% · guest 86.4%6:00 · the hosts 13.6% · guest 86.4%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 48.9% · guest 51.1%12:00 · the hosts 48.9% · guest 51.1%15:00 · the hosts 25.5% · guest 74.5%15:00 · the hosts 25.5% · guest 74.5%18:00 · the hosts 27.9% · guest 72.1%18:00 · the hosts 27.9% · guest 72.1%21:00 · the hosts 10% · guest 90%21:00 · the hosts 10% · guest 90%24:00 · the hosts 5.4% · guest 94.6%24:00 · the hosts 5.4% · guest 94.6%27:00 · the hosts 26.9% · guest 73.1%27:00 · the hosts 26.9% · guest 73.1%30:00 · the hosts 9.7% · guest 90.3%30:00 · the hosts 9.7% · guest 90.3%33:00 · the hosts 27.3% · guest 72.7%33:00 · the hosts 27.3% · guest 72.7%36:00 · the hosts 30.7% · guest 69.3%36:00 · the hosts 30.7% · guest 69.3%39:00 · the hosts 0% · guest 0%39:00 · the hosts 0% · guest 0%
Sharpest disagreement ▶ 37:08 Dismissing video world models as unsubstantiated hype

Alex flatly rejects Sarah's suggestion regarding video world models, calling it a 'great narrative' devoid of strong scientific evidence.

Hardest push from the hosts ▶ 19:29 Sarah challenges fundamental limits on model planning

Sarah directly questions Alex's definitive claim that models can never match biological long-horizon reasoning by asking if architectural breakthroughs in planning solve it.

Biggest teaching moment ▶ 24:20 Alex explains benchmark contamination via GSM-1K

Alex educates listeners on how leading frontier models overfit academic benchmarks, citing Scale's held-out GSM-1K math evaluation results.

The host holds their own ▶ 17:43 Elad cites Med-PaLM 2 physician outperformance

Elad demonstrates domain expertise by citing Google's Med-PaLM 2 study showing AI outperforming general physicians, challenging Alex's assertion on the perpetual need for human expertise.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
The Evolution of Scale AI: From Autonomous Vehicles to Generative AI 3200 Sarah sets the context recalling Scale's early days and asks Alex to describe the evolution of the business from autonomous vehicles to defense and LLM RLHF. Alex provides an expansive historical overview in an agreeable, narrative style.
Navigating Data Scarcity and Frontier Data Production 4300 Elad asks about emerging enterprise and sovereign AI demands and draws a parallel to Google's early mission to digitize books. Alex details the transition from scraping easy web data to producing high-signal frontier expert data.
Enterprise Proprietary Data and Hybrid Human-AI Synthetic Pipelines 4200 Sarah shares investor observations that incumbent enterprise software lacks the structured data required for AI models. Alex agrees, highlighting proprietary enterprise volume and introducing hybrid human-AI synthetic data pipelines.
Centaur Intelligence and the Durability of Human Expertise 6334 The hosts press Alex on the limits of centaur intelligence. Elad cites Google's Med-PaLM 2 outperforming general practitioners to ask when human expertise becomes obsolete, while Sarah probes architectural breakthroughs in planning.
Scale AI's $1B Fundraise and Strategic Industry Role 2100 Sarah congratulates Alex on Scale's $1B fundraise at a $14B valuation and inquires about strategic investors. Alex outlines Scale's role as neutral infrastructure across the entire AI stack.
Evaluating Frontier AI Systems and the GSM-1K Benchmark 4410 Sarah asks what makes evaluating frontier models difficult. Alex explains benchmark contamination and introduces Scale's GSM-1K study, which demonstrated that multiple models overfitted existing benchmarks.
Application Layer Dynamics and Scale's Product Announcements 5201 Alex discusses the application layer hype cycle around GPT-4 and Scale's enterprise/defense tooling. Elad references Meta research proving smaller, higher-quality datasets produce superior models.
Multimodality, Lab Convergence, and the Need for Smarter Models 5111 The hosts and guest discuss the rapid convergence between Google Astra and OpenAI GPT-4o. Sarah offers dual explanations of natural technical convergence versus competitive intelligence, and Alex laments the lack of fundamentally smarter reasoning models.
The Path to AGI and Organizational Agility 5333 Alex presents his contrarian view that AGI progress mirrors curing individual cancers rather than a single vaccine. Sarah pushes back on video world models as a general foundation, which Alex rejects as narrative lacking scientific evidence.

Statements from this episode (20)

Assertion Partly supported
Alexandr Wang: Scale AI built the first sensor-fusion data engine for AVs.
“And so we built, The very first data engine that supported sensor fused data. So support a combination of two D data plus three D data. So lidars plus cameras that were built on onto the vehicles. And then that very quickly became an industry standard across a…”
Alexandr Wang May 22, 2024 ▶ 4:26
Assertion Supported
Alexandr Wang: Scale built data infrastructure for DoD's first AI program.
“So we built the very first data engines to support government data. This would support mostly geospatial and satellite and over, other overhead imagery. This ended up fueling the first AI program of record for the US DoD.”
Alexandr Wang May 22, 2024 ▶ 5:28
Assertion Supported
Alexandr Wang: Scale partnered with OpenAI on the first GPT-2 RLHF experiments.
“So we partnered with OpenAI at that time to do the very first experiments on RLHF on top of GPT-II.”
Alexandr Wang May 22, 2024 ▶ 5:57
Assertion Not checkable as stated
Alexandr Wang: Scale AI fuels basically every major LLM developer.
“Today, you know, fast forward to today our data foundry fuels basically every major large language model in the industry. Work with OpenAI, Meta, Microsoft, many of the other players.”
Alexandr Wang May 22, 2024 ▶ 6:59
Insight
Alexandr Wang: Data abundance is the fundamental bottleneck for post-GPT-4 models.
“The key to the scaling of these large language models and the, you know, these language models in general is the ability to scale data. And I think that one of the fundamental bottlenecks to, you know, what's, what's in the way of us getting from GPT-IV to GPT…”
Alexandr Wang May 22, 2024 ▶ 9:09
Assertion Not checkable as stated
Alexandr Wang: The AI industry has exhausted all easy internet training data.
“And we've sort of, as a community, we have, we've had easy data, which is all the data on the internet and we've kind of exhausted all the easy data, and now it's about, you know, forward data production that has high supervisory signal that is basically very …”
Alexandr Wang May 22, 2024 ▶ 9:34
Assertion Not checkable as stated
Alexandr Wang: Frontier AI models can no longer learn much from Reddit.
“It's not Any more the case that these models can learn that much more from, you know, various comments on Reddit or whatnot. They need, ah, they need truly frontier data.”
Alexandr Wang May 22, 2024 ▶ 10:02
Insight
Sarah Guo: Incumbent software vendors lack the data needed to train models.
“You go to many of the incumbent software and services vendors, and despite having done this task, you know, for years, they have not actually captured the information you'd want to teach a model.”
Sarah Guo May 22, 2024 ▶ 13:53
Assertion Supported
Alexandr Wang: JPMorgan has 150 petabytes of data; GPT-4 used under one.
“JP Morgan's proprietary data set is a 150 petabytes of data. GPT-IV is trained on less than one petabyte. Of data.”
Alexandr Wang May 22, 2024 ▶ 14:29
Disclosure
Alexandr Wang: Scale AI's core thesis is hybrid human-AI synthetic data.
“And our perspective is that the critical thing is, is what we call hybrid human AI synthetic data. So how can you build hybrid human AI systems such that AI are doing a lot of the heavy lifting, but human experts and people, you know, the basically best, brigh…”
Alexandr Wang May 22, 2024 ▶ 14:58
Prediction Not checkable as stated
Alexandr Wang: Human-AI teams will outperform standalone models for a long time.
“The question is, is a human plus a model together going to be able to produce better output than a model alone? And I think that'll be the case for A very, very, very long time. That, that humans are still, you know, human intelligence is complementary to mach…”
Alexandr Wang May 22, 2024 ▶ 15:51
Assertion Supported
Alexandr Wang: Academic AI benchmarks are contaminated and models are overfit.
“The academic benchmarks that are what the industry used to measure the performance of these algorithms are fraught with issues. Many of the models are overfit on these benchmarks. They're sort of in the training data sets of these models.”
Alexandr Wang May 22, 2024 ▶ 24:05
Assertion Supported
Alexandr Wang: Held-out benchmarks reveal several AI models underperform their reported scores.
“So we, one of the things we did is we published DSM-I-K, which was a held out eval. So we basically produced a new evaluation of the math capabilities of models. That there's no way it would ever exist in the training data set to really see how much of the, ho…”
Alexandr Wang May 22, 2024 ▶ 24:20
Opinion
Alexandr Wang: GPT-4 was too early a model to sustain application hype.
“GPT-IV, I think, as a model, was a little early of a technology for us to have this entire hype wave around, and I think we, you know, the community very quickly discovered all the limitations of GPT-IV... It was probably a few generations too early of a model…”
Alexandr Wang May 22, 2024 ▶ 26:19
Insight
Alexandr Wang: High-quality frontier data is 10,000x more valuable than enterprise data.
“One of the things that every, you know, all the model developers understand well, but the enterprises understand super well is that you know, not all data is created equal and high quality data or frontier data is, is, can be, you know, 10,000 times more valua…”
Alexandr Wang May 22, 2024 ▶ 28:40
Disclosure
Alexandr Wang: Scale AI will launch recurring held-out LLM benchmark leaderboards.
“So one is that we're going to launch these private held out evaluations and have leaderboards associated with these evals for the leading LLMs in the ecosystem. And we're going to rerun this contest periodically. So every few months we're going to do a new set…”
Alexandr Wang May 22, 2024 ▶ 30:02
Insight
Alexandr Wang: Multimodality faces a scarcity of quality data for personal agents.
“So multimodality as an entire space is one where for the same reasons that we've like exhaust a lot of the internet data, there's a lot of scarcity for good multimodal data that can empower these personal agents and these personal Use cases.”
Alexandr Wang May 22, 2024 ▶ 32:30
Opinion
Alexandr Wang: Multimodality is a lateral move; the industry needs smarter models.
“So, you know, we got multi-modality capability. That's exciting. It's more of a lateral expansion of the models, and the industry needs smarter models. We need GPT-V, or we need Gemini-II, or whatever that, those models are going to be. and so to me it was, y…”
Alexandr Wang May 22, 2024 ▶ 34:16
Prediction Not checkable as stated
Alexandr Wang: Achieving AGI will take multiple decades of solving individual problems.
“My biggest belief here is that the path to AGI is is one that looks a lot more like curing cancer than developing a vaccine. And what I mean by that is I think that the path to build AGI is going to be in, in, you know, you're going to have to solve a bunch of…”
Alexandr Wang May 22, 2024 ▶ 34:53
Assertion Partly supported
Alexandr Wang: Cross-modality training produces no positive transfer between video and text.
“I think the main thing, fundamentally, is I think there's very limited generality that we get from these models and even for multimodality, for example my understanding there's no positive transfer from learning in one modality to other modalities. So like tra…”
Alexandr Wang May 22, 2024 ▶ 36:37
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.