Aug 22, 2024 · 42m · no-priors

No Priors Ep. 77 | With Foundry CEO and Founder Jared Quincy Davis

Jared Quincy Davis · 32m spoken Sarah Guo · 4m spoken Elad Gil · 2m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Foundry founder and CEO Jared Quincy Davis joins Sarah Guo and Elad Gil on No Priors to discuss reimagining AI cloud infrastructure from first principles. He details how modern GPU hardware bottlenecks and loss of elasticity can be solved through intelligent orchestration, while exploring the industry's architectural transition toward Compound AI Systems.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 17.1% of the talking time here. How this is scored →

The hosts as informed peer 4.4 Guest teaching 4.8 Guest disagreement 0.7 The hosts pushing back 0.7
05100:0015:0030:000:35–2:47 · The hosts as informed peer 2/10 Inspirations from AlphaFold and ChatGPT: Democratizing Compute Leverage Sarah opens with a prompt about Foundry's genesis. Jared explains the asymmetrical compute advantage behind AlphaFold 2 and ChatGPT, reframing the David vs Goliath narrative around computational leverage.2:47–7:29 · The hosts as informed peer 5/10 Foundry's AI-Native Cloud Infrastructure and Economic Advantages Elad categorizes three different types of GPU cloud users to probe true utilization. Jared provides an in-depth breakdown of hardware failure rates and the necessity of reserving a 10 to 20 percent healing buffer in clusters.7:30–9:46 · The hosts as informed peer 5/10 The 'Large Regime' and Distributed Systems Networking Challenges Elad asks whether high failure rates stem from QC issues or architectural complexity. Jared introduces his definition of the 'large regime' where model weights exceed single-node memory, turning model execution into a distributed systems problem.9:47–16:09 · The hosts as informed peer 7/10 Historical Cloud Paradigms and the Loss of Elasticity in AI Jared asserts that early cloud computing had few initial believers and that current AI cloud is merely co-location. Elad pushes back based on his firsthand experience building startups in 2006, arguing startups instantly saw AWS as magic while enterprises hesitated.16:09–19:12 · The hosts as informed peer 7/10 Evolution of Infrastructure Abstractions and AI Market Immaturity Sarah demonstrates strong domain expertise by articulating the historical evolution from on-prem closet servers to colo, hosting, virtualization, and serverless, contrasting it with AI cloud's regression to rigid reservations.19:12–23:54 · The hosts as informed peer 2/10 Foundry's Spot Usability Engine and the Parking Lot Analogy Sarah asks about Foundry's latest product launch. Jared uses an extended parking lot analogy to explain spot usability engines and automated preemption management on GPU clusters.23:54–29:50 · The hosts as informed peer 6/10 Global Compute Distribution and the MARS Resiliency Suite Jared quizzes the hosts on peak Ethereum GPU capacity. Elad correctly estimates tens of millions of V100 equivalents and connects it to historical Bitcoin mining compute exceeding Google data centers.29:50–36:17 · The hosts as informed peer 5/10 The Transition from Monolithic Pre-Training to Compound AI Systems Sarah inquires about the strategic viability of alternatives to massive monolithic clusters. Jared explains the paradigm shift toward compound AI systems, horizontal scaling, and test-time compute exemplified by AlphaCode-2 and Llama 3.36:17–40:18 · The hosts as informed peer 5/10 Research on Compound AI Systems and Verifiable Task Bootstrapping Elad brings up Jared's recent research paper on compound AI system design. Jared explains verifiable task bootstrapping, showing how parallel candidate generation and verification achieved a 10x improvement in prime factorization.40:18–42:24 · The hosts as informed peer 4/10 Extending Compound Architectures to Open-Ended Tasks and Future Directions Sarah prompts Jared on applying compound architectures to open-ended tasks. Jared describes multi-model ensembling across frontier LLMs combined with heuristic verifiers and simulators.42:25–42:40 · The hosts as informed peer 0/10 Episode Conclusion and Channel Information Standard podcast outro with channel information and subscription links.0:35–2:47 · Guest teaching 4/10 Inspirations from AlphaFold and ChatGPT: Democratizing Compute Leverage Sarah opens with a prompt about Foundry's genesis. Jared explains the asymmetrical compute advantage behind AlphaFold 2 and ChatGPT, reframing the David vs Goliath narrative around computational leverage.2:47–7:29 · Guest teaching 6/10 Foundry's AI-Native Cloud Infrastructure and Economic Advantages Elad categorizes three different types of GPU cloud users to probe true utilization. Jared provides an in-depth breakdown of hardware failure rates and the necessity of reserving a 10 to 20 percent healing buffer in clusters.7:30–9:46 · Guest teaching 6/10 The 'Large Regime' and Distributed Systems Networking Challenges Elad asks whether high failure rates stem from QC issues or architectural complexity. Jared introduces his definition of the 'large regime' where model weights exceed single-node memory, turning model execution into a distributed systems problem.9:47–16:09 · Guest teaching 5/10 Historical Cloud Paradigms and the Loss of Elasticity in AI Jared asserts that early cloud computing had few initial believers and that current AI cloud is merely co-location. Elad pushes back based on his firsthand experience building startups in 2006, arguing startups instantly saw AWS as magic while enterprises hesitated.16:09–19:12 · Guest teaching 3/10 Evolution of Infrastructure Abstractions and AI Market Immaturity Sarah demonstrates strong domain expertise by articulating the historical evolution from on-prem closet servers to colo, hosting, virtualization, and serverless, contrasting it with AI cloud's regression to rigid reservations.19:12–23:54 · Guest teaching 6/10 Foundry's Spot Usability Engine and the Parking Lot Analogy Sarah asks about Foundry's latest product launch. Jared uses an extended parking lot analogy to explain spot usability engines and automated preemption management on GPU clusters.23:54–29:50 · Guest teaching 5/10 Global Compute Distribution and the MARS Resiliency Suite Jared quizzes the hosts on peak Ethereum GPU capacity. Elad correctly estimates tens of millions of V100 equivalents and connects it to historical Bitcoin mining compute exceeding Google data centers.29:50–36:17 · Guest teaching 7/10 The Transition from Monolithic Pre-Training to Compound AI Systems Sarah inquires about the strategic viability of alternatives to massive monolithic clusters. Jared explains the paradigm shift toward compound AI systems, horizontal scaling, and test-time compute exemplified by AlphaCode-2 and Llama 3.36:17–40:18 · Guest teaching 6/10 Research on Compound AI Systems and Verifiable Task Bootstrapping Elad brings up Jared's recent research paper on compound AI system design. Jared explains verifiable task bootstrapping, showing how parallel candidate generation and verification achieved a 10x improvement in prime factorization.40:18–42:24 · Guest teaching 5/10 Extending Compound Architectures to Open-Ended Tasks and Future Directions Sarah prompts Jared on applying compound architectures to open-ended tasks. Jared describes multi-model ensembling across frontier LLMs combined with heuristic verifiers and simulators.42:25–42:40 · Guest teaching 0/10 Episode Conclusion and Channel Information Standard podcast outro with channel information and subscription links.0:35–2:47 · Guest disagreement 1/10 Inspirations from AlphaFold and ChatGPT: Democratizing Compute Leverage Sarah opens with a prompt about Foundry's genesis. Jared explains the asymmetrical compute advantage behind AlphaFold 2 and ChatGPT, reframing the David vs Goliath narrative around computational leverage.2:47–7:29 · Guest disagreement 1/10 Foundry's AI-Native Cloud Infrastructure and Economic Advantages Elad categorizes three different types of GPU cloud users to probe true utilization. Jared provides an in-depth breakdown of hardware failure rates and the necessity of reserving a 10 to 20 percent healing buffer in clusters.7:30–9:46 · Guest disagreement 1/10 The 'Large Regime' and Distributed Systems Networking Challenges Elad asks whether high failure rates stem from QC issues or architectural complexity. Jared introduces his definition of the 'large regime' where model weights exceed single-node memory, turning model execution into a distributed systems problem.9:47–16:09 · Guest disagreement 2/10 Historical Cloud Paradigms and the Loss of Elasticity in AI Jared asserts that early cloud computing had few initial believers and that current AI cloud is merely co-location. Elad pushes back based on his firsthand experience building startups in 2006, arguing startups instantly saw AWS as magic while enterprises hesitated.16:09–19:12 · Guest disagreement 0/10 Evolution of Infrastructure Abstractions and AI Market Immaturity Sarah demonstrates strong domain expertise by articulating the historical evolution from on-prem closet servers to colo, hosting, virtualization, and serverless, contrasting it with AI cloud's regression to rigid reservations.19:12–23:54 · Guest disagreement 1/10 Foundry's Spot Usability Engine and the Parking Lot Analogy Sarah asks about Foundry's latest product launch. Jared uses an extended parking lot analogy to explain spot usability engines and automated preemption management on GPU clusters.23:54–29:50 · Guest disagreement 1/10 Global Compute Distribution and the MARS Resiliency Suite Jared quizzes the hosts on peak Ethereum GPU capacity. Elad correctly estimates tens of millions of V100 equivalents and connects it to historical Bitcoin mining compute exceeding Google data centers.29:50–36:17 · Guest disagreement 1/10 The Transition from Monolithic Pre-Training to Compound AI Systems Sarah inquires about the strategic viability of alternatives to massive monolithic clusters. Jared explains the paradigm shift toward compound AI systems, horizontal scaling, and test-time compute exemplified by AlphaCode-2 and Llama 3.36:17–40:18 · Guest disagreement 0/10 Research on Compound AI Systems and Verifiable Task Bootstrapping Elad brings up Jared's recent research paper on compound AI system design. Jared explains verifiable task bootstrapping, showing how parallel candidate generation and verification achieved a 10x improvement in prime factorization.40:18–42:24 · Guest disagreement 0/10 Extending Compound Architectures to Open-Ended Tasks and Future Directions Sarah prompts Jared on applying compound architectures to open-ended tasks. Jared describes multi-model ensembling across frontier LLMs combined with heuristic verifiers and simulators.42:25–42:40 · Guest disagreement 0/10 Episode Conclusion and Channel Information Standard podcast outro with channel information and subscription links.0:35–2:47 · The hosts pushing back 0/10 Inspirations from AlphaFold and ChatGPT: Democratizing Compute Leverage Sarah opens with a prompt about Foundry's genesis. Jared explains the asymmetrical compute advantage behind AlphaFold 2 and ChatGPT, reframing the David vs Goliath narrative around computational leverage.2:47–7:29 · The hosts pushing back 1/10 Foundry's AI-Native Cloud Infrastructure and Economic Advantages Elad categorizes three different types of GPU cloud users to probe true utilization. Jared provides an in-depth breakdown of hardware failure rates and the necessity of reserving a 10 to 20 percent healing buffer in clusters.7:30–9:46 · The hosts pushing back 1/10 The 'Large Regime' and Distributed Systems Networking Challenges Elad asks whether high failure rates stem from QC issues or architectural complexity. Jared introduces his definition of the 'large regime' where model weights exceed single-node memory, turning model execution into a distributed systems problem.9:47–16:09 · The hosts pushing back 5/10 Historical Cloud Paradigms and the Loss of Elasticity in AI Jared asserts that early cloud computing had few initial believers and that current AI cloud is merely co-location. Elad pushes back based on his firsthand experience building startups in 2006, arguing startups instantly saw AWS as magic while enterprises hesitated.16:09–19:12 · The hosts pushing back 0/10 Evolution of Infrastructure Abstractions and AI Market Immaturity Sarah demonstrates strong domain expertise by articulating the historical evolution from on-prem closet servers to colo, hosting, virtualization, and serverless, contrasting it with AI cloud's regression to rigid reservations.19:12–23:54 · The hosts pushing back 0/10 Foundry's Spot Usability Engine and the Parking Lot Analogy Sarah asks about Foundry's latest product launch. Jared uses an extended parking lot analogy to explain spot usability engines and automated preemption management on GPU clusters.23:54–29:50 · The hosts pushing back 1/10 Global Compute Distribution and the MARS Resiliency Suite Jared quizzes the hosts on peak Ethereum GPU capacity. Elad correctly estimates tens of millions of V100 equivalents and connects it to historical Bitcoin mining compute exceeding Google data centers.29:50–36:17 · The hosts pushing back 0/10 The Transition from Monolithic Pre-Training to Compound AI Systems Sarah inquires about the strategic viability of alternatives to massive monolithic clusters. Jared explains the paradigm shift toward compound AI systems, horizontal scaling, and test-time compute exemplified by AlphaCode-2 and Llama 3.36:17–40:18 · The hosts pushing back 0/10 Research on Compound AI Systems and Verifiable Task Bootstrapping Elad brings up Jared's recent research paper on compound AI system design. Jared explains verifiable task bootstrapping, showing how parallel candidate generation and verification achieved a 10x improvement in prime factorization.40:18–42:24 · The hosts pushing back 0/10 Extending Compound Architectures to Open-Ended Tasks and Future Directions Sarah prompts Jared on applying compound architectures to open-ended tasks. Jared describes multi-model ensembling across frontier LLMs combined with heuristic verifiers and simulators.42:25–42:40 · The hosts pushing back 0/10 Episode Conclusion and Channel Information Standard podcast outro with channel information and subscription links.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 18.9% · guest 81.1%0:00 · the hosts 18.9% · guest 81.1%3:00 · the hosts 17% · guest 83%3:00 · the hosts 17% · guest 83%6:00 · the hosts 6.2% · guest 93.8%6:00 · the hosts 6.2% · guest 93.8%9:00 · the hosts 16.4% · guest 83.6%9:00 · the hosts 16.4% · guest 83.6%12:00 · the hosts 41.4% · guest 58.6%12:00 · the hosts 41.4% · guest 58.6%15:00 · the hosts 50.3% · guest 49.7%15:00 · the hosts 50.3% · guest 49.7%18:00 · the hosts 5% · guest 95%18:00 · the hosts 5% · guest 95%21:00 · the hosts 7.7% · guest 92.3%21:00 · the hosts 7.7% · guest 92.3%24:00 · the hosts 19.3% · guest 80.7%24:00 · the hosts 19.3% · guest 80.7%27:00 · the hosts 11.2% · guest 88.8%27:00 · the hosts 11.2% · guest 88.8%30:00 · the hosts 15.5% · guest 84.5%30:00 · the hosts 15.5% · guest 84.5%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%36:00 · the hosts 16.3% · guest 83.7%36:00 · the hosts 16.3% · guest 83.7%39:00 · the hosts 6% · guest 94%39:00 · the hosts 6% · guest 94%42:00 · the hosts 53.7% · guest 46.3%42:00 · the hosts 53.7% · guest 46.3%
Sharpest disagreement ▶ 10:13 Jared rejects the premise of modern AI cloud

Jared forcefully rejects the notion that modern GPU providers are true clouds, arguing they have devolved into simple co-location facilities.

Hardest push from the hosts ▶ 12:19 Elad challenges Jared on early cloud adoption skepticism

Elad politely but firmly counters Jared's claim that early cloud had no believers, citing his personal experience founding startups in 2006 that saw AWS as an immediate breakthrough.

Biggest teaching moment ▶ 4:27 Jared explains GPU cluster failure dynamics and healing buffers

Jared explains the physical realities of modern DGX systems, clarifying why failure cascades necessitate keeping 10 to 20 percent of GPUs idle as healing buffers.

The host holds their own ▶ 16:34 Sarah maps out the compute abstraction continuum

Sarah demonstrates deep technical authority by systematically deconstructing the compute evolution from on-prem closet servers to serverless, diagnosing exactly where AI infrastructure currently lags.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Inspirations from AlphaFold and ChatGPT: Democratizing Compute Leverage 2410 Sarah opens with a prompt about Foundry's genesis. Jared explains the asymmetrical compute advantage behind AlphaFold 2 and ChatGPT, reframing the David vs Goliath narrative around computational leverage.
Foundry's AI-Native Cloud Infrastructure and Economic Advantages 5611 Elad categorizes three different types of GPU cloud users to probe true utilization. Jared provides an in-depth breakdown of hardware failure rates and the necessity of reserving a 10 to 20 percent healing buffer in clusters.
The 'Large Regime' and Distributed Systems Networking Challenges 5611 Elad asks whether high failure rates stem from QC issues or architectural complexity. Jared introduces his definition of the 'large regime' where model weights exceed single-node memory, turning model execution into a distributed systems problem.
Historical Cloud Paradigms and the Loss of Elasticity in AI 7525 Jared asserts that early cloud computing had few initial believers and that current AI cloud is merely co-location. Elad pushes back based on his firsthand experience building startups in 2006, arguing startups instantly saw AWS as magic while enterprises hesitated.
Evolution of Infrastructure Abstractions and AI Market Immaturity 7300 Sarah demonstrates strong domain expertise by articulating the historical evolution from on-prem closet servers to colo, hosting, virtualization, and serverless, contrasting it with AI cloud's regression to rigid reservations.
Foundry's Spot Usability Engine and the Parking Lot Analogy 2610 Sarah asks about Foundry's latest product launch. Jared uses an extended parking lot analogy to explain spot usability engines and automated preemption management on GPU clusters.
Global Compute Distribution and the MARS Resiliency Suite 6511 Jared quizzes the hosts on peak Ethereum GPU capacity. Elad correctly estimates tens of millions of V100 equivalents and connects it to historical Bitcoin mining compute exceeding Google data centers.
The Transition from Monolithic Pre-Training to Compound AI Systems 5710 Sarah inquires about the strategic viability of alternatives to massive monolithic clusters. Jared explains the paradigm shift toward compound AI systems, horizontal scaling, and test-time compute exemplified by AlphaCode-2 and Llama 3.
Research on Compound AI Systems and Verifiable Task Bootstrapping 5600 Elad brings up Jared's recent research paper on compound AI system design. Jared explains verifiable task bootstrapping, showing how parallel candidate generation and verification achieved a 10x improvement in prime factorization.
Extending Compound Architectures to Open-Ended Tasks and Future Directions 4500 Sarah prompts Jared on applying compound architectures to open-ended tasks. Jared describes multi-model ensembling across frontier LLMs combined with heuristic verifiers and simulators.
Episode Conclusion and Channel Information 0000 Standard podcast outro with channel information and subscription links.

Statements from this episode (20)

Assertion Partly supported
Davis: DeepMind built AlphaFold 2 with a team of just 18 people
“I think that one of the things that was so remarkable to me about AlphaFold II is, initially it was a really small team, you know, three, and then later, 18 people or so, and they solved what was kind of a fifty-year grand challenge in biology, which is a pret…”
Jared Quincy Davis Aug 22, 2024 ▶ 0:45
Assertion Partly supported
Davis: OpenAI Had 400 Employees and $13B of Compute When Launching ChatGPT
“In OpenAI's case, you know, there were only 400 people, but had thirteen billion dollars worth of compute, you know, which is quite a bit of computational scale there.”
Jared Quincy Davis Aug 22, 2024 ▶ 1:29
Prediction Not checkable as stated
Davis: Democratizing AI Compute Will Increase AlphaFold-Level Breakthroughs 10x to 100x
“I think it would increase the frequency of events like AlphaFold II by 10 X, a hundred X, or maybe even more super linearly.”
Jared Quincy Davis Aug 22, 2024 ▶ 2:18
Assertion Not checkable as stated
Davis: AI teams hold 10% to 20% of GPUs idle as healing buffer
“And so one of the consequences of that is that it's very common now to hold aside 10 to 20% minimum of the GPUs that a team has as buffer, as healing buffer, in case of a failure so you can slot something else in to keep the training workload running, right?”
Jared Quincy Davis Aug 22, 2024 ▶ 5:01
Assertion Not checkable as stated
Davis: Large-scale GPU pre-training utilization is sub-80% and often below 50%
“And so even for a lot of the more sophisticated orgs running large pre-trainings at scale, The utilization sub-eighty percent, sometimes less than 50%, actually, depending on how bad of a batch they have and the frequency of failure in the cluster.”
Jared Quincy Davis Aug 22, 2024 ▶ 5:16
Assertion Supported
Davis: Probability of a GPU supercomputer running weeks without failure is zero
“And so because you have millions, perhaps, of individual components in this supercomputer, the probability that it will run for weeks on end, and this is basically a verbatim quote from Jensen's keynote is basically zero.”
Jared Quincy Davis Aug 22, 2024 ▶ 6:47
Opinion
Davis: NVIDIA's Mellanox buyout is among the best acquisitions in history
“And their acquisition of Mellanox was one of the better of all time, arguably, from a market cap creation perspective.”
Jared Quincy Davis Aug 22, 2024 ▶ 9:21
Opinion
Davis: Current AI cloud is basically co-location, not real cloud
“I think current AI cloud is not cloud in the originally intended sense by any means. [627] Jared Quincy Davis: so we should pull on that thread, but I'd say right now it's basically co-location. [631] Jared Quincy Davis: Yeah, it's basically co-location, right…”
Jared Quincy Davis Aug 22, 2024 ▶ 10:23
Assertion Not checkable as stated
Davis: AI cloud customers are forced into unwanted 3-year GPU contracts
“In AI cloud today, you're kind of forced to get really long-term reservations often three years for a fixed amount of capacity. [944] Jared Quincy Davis: No one really wants 64 GPUs for three years or a thousand GPUs for three years. [949] Jared Quincy Davis: …”
Jared Quincy Davis Aug 22, 2024 ▶ 15:37
Assertion Supported
Davis: AI compute market lacks futures and hedging mechanisms found in commodities
“The markets aren't mature enough that there's any analog. So what we have in other domains, like in commodities markets like wheat, oil, et cetera, where you can buy options and futures and hedge and sell back and things like that. It's kind of a pretty, Still…”
Jared Quincy Davis Aug 22, 2024 ▶ 18:23
Assertion Not checkable as stated
Davis: Public clouds own only basis points of global GPU capacity
“What percentage of the world's GPU petaflock capacity, or exaflock capacity, is kind of owned by the major public clouds. And I've asked this many people, and I typically have gotten guesses, you know, in the high tens of percents, and the only time I got a lo…”
Jared Quincy Davis Aug 22, 2024 ▶ 24:27
Assertion Supported
Davis: OpenAI trained GPT-3 on 10,000 V100s for 14.6 days
“GPT-III was trained on 10,000 V-one-hundred GPUs in an interconnected cluster in Azure for about 14.6 days.”
Jared Quincy Davis Aug 22, 2024 ▶ 25:18
Assertion Supported
Davis: Peak Ethereum mining equaled 10 to 20 million V100 GPUs
“The very peak, the tippy top of Ethereum. How many V-One hundred equivalents were there, given there were 10,000 for two weeks for GPT-III? ... It was about 10 to twenty million.”
Jared Quincy Davis Aug 22, 2024 ▶ 26:15
Assertion Partly supported
Davis: iPhone 15 Pro has more FP16 FLOPS than an NVIDIA V100
“Actually, an iPhone 15 pro now is actually stronger than a V-one hundred, as a funny example. It has about 35 teraflops in every 16, I believe where a V-one hundred is around 30”
Jared Quincy Davis Aug 22, 2024 ▶ 27:50
Assertion Not checkable as stated
Davis: NVIDIA H100 GPU utilization is 25% or lower
“By many measures, utilization of these, even H-one hundred systems are kind of state-of-the-art, the most viable, the most precious, et cetera, Is, in my case, it's 20%, 25% or lower, according to some you know, pretty high quality data I've seen from some gre…”
Jared Quincy Davis Aug 22, 2024 ▶ 28:13
Prediction Not checkable as stated
Davis: Future AI infrastructure will rely far less on massive GPU clusters
“I think this actually points the way towards like what the infrastructure future might look like. And I think it looks a lot less like everything requiring these big clusters.”
Jared Quincy Davis Aug 22, 2024 ▶ 32:22
Insight
Davis: AI workloads are shifting from large pre-training to batch inference
“I think people are getting more sophisticated at thinking about cost in a more of a life cycle way, and that's actually leading to the workload shifting from large pre-training more and more towards things like batch inference. Which is actually a really, real…”
Jared Quincy Davis Aug 22, 2024 ▶ 35:18
Assertion Supported
Davis: Compound AI approach increased prime factorization accuracy from 3.7% to 36.6%
“And so, you know, we did kind of some preliminary investigations here, and we were able to, in one case, a prime factorization, you know, kind of 10 x the performance, go from 3.7% to 36.6%. On prime factorization, which is pretty hard, kind of factor, you kno…”
Jared Quincy Davis Aug 22, 2024 ▶ 39:05
Assertion Supported
Davis: Compound AI approach yielded a 3% MMLU performance bump
“The MMLU performance bump was about three percent. And to put that in perspective, the gap between some of the previous best models is often less than one percent. Between, for example, Gemini, 1.5, and Lama 3.1, and things like that. So, actually, 2.8% or thr…”
Jared Quincy Davis Aug 22, 2024 ▶ 39:55
Prediction Not checkable as stated
Davis: AI systems will make millions of heterogeneous model calls per question
“I think that what we'll see people doing is kind of composing, this sounds funny, but massive networks, Where maybe each stage in the network will basically be maybe some best of A, best of K component with many, many calls to different language models, you kn…”
Jared Quincy Davis Aug 22, 2024 ▶ 40:39
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.