Training Data

topic on 14 shows · 25 statements across 25 episodes

Acquired More or Less Another Podcast Latent Space Lenny's Podcast No Priors the Official SaaStr Podcast A Product Market Fit Show Sourcery the MAD Podcast the a16z Podcast Big Technology All-In 20VC

25 statements about Training Data, every show

a16z Disclosure
Chi: Vals AI committed to never sell training data to labs
“At VALS, one very early decision we made was the decision to never sell training data to labs. It's often a place that we're pushed. When we start working with a new lab to actually source and sell for them a bunch of training data.”
Ryan Chi Sep 9, 2026 ▶ 8:58 Inside the Race to Measure Frontier Intelligence
20VC Insight
Arora: Enterprise AI Success Depends on Fast Training Data Creation
“I think the enterprise's ability to absorb this or digest this or perhaps leverage this to their advantage depends on their ability to create training data as fast as they can. And I think not enough people are focused on training data.”
Nikesh Arora Aug 5, 2026 ▶ 1:05:12 Leo Aschenbrenner's Situational Awareness Blows Up | Moonshot AI Raises $3.5B at $35B
Antani: AI Models Are Disposable; Data and Harnesses Are Durable
“At the end of the day, the models in AI are disposable. The weights are going to change. The models are going to come out. That doesn't matter. What's durable in AI is the harness and the training data.”
Snehal Antani Jul 29, 2026 ▶ 14:33 How AI's Top New Models Transform Cybersecurity (And Where They Don't) — With Snehal Antani
AI model builders demand perpetual data licenses to enable continuous model retraining
“One of the things we've seen is for high quality data, the model builders don't want to part with it. It's not the, we train once, we don't need the data ever again. It's we train once and if it's good, we want it forever because if we retrain our models from …”
Bobby Samuels Apr 13, 2026 ▶ 27:23 He raised a $10M seed with no revenue—then grew 30x to $30M in year two. | Bobby Samuels, Founder... · PMF Show
LATENT SPACE Assertion Not checkable as stated
Most Protein Design Validation Uses Targets With Training Data Overlap
“One of the things that, you know, we found, ah, with the field was that a lot of the validation, especially outside of the validation that was done on specific problems, was done on targets that have a lot of, you know, known interactions in, in the training d…”
Gabriele Corso Feb 12, 2026 ▶ 1:19:49 🔬Generating Molecules, Not Just Models
SAASTR Prediction Not checkable as stated
Lacour: Internal GTM AI Models Will Degrade as Company Data Ages
“At one point, this data will get still. So at one point, the AI models will get worse. So we're going to run into this question like, hey, how do we keep everything fresh?”
Philippe Lacour Jan 7, 2026 ▶ 18:24 The Step-By-Step Playbook for Building AI-Powered GTM Teams with Personio's CRO
MAD Assertion Not checkable as stated
Bourgeau: AI development is not running out of training data
“The other part of your question are we running out of data? I don't think so, so there's more.”
Sebastien Bourgeau Dec 18, 2025 ▶ 34:15 ”We’re Ahead of Where I Thought We’d Be” — Gemini 3 & the Future of AI
SOURCERY Insight
Agrawal: Frontier AI Training Data Is Saturated, Shifting the War to Talent
“The ingredients of call it being on the frontier are compute. CapEx, we spoke about it. Data which is sort of access accessible to everybody. We are saturated on, on, on the training data that exists. Talent. This is where the war is now.”
Apoorv Agrawal Aug 27, 2025 ▶ 42:31 Understanding OpenAI’s $500B Valuation · Sourcery with Molly O'Shea
BIG TECHNOLOGY Assertion Supported
New England Journal of Medicine cases avoid AI training data contamination
“Each week they put out a brand new case, which has never even been digitized, so there's no question that it's not in the training data.”
Mustafa Suleyman Jul 4, 2025 ▶ 10:08 Microsoft AI CEO Mustafa Suleyman: Our AI Doctor Outperforms Human Diagnosticians
LENNY'S PODCAST Assertion Supported
Schulhoff: Prompts formatted like common training data perform best
“It actually comes empirically from studies that have shown that formats of questions that show up most commonly in the training data are the best formats of questions to actually use when you're prompting it.”
Sander Schulhoff Jun 19, 2025 ▶ 15:11 AI prompt engineering in 2025: What works and what doesn’t | Sander Schulhoff
20VC Assertion Not checkable as stated
Scott: Claims about proprietary data's value for AI lack scientific backing
“Most of those assertions that people make are like most of the assertions that people make are just unfounded in any kind of science. Like no, no one's gotta, and like the measurements we do have show that there's a pretty big disconnect around what some peopl…”
Kevin Scott Mar 31, 2025 ▶ 11:50 Kevin Scott, CTO @ Microsoft: An Evaluation of Deepseek and How We Underestimate the Chinese · 20VC with Harry Stebbings
Bryk: Search engines do not need PhD-level training data
“With search, you're asking, like, simple questions about billions of things, like, is this a startup? Did this person write a blog post about search? You know, those are actually simple questions. You don't need, like, PhD level training data.”
Will Bryk Jan 10, 2025 ▶ 49:13 Beating Google at Search with Neural PageRank and $5M of H200s — with Will Bryk of Exa.ai
Customer inference workloads rarely align with foundation model training distributions
“The data distribution in their inference workload doesn't align with the data distribution in the training data for the model, right? It's a given, actually. If you think about this, because researchers have to guesstimate what is important, what's not importa…”
Lin Qiao Nov 25, 2024 ▶ 15:29 Why Compound AI + Open Source will beat Closed AI — with Lin Qiao, CEO of Fireworks AI
20VC Insight
Gomez: AI models are hyper-sensitive to single bad data examples
“Like a single bad example, right, amongst, like, billions. Like, it's so sensitive. Like, it is a bit surreal how sensitive The models are to their data. Everyone underrates it.”
Aidan Gomez Aug 19, 2024 ▶ 56:23 Aidan Gomez: What No One Understands About Foundation Models | E1191 · 20VC with Harry Stebbings
ALL-IN Prediction Not checkable as stated
Altman: AI will not become an arms race for training data
“I definitely don't think it'll be an arms race for data because when the models get smart enough at some point, it shouldn't be about more data, at least not for training.”
Sam Altman May 10, 2024 ▶ 12:32 Sam Altman: Getting Fired (and Re-Hired) by OpenAI, Agents, AI Copyright issues
ALL-IN Assertion Not checkable as stated
Friedberg: AI models do not store training data, but learn synthesized predictors
“The truth is that the models don't actually hold the data that they're trained on. They develop a bunch of synthesized predictors that they learn what the prediction could or should be from the data.”
David Friedberg Apr 12, 2024 ▶ 52:10 E174: Inflation stays hot, AI disclosure bill, Drone warfare, defense startups & more
MORE OR LESS Prediction Not checkable as stated
Sam Lessin: Generative AI content will be homogenizing, not original
“I think there's a very strong argument that it's very homogenizing. You know, yes, you'll get weirder and weirder porn, but it will all hit the same beats, right, that, like, have been trained on the underlying data.”
Sam Lessin Apr 5, 2024 ▶ 6:45 #41: Is AI Killing Media? · More or Less Podcast
Evans: Generative AI Infers Patterns Rather Than Storing Training Data
“What is generative AI trying to do? Which is, is not actually trying to reproduce any individual thing in the training data. So, you know, when we did image recognition models, 10 years ago, you give it a billion pictures of cats. The purpose is not that it ca…”
Benedict Evans Mar 3, 2024 ▶ 25:34 Google Gemini and AI bias
ALL-IN Prediction Not checkable as stated
Palihapitiya: Unique web and app data will yield new licensing revenue
“If you're an entrepreneur building a website or building an app that has really unique training data or really unique data, you'll be able to license and sell that. And that'll be an incremental revenue stream to everything you do in the near future.”
Chamath Palihapitiya Mar 1, 2024 ▶ 44:04 E168: Can Google save itself? Abolish HR, AI takes over Customer Support, Reddit IPO teardown
SAASTR Insight
Roberge: Top AI talent beats mediocre teams with trillions of data records
“So if you had to choose between having an A plus AI tech talent team with millions of records of data, and you were going against a mediocre AI tech talent team with trillions of data that the startup wins. And that starts to make sense with the diminishing re…”
Mark Roberge Nov 17, 2023 ▶ 20:44 Who Will Win the Go-To-Market AI Race? with Stage 2 Capital Co-founder Mark Roberge
NO PRIORS Insight
Kelly: Biological AI's main bottleneck is availability of training data
“The big limitation in bio is the availability of data to train these things, right? And so you have this tough situation where, like, everyone is doing these models of training on the same data, right?”
Jason Kelly Sep 28, 2023 ▶ 18:55 No Priors Ep. 34 | With Ginkgo Bioworks Co-Founder and CEO Jason Kelly
ACQUIRED Insight
Language structure and knowledge are embedded directly within raw text data
“The very structure of language and the way to interpret knowledge is actually embedded in the training data itself rather than requiring labeling.”
Ben Gilbert Sep 6, 2023 ▶ 29:02 Nvidia Part III: The Dawn of the AI Era (2022-2023) (Audio) · Acquired
NO PRIORS Assertion Not checkable as stated
Biewald: Yahoo Search Success Depended Entirely On Local Training Data Quality
“The model that I'm building is like the same for each country. It's the training data though is different. So some countries would take the training data collection process really seriously and they'd get a great model. And some would just like really half-ass…”
Lukas Biewald Aug 3, 2023 ▶ 6:09 No Priors Ep. 26 | With Weights & Biases CEO Lukas Biewald
NO PRIORS Insight
AI Data Needs Scale With the Square Root of Compute
“The data requirements tend to go up like with the square root of the amount of computation, because you're going to train a bigger model and then you're going to throw more data at it.”
Noam Shazeer Apr 25, 2023 ▶ 10:54 No Priors Ep. 12 | With Noam Shazeer
MAD Assertion Not checkable as stated
Changing deep learning target categories requires retraining models from scratch
“But if you want to change your columns, your categories, then you have to redo all your multiple choice tests and then retrain the system. And if you want to change the categories, you have to start over.”
Ben Vigoda May 22, 2018 ▶ 4:38 A New Approach to Machine Intelligence // Ben Vigoda, Gamalon (FirstMark's Data Driven)

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.