Mar 3, 2024 · 32m · another-podcast

Google Gemini and AI bias

Benedict Evans · 23m spoken Toni Cowan-Brown · 6m spoken
0:00 / 0:00

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Hosts Benedict Evans and Toni Karen Brown analyze the technical mechanics of artificial intelligence bias, using Google Gemini's overcorrection controversies to explore dataset curation, proxy variables, and the systemic challenges of AI governance.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 78.6% of the talking time here. How this is scored →

The hosts as informed peer 8.1 Guest teaching 1.7 Guest disagreement 1.3 The hosts pushing back 2.9
05100:0010:0020:0030:000:57–3:23 · The hosts as informed peer 8/10 Google Gemini Diversity Flaw and Data Patterns Benedict Evans opens with extensive historical context on AI bias and explains how Gemini's diverse Nazi outputs stemmed from pattern-matching failures rather than simple tech-bro demographic bias. Toni acts largely as an agreeable conversational prompt.3:24–7:46 · The hosts as informed peer 9/10 Algorithmic Proxies in Recruitment and Insurance Pricing Evans delivers a masterclass on algorithmic bias, drawing on concrete industry examples including Google hiring algorithms, auto insurance risk categorization, and the skin cancer ruler detector issue. Toni interjects brief validating observations without contesting Evans's framing.7:46–11:57 · The hosts as informed peer 8/10 Representation Discrepancies and Dataset Auditing Challenges Toni brings in a specific Bloomberg study regarding severe gender representation skew in Stable Diffusion and questions dataset regulation feasibility. Evans counters by explaining the immense technical and practical absurdity of auditing billions of scraped images.11:57–19:49 · The hosts as informed peer 9/10 The Myth of Raw Data and Engineering Diversity Evans rejects the simplistic notion that workforce diversity alone fixes dataset bugs, citing the UK Post Office Horizon scandal and Simenon scale analogies. Toni concedes her own earlier misconceptions around smartwatch sensor testing versus complex dataset bias.19:50–24:45 · The hosts as informed peer 8/10 Open Source Models, Misinformation, and Editorial Representation Evans breaks down how open-source models shift governance from creation to distribution and unpacks the editorial trade-offs between realistic historical representations and synthetic parity. Toni contributes reflections on political campaigning data and corporate Barbie tropes.24:46–30:36 · The hosts as informed peer 8/10 Machine Cognition Misconceptions and Historical Computing Errors Evans addresses misconceptions about machine cognition, drawing parallels between LLM pattern synthesis and historical database errors in conventional computing. Toni references Evans's own 2019 essay to ask if modern AI harms are fundamentally novel.30:36–32:29 · The hosts as informed peer 7/10 Regulatory Scramble and Final Reflections on AI Bias Evans summarizes the regulatory lag problem, emphasizing that generative AI generates new artifacts rather than merely classifying past data. Toni agrees and concludes that reducing AI bias strictly to builder diversity is incomplete.0:57–3:23 · Guest teaching 1/10 Google Gemini Diversity Flaw and Data Patterns Benedict Evans opens with extensive historical context on AI bias and explains how Gemini's diverse Nazi outputs stemmed from pattern-matching failures rather than simple tech-bro demographic bias. Toni acts largely as an agreeable conversational prompt.3:24–7:46 · Guest teaching 1/10 Algorithmic Proxies in Recruitment and Insurance Pricing Evans delivers a masterclass on algorithmic bias, drawing on concrete industry examples including Google hiring algorithms, auto insurance risk categorization, and the skin cancer ruler detector issue. Toni interjects brief validating observations without contesting Evans's framing.7:46–11:57 · Guest teaching 3/10 Representation Discrepancies and Dataset Auditing Challenges Toni brings in a specific Bloomberg study regarding severe gender representation skew in Stable Diffusion and questions dataset regulation feasibility. Evans counters by explaining the immense technical and practical absurdity of auditing billions of scraped images.11:57–19:49 · Guest teaching 2/10 The Myth of Raw Data and Engineering Diversity Evans rejects the simplistic notion that workforce diversity alone fixes dataset bugs, citing the UK Post Office Horizon scandal and Simenon scale analogies. Toni concedes her own earlier misconceptions around smartwatch sensor testing versus complex dataset bias.19:50–24:45 · Guest teaching 2/10 Open Source Models, Misinformation, and Editorial Representation Evans breaks down how open-source models shift governance from creation to distribution and unpacks the editorial trade-offs between realistic historical representations and synthetic parity. Toni contributes reflections on political campaigning data and corporate Barbie tropes.24:46–30:36 · Guest teaching 2/10 Machine Cognition Misconceptions and Historical Computing Errors Evans addresses misconceptions about machine cognition, drawing parallels between LLM pattern synthesis and historical database errors in conventional computing. Toni references Evans's own 2019 essay to ask if modern AI harms are fundamentally novel.30:36–32:29 · Guest teaching 1/10 Regulatory Scramble and Final Reflections on AI Bias Evans summarizes the regulatory lag problem, emphasizing that generative AI generates new artifacts rather than merely classifying past data. Toni agrees and concludes that reducing AI bias strictly to builder diversity is incomplete.0:57–3:23 · Guest disagreement 1/10 Google Gemini Diversity Flaw and Data Patterns Benedict Evans opens with extensive historical context on AI bias and explains how Gemini's diverse Nazi outputs stemmed from pattern-matching failures rather than simple tech-bro demographic bias. Toni acts largely as an agreeable conversational prompt.3:24–7:46 · Guest disagreement 1/10 Algorithmic Proxies in Recruitment and Insurance Pricing Evans delivers a masterclass on algorithmic bias, drawing on concrete industry examples including Google hiring algorithms, auto insurance risk categorization, and the skin cancer ruler detector issue. Toni interjects brief validating observations without contesting Evans's framing.7:46–11:57 · Guest disagreement 2/10 Representation Discrepancies and Dataset Auditing Challenges Toni brings in a specific Bloomberg study regarding severe gender representation skew in Stable Diffusion and questions dataset regulation feasibility. Evans counters by explaining the immense technical and practical absurdity of auditing billions of scraped images.11:57–19:49 · Guest disagreement 2/10 The Myth of Raw Data and Engineering Diversity Evans rejects the simplistic notion that workforce diversity alone fixes dataset bugs, citing the UK Post Office Horizon scandal and Simenon scale analogies. Toni concedes her own earlier misconceptions around smartwatch sensor testing versus complex dataset bias.19:50–24:45 · Guest disagreement 1/10 Open Source Models, Misinformation, and Editorial Representation Evans breaks down how open-source models shift governance from creation to distribution and unpacks the editorial trade-offs between realistic historical representations and synthetic parity. Toni contributes reflections on political campaigning data and corporate Barbie tropes.24:46–30:36 · Guest disagreement 1/10 Machine Cognition Misconceptions and Historical Computing Errors Evans addresses misconceptions about machine cognition, drawing parallels between LLM pattern synthesis and historical database errors in conventional computing. Toni references Evans's own 2019 essay to ask if modern AI harms are fundamentally novel.30:36–32:29 · Guest disagreement 1/10 Regulatory Scramble and Final Reflections on AI Bias Evans summarizes the regulatory lag problem, emphasizing that generative AI generates new artifacts rather than merely classifying past data. Toni agrees and concludes that reducing AI bias strictly to builder diversity is incomplete.0:57–3:23 · The hosts pushing back 2/10 Google Gemini Diversity Flaw and Data Patterns Benedict Evans opens with extensive historical context on AI bias and explains how Gemini's diverse Nazi outputs stemmed from pattern-matching failures rather than simple tech-bro demographic bias. Toni acts largely as an agreeable conversational prompt.3:24–7:46 · The hosts pushing back 2/10 Algorithmic Proxies in Recruitment and Insurance Pricing Evans delivers a masterclass on algorithmic bias, drawing on concrete industry examples including Google hiring algorithms, auto insurance risk categorization, and the skin cancer ruler detector issue. Toni interjects brief validating observations without contesting Evans's framing.7:46–11:57 · The hosts pushing back 4/10 Representation Discrepancies and Dataset Auditing Challenges Toni brings in a specific Bloomberg study regarding severe gender representation skew in Stable Diffusion and questions dataset regulation feasibility. Evans counters by explaining the immense technical and practical absurdity of auditing billions of scraped images.11:57–19:49 · The hosts pushing back 5/10 The Myth of Raw Data and Engineering Diversity Evans rejects the simplistic notion that workforce diversity alone fixes dataset bugs, citing the UK Post Office Horizon scandal and Simenon scale analogies. Toni concedes her own earlier misconceptions around smartwatch sensor testing versus complex dataset bias.19:50–24:45 · The hosts pushing back 3/10 Open Source Models, Misinformation, and Editorial Representation Evans breaks down how open-source models shift governance from creation to distribution and unpacks the editorial trade-offs between realistic historical representations and synthetic parity. Toni contributes reflections on political campaigning data and corporate Barbie tropes.24:46–30:36 · The hosts pushing back 3/10 Machine Cognition Misconceptions and Historical Computing Errors Evans addresses misconceptions about machine cognition, drawing parallels between LLM pattern synthesis and historical database errors in conventional computing. Toni references Evans's own 2019 essay to ask if modern AI harms are fundamentally novel.30:36–32:29 · The hosts pushing back 1/10 Regulatory Scramble and Final Reflections on AI Bias Evans summarizes the regulatory lag problem, emphasizing that generative AI generates new artifacts rather than merely classifying past data. Toni agrees and concludes that reducing AI bias strictly to builder diversity is incomplete.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 80.2% · guest 19.8%0:00 · the hosts 80.2% · guest 19.8%3:00 · the hosts 99.1% · guest 0.9%3:00 · the hosts 99.1% · guest 0.9%6:00 · the hosts 75.3% · guest 24.7%6:00 · the hosts 75.3% · guest 24.7%9:00 · the hosts 78.8% · guest 21.2%9:00 · the hosts 78.8% · guest 21.2%12:00 · the hosts 79.8% · guest 20.2%12:00 · the hosts 79.8% · guest 20.2%15:00 · the hosts 76.2% · guest 23.8%15:00 · the hosts 76.2% · guest 23.8%18:00 · the hosts 75.6% · guest 24.4%18:00 · the hosts 75.6% · guest 24.4%21:00 · the hosts 71.4% · guest 28.6%21:00 · the hosts 71.4% · guest 28.6%24:00 · the hosts 81.8% · guest 18.2%24:00 · the hosts 81.8% · guest 18.2%27:00 · the hosts 78.4% · guest 21.6%27:00 · the hosts 78.4% · guest 21.6%30:00 · the hosts 65.6% · guest 34.4%30:00 · the hosts 65.6% · guest 34.4%
Sharpest disagreement ▶ 16:21 Toni asserts emotional severity of non-consensual deepfakes

Toni pushes back on Evans comparing generative fake nudes to 30 years of Photoshop manipulation, arguing that personal deepfakes carry unique visceral harm.

Hardest push from the hosts ▶ 13:08 Evans rejects workforce diversity as the primary fix

Evans firmly dismisses the common assertion that hiring more women and minorities prevents technical bias like ruler correlation in training sets.

Biggest teaching moment ▶ 7:46 Toni introduces Bloomberg Stable Diffusion bias metrics

Toni introduces hard empirical data showing female doctor representation in Stable Diffusion was drastically lower than real-world employment statistics.

The host holds their own ▶ 4:50 Evans details the dermatologist ruler detection flaw

Evans articulates the nuanced failure modes of machine learning by explaining how models inadvertently classify dermatologists' rulers rather than skin blemishes.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Google Gemini Diversity Flaw and Data Patterns 8112 Benedict Evans opens with extensive historical context on AI bias and explains how Gemini's diverse Nazi outputs stemmed from pattern-matching failures rather than simple tech-bro demographic bias. Toni acts largely as an agreeable conversational prompt.
Algorithmic Proxies in Recruitment and Insurance Pricing 9112 Evans delivers a masterclass on algorithmic bias, drawing on concrete industry examples including Google hiring algorithms, auto insurance risk categorization, and the skin cancer ruler detector issue. Toni interjects brief validating observations without contesting Evans's framing.
Representation Discrepancies and Dataset Auditing Challenges 8324 Toni brings in a specific Bloomberg study regarding severe gender representation skew in Stable Diffusion and questions dataset regulation feasibility. Evans counters by explaining the immense technical and practical absurdity of auditing billions of scraped images.
The Myth of Raw Data and Engineering Diversity 9225 Evans rejects the simplistic notion that workforce diversity alone fixes dataset bugs, citing the UK Post Office Horizon scandal and Simenon scale analogies. Toni concedes her own earlier misconceptions around smartwatch sensor testing versus complex dataset bias.
Open Source Models, Misinformation, and Editorial Representation 8213 Evans breaks down how open-source models shift governance from creation to distribution and unpacks the editorial trade-offs between realistic historical representations and synthetic parity. Toni contributes reflections on political campaigning data and corporate Barbie tropes.
Machine Cognition Misconceptions and Historical Computing Errors 8213 Evans addresses misconceptions about machine cognition, drawing parallels between LLM pattern synthesis and historical database errors in conventional computing. Toni references Evans's own 2019 essay to ask if modern AI harms are fundamentally novel.
Regulatory Scramble and Final Reflections on AI Bias 7111 Evans summarizes the regulatory lag problem, emphasizing that generative AI generates new artifacts rather than merely classifying past data. Toni agrees and concludes that reducing AI bias strictly to builder diversity is incomplete.

Statements from this episode (18)

Assertion Supported
Evans: Google Gemini generated racially diverse depictions of Nazi stormtroopers
“And so it turned out that if you typed in, give me pictures of Nazi stormtroopers, you got A white person in Nazi uniform, and a black person, and a Chinese person, and an Asian woman, and so on.”
Benedict Evans Mar 3, 2024 ▶ 1:28
Assertion Supported
Evans: Regulators in some countries ban gender-based car insurance pricing
“There are some countries where you're not allowed to give women better insurance or at car insurance prices. Even though women are much less likely to drive fast and crash their car, but you're not allowed to use gender.”
Benedict Evans Mar 3, 2024 ▶ 4:31
Insight
Evans: AI risk models can recreate protected demographic categories without knowing them
“It might be giving, it might have internally created a category and given those categories of people, lower risk and lower, lower insurance prices, and those might be women, but it might, it wouldn't have a concept of women and you might not know if it's doing…”
Benedict Evans Mar 3, 2024 ▶ 4:48
Assertion Partly supported
Evans: DeepMind AI Detected Retinal Sex Differences Unknown to Doctors
“I think deep mind was working on a project that was looking at retinas and this system spontaneously started recognizing that there was a difference between male men's retinas and women's retinas, which is not a difference that doctors actually knew about.”
Benedict Evans Mar 3, 2024 ▶ 6:12
Insight
Evans: AI Models Inherently Extract Both Desirable Signals and Unwanted Patterns
“These systems are basically pattern matching pattern recognition systems at a very, very simplistic level. There might be stuff in the pattern that you want it to use. There might be stuff in the pattern. There might be pattern that you didn't know was there. …”
Benedict Evans Mar 3, 2024 ▶ 6:40
Assertion Supported
Brown: Stable Diffusion generated 7% female doctors, compared to 39% in reality
“And with the doctors one, the reality was that 39% of doctors in America are women, but in the stable diffusion results, it was only seven percent.”
Toni Cowan-Brown Mar 3, 2024 ▶ 8:08
Opinion
Evans: AI regulation should target outcomes rather than technical dataset construction
“You know, that's an, you should not be regulating at that level of specificity, you should be regulating for outcomes, I think.”
Benedict Evans Mar 3, 2024 ▶ 9:27
Insight
Evans: Mandating training data disclosure will not make AI models auditable
“So it's sort of easy to say, oh, well, you have to publish what your training data is, but that's not Going to produce something that's not going to produce something that's easy for you to audit.”
Benedict Evans Mar 3, 2024 ▶ 11:28
Insight
Evans: Engineering Diversity Alone Cannot Fix AI Dataset Bias
“If you think that you're going to solve this by hiring more women or more black people, you're kind of not understanding what's going on.”
Benedict Evans Mar 3, 2024 ▶ 13:34
Insight
Evans: Software bugs ruining lives represent institutional failures, not just engineering bugs
“And if there are bugs in the computer that are ruining people's lives, then that's an institutional problem, not an engineering problem. I mean, it's an engineering problem too, but the backstop has to be a willingness to consider that the computer might be wr…”
Benedict Evans Mar 3, 2024 ▶ 14:37
Assertion Supported
Evans: Social platforms previously blocked COVID lab leak and mask efficacy posts
“I mean, the kind of the case study here is when social networks decided to block anybody Suggesting that masks weren't very useful during COVID or block anybody suggesting that maybe COVID had leaked from a Chinese biology research lab. And I think now, I don'…”
Benedict Evans Mar 3, 2024 ▶ 15:08
Insight
Evans: Massive scale transforms generative AI tools qualitatively compared to Photoshop
“There's a difference when you do, something can be the same theoretically, just at greatest in principle, but a much greater scale and practical and accessible and everything else. And in fact, making something much greater scale and much more accessible, but …”
Benedict Evans Mar 3, 2024 ▶ 18:27
Prediction Not checkable as stated
Evans: Tech giants cannot stop open-source AI image generation
“One of them is that the models are open source and the compute will get better and the models will get more efficient. And so if you want to do the scenario or just outlined, it's not like Facebook or Google or Adobe can stop you.”
Benedict Evans Mar 3, 2024 ▶ 19:55
Insight
Evans: AI image regulation can only intervene on distribution, not creation
“What they can do, the intervention point is in the distribution. It's not in the creation. Can't stop people making those images. You can think about how they go viral, how you distribute them, what the source is.”
Benedict Evans Mar 3, 2024 ▶ 20:12
Insight
Evans: Most visual misinformation relies on fake captions, not AI images
“You don't need to make a fake image. The fake is in the fake is in the caption. There's not a shortage of, there's not a fake is in the. Yeah. What is it that I'm saying? What is it that I'm seeing here? So, you know, if you want to make a picture of an airstr…”
Benedict Evans Mar 3, 2024 ▶ 20:44
Insight
Evans: AI prompt moderation will repeat social media governance battles
“You re-run all of the arguments we had around content moderation and social media of what should people be allowed to say. You either say more or less anything, which is Elon Musk's idea, or you end up with a long list of criteria, which is where Facebook has …”
Benedict Evans Mar 3, 2024 ▶ 23:56
Insight
Evans: Generative AI Infers Patterns Rather Than Storing Training Data
“What is generative AI trying to do? Which is, is not actually trying to reproduce any individual thing in the training data. So, you know, when we did image recognition models, 10 years ago, you give it a billion pictures of cats. The purpose is not that it ca…”
Benedict Evans Mar 3, 2024 ▶ 25:34
Opinion
Evans: AI Doomers Wrongly Treat Harmful Computer Decisions as Novel
“Screwing up with a computer is not new. This is my point. Ruining people's lives. This is kind of the kind of mistake that the AI dealers make is like, we haven't had computers before. We've never before have we had computers that were making decisions and cou…”
Benedict Evans Mar 3, 2024 ▶ 30:03
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.