Oct 30, 2024 · 30m · a16z

Can We Detect a Deepfake?

Vijay Balasubramaniyan · 19m spoken Martin Casado · 7m spoken Joe Biden (AI Deepfake Audio) · 6s spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of The a16z Podcast, host Martin Casado and Pindrop CEO Vijay Balasubramaniyan explore the sharp rise in AI voice deepfakes, examining technical detection methods, defense economics, and practical regulatory frameworks for combating synthetic media fraud.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The host as informed peer 4.5 Guest teaching 4.7 Guest disagreement 0.9 The host pushing back 1.9
05100:0010:0020:0030:000:40–4:15 · The host as informed peer 3/10 Legal Disclaimer and Disclosure Statement Martin prompts Vijay to explain if deepfakes are simply riding the generative AI wave. Vijay educates him on Generative Adversarial Networks (GANs) operating as a reverse Turing test and how voice cloning time dropped from 20 hours to 15 seconds.4:15–7:35 · The host as informed peer 3/10 The Grandparent Scam, Ori Ori Sagi, and LLM Integration Martin shares a personal phone scam story involving his parents and asks if AI fakes are fundamentally different. Vijay explains the historical Japanese Ori Ori Sagi scam and how combining voice clones with LLMs automates convincing fraud at scale.7:35–10:00 · The host as informed peer 2/10 Real-World Threats: Political Misinformation and the Biden Robocall Martin asks for real-world deepfake anecdotes beyond personal scams. Vijay details how his team identified the fake Biden robocall during the New Hampshire primary and traced the exact AI app used to generate it.10:00–13:09 · The host as informed peer 3/10 The Surge of Deepfakes in Enterprise and Financial Sector Fraud Martin asks whether deepfakes are fundamentally detectable or if defense is hopeless. Vijay reveals a 1400% surge in enterprise deepfakes and explains that analyzing 8,000 audio samples per second enables a 99% detection accuracy rate.13:09–15:42 · The host as informed peer 6/10 Watermarking Flaws and Disinformation in News Media Martin challenges policy proposals relying on watermarking, arguing attackers will not comply. Vijay validates Martin's point with data showing the Biden call watermark degraded to 2% over telephony networks, and surprises Martin by revealing 90% of war footage sent to news outlets is fake.15:42–19:16 · The host as informed peer 4/10 Detecting Zero-Day Deepfakes and Adversarial GAN Stress-Testing Martin interrupts to clarify how GAN stress-testing works against detection models. Vijay explains how deepfakes leave residual fake prints and how forcing generator models to evade detection causes perceptual speech quality to diverge.19:16–21:37 · The host as informed peer 6/10 The Economics of AI Security: Detection vs. Generation Costs Martin frames the security problem around marginal costs of generation versus detection. Vijay reveals detection models are 100 times smaller and two orders of magnitude cheaper than generation models, handing defenders an economic advantage.21:37–24:31 · The host as informed peer 5/10 Physical Vocal Constraints and Defender Advantages Martin hypothesizes that detection is cheaper because generation must satisfy multiple human perceptual requirements. Vijay confirms this by highlighting physical human vocal tract mechanics that generic synthetic models fail to replicate perfectly.24:31–28:56 · The host as informed peer 7/10 Policy and Regulatory Guidance for Lawmakers Martin leads a regulatory discussion on policy guidance for lawmakers balancing security and innovation. He proposes a framework separating illegal fraud, unwanted nuisance content, and legitimate creative uses based on spam regulation analogs.28:56–29:51 · The host as informed peer 6/10 Episode Summary and Key Takeaways Martin synthesizes the interview into core takeaways regarding deepfake prevalence, inherent detectability, and platform accountability. Vijay warmly validates Martin's summary as a beautiful synopsis.0:40–4:15 · Guest teaching 5/10 Legal Disclaimer and Disclosure Statement Martin prompts Vijay to explain if deepfakes are simply riding the generative AI wave. Vijay educates him on Generative Adversarial Networks (GANs) operating as a reverse Turing test and how voice cloning time dropped from 20 hours to 15 seconds.4:15–7:35 · Guest teaching 6/10 The Grandparent Scam, Ori Ori Sagi, and LLM Integration Martin shares a personal phone scam story involving his parents and asks if AI fakes are fundamentally different. Vijay explains the historical Japanese Ori Ori Sagi scam and how combining voice clones with LLMs automates convincing fraud at scale.7:35–10:00 · Guest teaching 5/10 Real-World Threats: Political Misinformation and the Biden Robocall Martin asks for real-world deepfake anecdotes beyond personal scams. Vijay details how his team identified the fake Biden robocall during the New Hampshire primary and traced the exact AI app used to generate it.10:00–13:09 · Guest teaching 6/10 The Surge of Deepfakes in Enterprise and Financial Sector Fraud Martin asks whether deepfakes are fundamentally detectable or if defense is hopeless. Vijay reveals a 1400% surge in enterprise deepfakes and explains that analyzing 8,000 audio samples per second enables a 99% detection accuracy rate.13:09–15:42 · Guest teaching 6/10 Watermarking Flaws and Disinformation in News Media Martin challenges policy proposals relying on watermarking, arguing attackers will not comply. Vijay validates Martin's point with data showing the Biden call watermark degraded to 2% over telephony networks, and surprises Martin by revealing 90% of war footage sent to news outlets is fake.15:42–19:16 · Guest teaching 6/10 Detecting Zero-Day Deepfakes and Adversarial GAN Stress-Testing Martin interrupts to clarify how GAN stress-testing works against detection models. Vijay explains how deepfakes leave residual fake prints and how forcing generator models to evade detection causes perceptual speech quality to diverge.19:16–21:37 · Guest teaching 4/10 The Economics of AI Security: Detection vs. Generation Costs Martin frames the security problem around marginal costs of generation versus detection. Vijay reveals detection models are 100 times smaller and two orders of magnitude cheaper than generation models, handing defenders an economic advantage.21:37–24:31 · Guest teaching 5/10 Physical Vocal Constraints and Defender Advantages Martin hypothesizes that detection is cheaper because generation must satisfy multiple human perceptual requirements. Vijay confirms this by highlighting physical human vocal tract mechanics that generic synthetic models fail to replicate perfectly.24:31–28:56 · Guest teaching 3/10 Policy and Regulatory Guidance for Lawmakers Martin leads a regulatory discussion on policy guidance for lawmakers balancing security and innovation. He proposes a framework separating illegal fraud, unwanted nuisance content, and legitimate creative uses based on spam regulation analogs.28:56–29:51 · Guest teaching 1/10 Episode Summary and Key Takeaways Martin synthesizes the interview into core takeaways regarding deepfake prevalence, inherent detectability, and platform accountability. Vijay warmly validates Martin's summary as a beautiful synopsis.0:40–4:15 · Guest disagreement 1/10 Legal Disclaimer and Disclosure Statement Martin prompts Vijay to explain if deepfakes are simply riding the generative AI wave. Vijay educates him on Generative Adversarial Networks (GANs) operating as a reverse Turing test and how voice cloning time dropped from 20 hours to 15 seconds.4:15–7:35 · Guest disagreement 1/10 The Grandparent Scam, Ori Ori Sagi, and LLM Integration Martin shares a personal phone scam story involving his parents and asks if AI fakes are fundamentally different. Vijay explains the historical Japanese Ori Ori Sagi scam and how combining voice clones with LLMs automates convincing fraud at scale.7:35–10:00 · Guest disagreement 1/10 Real-World Threats: Political Misinformation and the Biden Robocall Martin asks for real-world deepfake anecdotes beyond personal scams. Vijay details how his team identified the fake Biden robocall during the New Hampshire primary and traced the exact AI app used to generate it.10:00–13:09 · Guest disagreement 1/10 The Surge of Deepfakes in Enterprise and Financial Sector Fraud Martin asks whether deepfakes are fundamentally detectable or if defense is hopeless. Vijay reveals a 1400% surge in enterprise deepfakes and explains that analyzing 8,000 audio samples per second enables a 99% detection accuracy rate.13:09–15:42 · Guest disagreement 1/10 Watermarking Flaws and Disinformation in News Media Martin challenges policy proposals relying on watermarking, arguing attackers will not comply. Vijay validates Martin's point with data showing the Biden call watermark degraded to 2% over telephony networks, and surprises Martin by revealing 90% of war footage sent to news outlets is fake.15:42–19:16 · Guest disagreement 1/10 Detecting Zero-Day Deepfakes and Adversarial GAN Stress-Testing Martin interrupts to clarify how GAN stress-testing works against detection models. Vijay explains how deepfakes leave residual fake prints and how forcing generator models to evade detection causes perceptual speech quality to diverge.19:16–21:37 · Guest disagreement 1/10 The Economics of AI Security: Detection vs. Generation Costs Martin frames the security problem around marginal costs of generation versus detection. Vijay reveals detection models are 100 times smaller and two orders of magnitude cheaper than generation models, handing defenders an economic advantage.21:37–24:31 · Guest disagreement 1/10 Physical Vocal Constraints and Defender Advantages Martin hypothesizes that detection is cheaper because generation must satisfy multiple human perceptual requirements. Vijay confirms this by highlighting physical human vocal tract mechanics that generic synthetic models fail to replicate perfectly.24:31–28:56 · Guest disagreement 1/10 Policy and Regulatory Guidance for Lawmakers Martin leads a regulatory discussion on policy guidance for lawmakers balancing security and innovation. He proposes a framework separating illegal fraud, unwanted nuisance content, and legitimate creative uses based on spam regulation analogs.28:56–29:51 · Guest disagreement 0/10 Episode Summary and Key Takeaways Martin synthesizes the interview into core takeaways regarding deepfake prevalence, inherent detectability, and platform accountability. Vijay warmly validates Martin's summary as a beautiful synopsis.0:40–4:15 · The host pushing back 1/10 Legal Disclaimer and Disclosure Statement Martin prompts Vijay to explain if deepfakes are simply riding the generative AI wave. Vijay educates him on Generative Adversarial Networks (GANs) operating as a reverse Turing test and how voice cloning time dropped from 20 hours to 15 seconds.4:15–7:35 · The host pushing back 2/10 The Grandparent Scam, Ori Ori Sagi, and LLM Integration Martin shares a personal phone scam story involving his parents and asks if AI fakes are fundamentally different. Vijay explains the historical Japanese Ori Ori Sagi scam and how combining voice clones with LLMs automates convincing fraud at scale.7:35–10:00 · The host pushing back 1/10 Real-World Threats: Political Misinformation and the Biden Robocall Martin asks for real-world deepfake anecdotes beyond personal scams. Vijay details how his team identified the fake Biden robocall during the New Hampshire primary and traced the exact AI app used to generate it.10:00–13:09 · The host pushing back 2/10 The Surge of Deepfakes in Enterprise and Financial Sector Fraud Martin asks whether deepfakes are fundamentally detectable or if defense is hopeless. Vijay reveals a 1400% surge in enterprise deepfakes and explains that analyzing 8,000 audio samples per second enables a 99% detection accuracy rate.13:09–15:42 · The host pushing back 4/10 Watermarking Flaws and Disinformation in News Media Martin challenges policy proposals relying on watermarking, arguing attackers will not comply. Vijay validates Martin's point with data showing the Biden call watermark degraded to 2% over telephony networks, and surprises Martin by revealing 90% of war footage sent to news outlets is fake.15:42–19:16 · The host pushing back 3/10 Detecting Zero-Day Deepfakes and Adversarial GAN Stress-Testing Martin interrupts to clarify how GAN stress-testing works against detection models. Vijay explains how deepfakes leave residual fake prints and how forcing generator models to evade detection causes perceptual speech quality to diverge.19:16–21:37 · The host pushing back 2/10 The Economics of AI Security: Detection vs. Generation Costs Martin frames the security problem around marginal costs of generation versus detection. Vijay reveals detection models are 100 times smaller and two orders of magnitude cheaper than generation models, handing defenders an economic advantage.21:37–24:31 · The host pushing back 2/10 Physical Vocal Constraints and Defender Advantages Martin hypothesizes that detection is cheaper because generation must satisfy multiple human perceptual requirements. Vijay confirms this by highlighting physical human vocal tract mechanics that generic synthetic models fail to replicate perfectly.24:31–28:56 · The host pushing back 2/10 Policy and Regulatory Guidance for Lawmakers Martin leads a regulatory discussion on policy guidance for lawmakers balancing security and innovation. He proposes a framework separating illegal fraud, unwanted nuisance content, and legitimate creative uses based on spam regulation analogs.28:56–29:51 · The host pushing back 0/10 Episode Summary and Key Takeaways Martin synthesizes the interview into core takeaways regarding deepfake prevalence, inherent detectability, and platform accountability. Vijay warmly validates Martin's summary as a beautiful synopsis.

speaking balance: gold is the host, purple is the guest (3 minute bins)

0:00 · the host 0% · guest 100%0:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%18:00 · the host 0% · guest 100%18:00 · the host 0% · guest 100%21:00 · the host 0% · guest 100%21:00 · the host 0% · guest 100%24:00 · the host 0% · guest 100%24:00 · the host 0% · guest 100%27:00 · the host 0% · guest 100%27:00 · the host 0% · guest 100%30:00 · the host 0% · guest 100%30:00 · the host 0% · guest 100%
Sharpest disagreement ▶ 5:46 Rejecting 'more of the same' framing via financial harm and scale

Vijay counters Martin's suggestion that generative AI fakes are just more of the same non-urgent threats by detailing half-billion dollar losses in Japan and exponential scaling.

Hardest push from the host ▶ 13:10 Challenging watermarking as a primary defense

Martin explicitly rejects policy proposals advocating watermarking and cryptography, arguing that expecting malicious actors to comply with voluntary technical standards is fundamentally flawed.

Biggest teaching moment ▶ 14:55 Revealing 90% of war media received by news outlets is fake

Vijay stuns Martin by revealing that 90% of audio and video content received by news organizations from the Israel-Hamas war is fake or manipulated.

The host holds their own ▶ 27:43 Structuring regulatory tiers for synthetic media

Martin demonstrates deep policy expertise by synthesizing CAN-SPAM precedents into a legal framework that separates illegal fraud, unwanted opt-out content, and permitted speech.

the scores for every segment, with the reasoning behind each
ChapterTopicThe host as informed peerGuest teachingGuest disagreementThe host pushing backWhy
Legal Disclaimer and Disclosure Statement 3511 Martin prompts Vijay to explain if deepfakes are simply riding the generative AI wave. Vijay educates him on Generative Adversarial Networks (GANs) operating as a reverse Turing test and how voice cloning time dropped from 20 hours to 15 seconds.
The Grandparent Scam, Ori Ori Sagi, and LLM Integration 3612 Martin shares a personal phone scam story involving his parents and asks if AI fakes are fundamentally different. Vijay explains the historical Japanese Ori Ori Sagi scam and how combining voice clones with LLMs automates convincing fraud at scale.
Real-World Threats: Political Misinformation and the Biden Robocall 2511 Martin asks for real-world deepfake anecdotes beyond personal scams. Vijay details how his team identified the fake Biden robocall during the New Hampshire primary and traced the exact AI app used to generate it.
The Surge of Deepfakes in Enterprise and Financial Sector Fraud 3612 Martin asks whether deepfakes are fundamentally detectable or if defense is hopeless. Vijay reveals a 1400% surge in enterprise deepfakes and explains that analyzing 8,000 audio samples per second enables a 99% detection accuracy rate.
Watermarking Flaws and Disinformation in News Media 6614 Martin challenges policy proposals relying on watermarking, arguing attackers will not comply. Vijay validates Martin's point with data showing the Biden call watermark degraded to 2% over telephony networks, and surprises Martin by revealing 90% of war footage sent to news outlets is fake.
Detecting Zero-Day Deepfakes and Adversarial GAN Stress-Testing 4613 Martin interrupts to clarify how GAN stress-testing works against detection models. Vijay explains how deepfakes leave residual fake prints and how forcing generator models to evade detection causes perceptual speech quality to diverge.
The Economics of AI Security: Detection vs. Generation Costs 6412 Martin frames the security problem around marginal costs of generation versus detection. Vijay reveals detection models are 100 times smaller and two orders of magnitude cheaper than generation models, handing defenders an economic advantage.
Physical Vocal Constraints and Defender Advantages 5512 Martin hypothesizes that detection is cheaper because generation must satisfy multiple human perceptual requirements. Vijay confirms this by highlighting physical human vocal tract mechanics that generic synthetic models fail to replicate perfectly.
Policy and Regulatory Guidance for Lawmakers 7312 Martin leads a regulatory discussion on policy guidance for lawmakers balancing security and innovation. He proposes a framework separating illegal fraud, unwanted nuisance content, and legitimate creative uses based on spam regulation analogs.
Episode Summary and Key Takeaways 6100 Martin synthesizes the interview into core takeaways regarding deepfake prevalence, inherent detectability, and platform accountability. Vijay warmly validates Martin's summary as a beautiful synopsis.

Statements from this episode (17)

Assertion Supported
Balasubramaniyan: The 2019 Pelosi slurring video was a cheap fake
“All they did was slow down the audio, and you know, it wasn't a deep fake, it was actually a cheap fake, right?”
Vijay Balasubramaniyan Oct 30, 2024 ▶ 1:10
Assertion Supported
Balasubramaniyan: Voice cloning tools jumped from 120 to 350 in early 2024
“At the end of last year, there were 120 tools with which you can clone someone's voice. And by March of this year, it's become 350.”
Vijay Balasubramaniyan Oct 30, 2024 ▶ 2:58
Assertion Not checkable as stated
Balasubramaniyan: High-quality voice deepfakes require only 15 seconds of audio
“All it requires for me to clone your voice, Martine now, requires about Three to five seconds of your audio, and if I want a really high quality deep fake, it requires about 15 seconds of audio.”
Vijay Balasubramaniyan Oct 30, 2024 ▶ 3:23
Assertion Supported
Balasubramaniyan: Japanese 'Ori Ori Sagi' scam cost citizens $500M
“At that point in time, it had started costing Japan close to half a billion dollars in people losing their life savings To the scam.”
Vijay Balasubramaniyan Oct 30, 2024 ▶ 6:06
Insight
Balasubramaniyan: LLM hallucinations make models ideal tools for fraudsters
“Like in LLMs, hallucination is a problem. So the fact that you're making shit up is a bad idea. But if you have to make shit up to convince someone, LLM is great and perfect.”
Vijay Balasubramaniyan Oct 30, 2024 ▶ 7:12
Assertion Supported
Pindrop traced the fake Biden robocall to shut down the perpetrator
“So not only did we come in and say this was a deep fake, but we identified The, we have something called source tracing, which tells us which AI application was used to create this deep fake. So we identified the deep fake and then we worked with that AI appli…”
Vijay Balasubramaniyan Oct 30, 2024 ▶ 9:02
Assertion Not checkable as stated
Balasubramaniyan: Deepfake attacks grew from monthly to daily per customer
“In 2023, we were seeing essentially one deep fake a month in some customer, right? So it was just one deep fake a month, and some customer would face it. It wasn't a widespread problem. But this year, we've now seen one deep fake per customer per day.”
Vijay Balasubramaniyan Oct 30, 2024 ▶ 10:09
Assertion Partly supported
Balasubramaniyan: Deepfakes surged 1,400% in the first half of 2024
“There has been a 1400% increase in the amount of deep fakes we've seen this year in the first six months compared to all of last year.”
Vijay Balasubramaniyan Oct 30, 2024 ▶ 10:41
Assertion Not checkable as stated
Balasubramaniyan: Pindrop detects voice deepfakes with 99% accuracy, 1% false positives
“They're completely detectable. Right now we're detecting them with 99% detection rate with a one percent false positive rate.”
Vijay Balasubramaniyan Oct 30, 2024 ▶ 11:14
Opinion
Casado: Policy reliance on AI watermarking is ineffective against bad actors
“Using things like watermarking or cryptography, you know, which has always seemed kind of a strange idea to me. Cause you're kind of asking criminals to comply by something”
Martin Casado Oct 30, 2024 ▶ 13:10
Assertion Supported
Balasubramaniyan: Biden robocall watermark test yielded only 2% recovery
“Like the President Biden robocall that I referenced before, when it finally showed up, the system that actually generated it had a watermark in it. But when they tested it against that watermark, They only were able to extract two percent.”
Vijay Balasubramaniyan Oct 30, 2024 ▶ 13:53
Insight
Balasubramaniyan: Audio watermarks get stripped away by telecom transmission
“When you take that audio, play it across air, play it across telephony channels, the bits and bytes, they get stripped away. And so once they get stripped away, an audio is a very sparse channel. So even if you add it over and over again, it's not possible to …”
Vijay Balasubramaniyan Oct 30, 2024 ▶ 14:25
Insight
Deepfake generators face diverging objectives between realism and evasion
“If a deepfake system has to serve two masters, that is one, I need to make the speech legible and sound as much like Martine, and two, I need to deceive a deepfake detection system. Those two objective functions start diverging.”
Vijay Balasubramaniyan Oct 30, 2024 ▶ 17:33
Disclosure
Pindrop proved the LeBron James Paris Olympics audio was a deepfake
“We were called into one of these deep fakes where LeBron James apparently was saying bad things about the coach during the Paris Olympics. It wasn't LeBron James. It was a deep fake. We actually provided his organ his management team the necessary detail so th…”
Vijay Balasubramaniyan Oct 30, 2024 ▶ 18:08
Assertion Not checkable as stated
Balasubramaniyan: AI deepfake detection is 100x cheaper than generation
“It's way cheaper to detect deepfix, right? Because if you think about it, the closest example is Apple released its model, you know, that could run on device. And even that model, which is a small model in order to do lots of things like voice to text and thin…”
Vijay Balasubramaniyan Oct 30, 2024 ▶ 19:44
Insight
Balasubramaniyan: AI legislation is unenforceable without detection capabilities
“And so I think what was really good about both of those cases is they got really specific on one. What can the technology detect? Because if the technology can't detect it, you can't litigate.”
Vijay Balasubramaniyan Oct 30, 2024 ▶ 26:40
Opinion
Balasubramaniyan: Media platforms should be held accountable for labeling AI content
“And the only other thing that I'll say is right now, because we consume things through a lot of platforms, Platforms should be held accountable at some level to, you know, clearly demarcating what is real and what is not, right? Because otherwise it's going to…”
Vijay Balasubramaniyan Oct 30, 2024 ▶ 28:13
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.