Jun 16, 2025 · 24m · a16z

AI is Revolutionizing Web Security - Bots, Agents, & Real-Time Defense

David Mytton · 16m spoken Joel de la Garza · 6m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of the a16z podcast, Arcjet CEO David Mytton and partner Joel de la Garza discuss why legacy network-level blocking of AI traffic fails for modern web applications. They explore granular application-context security, multi-layered fingerprinting, cryptographic proofs, and low-latency edge AI inference needed to manage automated agents effectively.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The host as informed peer 5.3 Guest teaching 4.0 Guest disagreement 0.9 The host pushing back 1.6
05100:0010:0020:000:30–3:23 · The host as informed peer 4/10 Title Card: AI Bots at the Edge Joel opens by contrasting legacy IP-blocking methods with modern bot nuances, demonstrating familiarity with network security tools. David explains why DDoS solutions handle network traffic while application-layer context is needed for modern AI bots. The tone is highly collaborative.3:23–6:42 · The host as informed peer 5/10 Modern AI Solutions and Application Context Joel asks whether vendor AI claims are just rebranded network telemetry, referencing historical standards like robots.txt. David educates on OpenAI's multi-bot setup and why voluntary standards fail without application-layer enforcement. Joel reinforces the business revenue risks of over-blocking.6:42–9:07 · The host as informed peer 4/10 Understanding OpenAI Crawlers and Real-Time Query Agents David systematically categorizes OpenAI's crawlers, from model training to real-time doc retrieval and search indexing. Joel validates the explanations with concrete user examples like querying JFK's birthday. The dynamic remains entirely conversational and informative.9:07–13:27 · The host as informed peer 6/10 Computer-Use Agents, Headless Browsers, and Action Control Joel shares technical experience handling terabits of traffic across 450k IPs and questions how fine-grained control works at scale. David walks through defense layers, including residential proxies, User-Agent verification, and TLS fingerprinting hashes like JA3 and JA4.13:27–16:20 · The host as informed peer 6/10 OSI Stack Identity and Header Fingerprinting Joel actively maps fingerprinting concepts back to the OSI stack and compares emerging cryptographic pass keys to legacy Kerberos implementations. David details HTTP header order hashing (JA4H) and Apple's Privacy Pass.16:20–19:20 · The host as informed peer 5/10 AI Agents as Primary Web Consumers and Future Trends Joel offers a personal thesis that AI agents will become the main consumers of internet content. David agrees and expands on agent.txt standards and future criminal bot detection models.19:20–23:42 · The host as informed peer 7/10 Proof of Humanness and Real-Time Edge Inference Joel cites the 35-year NIST identity working group and compares inference price drops to historic cloud S3 storage trends. David outlines ultra-fast edge inference models, while Joel closes with an insightful observation about ad-tech applications.0:30–3:23 · Guest teaching 3/10 Title Card: AI Bots at the Edge Joel opens by contrasting legacy IP-blocking methods with modern bot nuances, demonstrating familiarity with network security tools. David explains why DDoS solutions handle network traffic while application-layer context is needed for modern AI bots. The tone is highly collaborative.3:23–6:42 · Guest teaching 4/10 Modern AI Solutions and Application Context Joel asks whether vendor AI claims are just rebranded network telemetry, referencing historical standards like robots.txt. David educates on OpenAI's multi-bot setup and why voluntary standards fail without application-layer enforcement. Joel reinforces the business revenue risks of over-blocking.6:42–9:07 · Guest teaching 5/10 Understanding OpenAI Crawlers and Real-Time Query Agents David systematically categorizes OpenAI's crawlers, from model training to real-time doc retrieval and search indexing. Joel validates the explanations with concrete user examples like querying JFK's birthday. The dynamic remains entirely conversational and informative.9:07–13:27 · Guest teaching 5/10 Computer-Use Agents, Headless Browsers, and Action Control Joel shares technical experience handling terabits of traffic across 450k IPs and questions how fine-grained control works at scale. David walks through defense layers, including residential proxies, User-Agent verification, and TLS fingerprinting hashes like JA3 and JA4.13:27–16:20 · Guest teaching 4/10 OSI Stack Identity and Header Fingerprinting Joel actively maps fingerprinting concepts back to the OSI stack and compares emerging cryptographic pass keys to legacy Kerberos implementations. David details HTTP header order hashing (JA4H) and Apple's Privacy Pass.16:20–19:20 · Guest teaching 3/10 AI Agents as Primary Web Consumers and Future Trends Joel offers a personal thesis that AI agents will become the main consumers of internet content. David agrees and expands on agent.txt standards and future criminal bot detection models.19:20–23:42 · Guest teaching 4/10 Proof of Humanness and Real-Time Edge Inference Joel cites the 35-year NIST identity working group and compares inference price drops to historic cloud S3 storage trends. David outlines ultra-fast edge inference models, while Joel closes with an insightful observation about ad-tech applications.0:30–3:23 · Guest disagreement 1/10 Title Card: AI Bots at the Edge Joel opens by contrasting legacy IP-blocking methods with modern bot nuances, demonstrating familiarity with network security tools. David explains why DDoS solutions handle network traffic while application-layer context is needed for modern AI bots. The tone is highly collaborative.3:23–6:42 · Guest disagreement 1/10 Modern AI Solutions and Application Context Joel asks whether vendor AI claims are just rebranded network telemetry, referencing historical standards like robots.txt. David educates on OpenAI's multi-bot setup and why voluntary standards fail without application-layer enforcement. Joel reinforces the business revenue risks of over-blocking.6:42–9:07 · Guest disagreement 0/10 Understanding OpenAI Crawlers and Real-Time Query Agents David systematically categorizes OpenAI's crawlers, from model training to real-time doc retrieval and search indexing. Joel validates the explanations with concrete user examples like querying JFK's birthday. The dynamic remains entirely conversational and informative.9:07–13:27 · Guest disagreement 1/10 Computer-Use Agents, Headless Browsers, and Action Control Joel shares technical experience handling terabits of traffic across 450k IPs and questions how fine-grained control works at scale. David walks through defense layers, including residential proxies, User-Agent verification, and TLS fingerprinting hashes like JA3 and JA4.13:27–16:20 · Guest disagreement 1/10 OSI Stack Identity and Header Fingerprinting Joel actively maps fingerprinting concepts back to the OSI stack and compares emerging cryptographic pass keys to legacy Kerberos implementations. David details HTTP header order hashing (JA4H) and Apple's Privacy Pass.16:20–19:20 · Guest disagreement 1/10 AI Agents as Primary Web Consumers and Future Trends Joel offers a personal thesis that AI agents will become the main consumers of internet content. David agrees and expands on agent.txt standards and future criminal bot detection models.19:20–23:42 · Guest disagreement 1/10 Proof of Humanness and Real-Time Edge Inference Joel cites the 35-year NIST identity working group and compares inference price drops to historic cloud S3 storage trends. David outlines ultra-fast edge inference models, while Joel closes with an insightful observation about ad-tech applications.0:30–3:23 · The host pushing back 1/10 Title Card: AI Bots at the Edge Joel opens by contrasting legacy IP-blocking methods with modern bot nuances, demonstrating familiarity with network security tools. David explains why DDoS solutions handle network traffic while application-layer context is needed for modern AI bots. The tone is highly collaborative.3:23–6:42 · The host pushing back 2/10 Modern AI Solutions and Application Context Joel asks whether vendor AI claims are just rebranded network telemetry, referencing historical standards like robots.txt. David educates on OpenAI's multi-bot setup and why voluntary standards fail without application-layer enforcement. Joel reinforces the business revenue risks of over-blocking.6:42–9:07 · The host pushing back 1/10 Understanding OpenAI Crawlers and Real-Time Query Agents David systematically categorizes OpenAI's crawlers, from model training to real-time doc retrieval and search indexing. Joel validates the explanations with concrete user examples like querying JFK's birthday. The dynamic remains entirely conversational and informative.9:07–13:27 · The host pushing back 2/10 Computer-Use Agents, Headless Browsers, and Action Control Joel shares technical experience handling terabits of traffic across 450k IPs and questions how fine-grained control works at scale. David walks through defense layers, including residential proxies, User-Agent verification, and TLS fingerprinting hashes like JA3 and JA4.13:27–16:20 · The host pushing back 2/10 OSI Stack Identity and Header Fingerprinting Joel actively maps fingerprinting concepts back to the OSI stack and compares emerging cryptographic pass keys to legacy Kerberos implementations. David details HTTP header order hashing (JA4H) and Apple's Privacy Pass.16:20–19:20 · The host pushing back 1/10 AI Agents as Primary Web Consumers and Future Trends Joel offers a personal thesis that AI agents will become the main consumers of internet content. David agrees and expands on agent.txt standards and future criminal bot detection models.19:20–23:42 · The host pushing back 2/10 Proof of Humanness and Real-Time Edge Inference Joel cites the 35-year NIST identity working group and compares inference price drops to historic cloud S3 storage trends. David outlines ultra-fast edge inference models, while Joel closes with an insightful observation about ad-tech applications.

speaking balance: gold is the host, purple is the guest (3 minute bins)

0:00 · the host 0% · guest 100%0:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%18:00 · the host 0% · guest 100%18:00 · the host 0% · guest 100%21:00 · the host 0% · guest 100%21:00 · the host 0% · guest 100%24:00 · the host 0% · guest 100%24:00 · the host 0% · guest 100%
Sharpest disagreement ▶ 17:33 Guest rejects assumption that non-human traffic implies malicious intent

In a generally non-combative interview, David mildly challenges conventional security assumptions by emphasizing that treating automated browser traffic as inherently malicious is an outdated paradigm.

Hardest push from the host ▶ 10:28 Host pushes on the reality of controlling traffic at internet scale

Joel presses David on how complex application-layer inspection can realistically function under massive terabit DDoS conditions based on his own real-world scale experience.

Biggest teaching moment ▶ 12:30 Guest explains reverse DNS lookups and TLS fingerprinting mechanics

David provides a detailed technical breakdown of how security teams verify bot identities using reverse DNS lookups and client request hashing algorithms.

The host holds their own ▶ 19:19 Host frames identity proofing around NIST working groups and gig-economy history

Joel demonstrates extensive domain authority by contextualizing AI proof-of-humanness within 35 years of NIST identity standardization efforts and gig-economy onboarding challenges.

the scores for every segment, with the reasoning behind each
ChapterTopicThe host as informed peerGuest teachingGuest disagreementThe host pushing backWhy
Title Card: AI Bots at the Edge 4311 Joel opens by contrasting legacy IP-blocking methods with modern bot nuances, demonstrating familiarity with network security tools. David explains why DDoS solutions handle network traffic while application-layer context is needed for modern AI bots. The tone is highly collaborative.
Modern AI Solutions and Application Context 5412 Joel asks whether vendor AI claims are just rebranded network telemetry, referencing historical standards like robots.txt. David educates on OpenAI's multi-bot setup and why voluntary standards fail without application-layer enforcement. Joel reinforces the business revenue risks of over-blocking.
Understanding OpenAI Crawlers and Real-Time Query Agents 4501 David systematically categorizes OpenAI's crawlers, from model training to real-time doc retrieval and search indexing. Joel validates the explanations with concrete user examples like querying JFK's birthday. The dynamic remains entirely conversational and informative.
Computer-Use Agents, Headless Browsers, and Action Control 6512 Joel shares technical experience handling terabits of traffic across 450k IPs and questions how fine-grained control works at scale. David walks through defense layers, including residential proxies, User-Agent verification, and TLS fingerprinting hashes like JA3 and JA4.
OSI Stack Identity and Header Fingerprinting 6412 Joel actively maps fingerprinting concepts back to the OSI stack and compares emerging cryptographic pass keys to legacy Kerberos implementations. David details HTTP header order hashing (JA4H) and Apple's Privacy Pass.
AI Agents as Primary Web Consumers and Future Trends 5311 Joel offers a personal thesis that AI agents will become the main consumers of internet content. David agrees and expands on agent.txt standards and future criminal bot detection models.
Proof of Humanness and Real-Time Edge Inference 7412 Joel cites the 35-year NIST identity working group and compares inference price drops to historic cloud S3 storage trends. David outlines ultra-fast edge inference models, while Joel closes with an insightful observation about ad-tech applications.

Statements from this episode (19)

Assertion Not checkable as stated
Mytton: DDoS protection has become a commodity handled by cloud providers
“The DDoS problem is still there, but it's just almost handled as a commodity these days. The network provider, your cloud provider, They'll just deal with it. And so when you're deploying an application, most of the time you just don't have to think about it.”
David Mytton Jun 16, 2025 ▶ 0:53
Insight
Mytton: Non-DDoS web security requires application-level context
“So a volumetric DDoS attack, you just want to block that at the network. You never want to see that traffic, but everything else needs the context of the application.”
David Mytton Jun 16, 2025 ▶ 2:26
Insight
Mytton: E-commerce sites should flag suspicious transactions, not block them
“If you're running an e-commerce operation, an online store, the worst thing you can do is block a transaction because then you've lost the revenue. Usually you want to then flag that order for review.”
David Mytton Jun 16, 2025 ▶ 2:59
Assertion Partly supported
Mytton: OpenAI operates four to five distinct bot types
“Particularly with AI coming in where something like OpenAI has four or five different types of bots, and some of them you might want to make a more restrictive decision over, but then others are going to be taking actions on behalf of a user search.”
David Mytton Jun 16, 2025 ▶ 4:03
Assertion Supported
Mytton: Newer bots use robots.txt to target restricted site areas
“And, but there are newer bots that are ignoring it or even sometimes using it as a way to find the parts of your site that you don't want it to access. And they will just do that anyway.”
David Mytton Jun 16, 2025 ▶ 6:24
Assertion Not checkable as stated
Mytton: Websites in AI search indexes gain more traffic and signups
“As we were seeing, sites are getting more signups, they're getting more traffic.”
David Mytton Jun 16, 2025 ▶ 7:46
Assertion Not checkable as stated
De la Garza: Traditional bot blocking was driven strictly by scale
“I mean, if I remember back why we made a lot of the decisions we made in blocking bots was strictly because of scale. So, you know, you've got 450,000 IP addresses sending you terabits of traffic through a link that only can do gigabit and you've got to just s…”
Joel de la Garza Jun 16, 2025 ▶ 10:30
Assertion Not checkable as stated
Mytton: Residential IP Proxies Render Basic ISP Blocks Ineffective
“And then you have the compounding factor of the abusers will purchase access to proxies which run on residential IP addresses. So you can't easily rely on the fact that it's part of a home ISP block anymore.”
David Mytton Jun 16, 2025 ▶ 11:49
Assertion Partly supported
Mytton: Googlebot and OpenAI Bots Can Be Validated via Reverse DNS
“And Googlebot, OpenAI, they tell you who they are, and then you can verify that by doing a reverse DNS lookup on the IP address.”
David Mytton Jun 16, 2025 ▶ 12:30
Prediction Not checkable as stated
De la Garza: Major vendors will each launch distinct request-signing standards
“Every, every large vendor is going to have their flavor. And if you're a shop and you're trying to sell to everybody, you've got to kind of work with all of them.”
Joel de la Garza Jun 16, 2025 ▶ 15:59
Prediction Not checkable as stated
De la Garza: AI agents will become primary consumers of web content
“And it seems like we're moving to a world where almost the layer you describe, the agent type activity you describe will become the primary consumer of everything on the internet.”
Joel de la Garza Jun 16, 2025 ▶ 16:42
Assertion Supported
Mytton: 50% of Current Internet Traffic Comes From Automated Bots
“50% of traffic is already bots, it's already automated”
David Mytton Jun 16, 2025 ▶ 16:52
Prediction Not checkable as stated
Mytton: Web Traffic From Autonomous AI Agents Will Explode
“We're going to see an explosion in the traffic that's coming from these tools”
David Mytton Jun 16, 2025 ▶ 17:08
Opinion
Mytton: Blanket Blocking of AI Traffic Is the Wrong Security Strategy
“Just blocking them just because they're AI is the wrong answer.”
David Mytton Jun 16, 2025 ▶ 17:12
Prediction Not checkable as stated
Mytton: AI crawlers will increasingly adhere to web rules within 18 months
“And so over the next 18 months, I think we'll see more of that, more of the AI crawlers that we want, following the rules, doing things in the right way. And it will start to split into making it a lot easier to detect the bots with criminal intent, and those …”
David Mytton Jun 16, 2025 ▶ 19:05
Assertion Not checkable as stated
De la Garza: NIST identity proofing group has stalled for 35 years
“There's a NIST working group on proofing identity that's been running, I think for 35 years, and like still hasn't really gotten to something that's implementable.”
Joel de la Garza Jun 16, 2025 ▶ 19:53
Prediction Held up
Mytton: Low-memory edge AI models will be deployed within years
“We're already seeing new edge models designed to be deployed to mobile devices and IOT that use very low amounts of system memory and can provide inference responses within milliseconds. I think those are going to start to be deployed to applications over the …”
David Mytton Jun 16, 2025 ▶ 21:32
Assertion Not checkable as stated
De la Garza: ChatGPT is 100% accurate at detecting suspicious emails
“To drop a suspicious email into a, into ChatGPT and ask if it's suspicious, and it's like a hundred percent accurate, right? Like if you want to like find sensitive information, you ask the LLM, is this sensitive information? And it's like a hundred percent ac…”
Joel de la Garza Jun 16, 2025 ▶ 22:21
Assertion Not checkable as stated
Mytton: Edge inference can filter click spam before ad auctions occur
“Super fast inference on the edge coming into a decision and for advertisers stopping click spam, that's a huge problem and being able to come to that decision before it even goes through your ad model and the auction system.”
David Mytton Jun 16, 2025 ▶ 23:54
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.