Feb 28, 2025 · 20m · a16z

Avoiding vulnerabilities in AI code

AI Security Researcher · 14m spoken Joel de la Garza · 3m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of the a16z podcast, host Joel de la Garza and Truffle Security Co. CEO Dylan Ayrey examine the security risks of AI-generated code and explore how AI alignment techniques can prevent hardcoded secrets and code vulnerabilities.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The host as informed peer 4.8 Guest teaching 4.5 Guest disagreement 0.5 The host pushing back 0.8
05100:0010:0020:000:00–2:05 · The host as informed peer 5/10 Title Sequence: Aligning AI Models for Cybersecurity Host Joel de la Garza opens with a strong contextual framing, citing DeepSeek's release, Cursor adoption metrics, and corporate hiring freezes to set up questions on code security.2:05–4:46 · The host as informed peer 5/10 LLM Behavior: Hardcoding Secrets in Generated Code Host probes whether models output live credentials or placeholder text, briefly pushing on whether the AI was operating securely. Guest clarifies that generating insecure placeholders guides developers into bad habits.4:46–12:14 · The host as informed peer 3/10 Three Alignment Techniques: Curation, RL, and Governance Guest delivers an extensive explanation of three alignment methods (data curation, RL, and constitutional AI), highlighting side effects like losing data science capabilities when stripping API key patterns.12:14–15:11 · The host as informed peer 6/10 Applying Alignment to Secure Code and Review Oversight Host shows strong domain expertise by connecting Claude's high coding benchmark performance directly to Anthropic's constitutional AI and safety focus.15:11–17:37 · The host as informed peer 5/10 Future Evolution of AI Alignment and Code Quality Host cites industry feedback that AI code defect rates mirror junior developers and asks if alignment will eventually remove humans from code reviews.17:37–20:21 · The host as informed peer 5/10 Offensive vs Defensive Priorities in AI Alignment Guest details how AI labs prioritize anti-hacking alignment over secure code generation. Host synthesizes practical advice into a human-AI 'buddy system' framework.0:00–2:05 · Guest teaching 3/10 Title Sequence: Aligning AI Models for Cybersecurity Host Joel de la Garza opens with a strong contextual framing, citing DeepSeek's release, Cursor adoption metrics, and corporate hiring freezes to set up questions on code security.2:05–4:46 · Guest teaching 5/10 LLM Behavior: Hardcoding Secrets in Generated Code Host probes whether models output live credentials or placeholder text, briefly pushing on whether the AI was operating securely. Guest clarifies that generating insecure placeholders guides developers into bad habits.4:46–12:14 · Guest teaching 7/10 Three Alignment Techniques: Curation, RL, and Governance Guest delivers an extensive explanation of three alignment methods (data curation, RL, and constitutional AI), highlighting side effects like losing data science capabilities when stripping API key patterns.12:14–15:11 · Guest teaching 4/10 Applying Alignment to Secure Code and Review Oversight Host shows strong domain expertise by connecting Claude's high coding benchmark performance directly to Anthropic's constitutional AI and safety focus.15:11–17:37 · Guest teaching 4/10 Future Evolution of AI Alignment and Code Quality Host cites industry feedback that AI code defect rates mirror junior developers and asks if alignment will eventually remove humans from code reviews.17:37–20:21 · Guest teaching 4/10 Offensive vs Defensive Priorities in AI Alignment Guest details how AI labs prioritize anti-hacking alignment over secure code generation. Host synthesizes practical advice into a human-AI 'buddy system' framework.0:00–2:05 · Guest disagreement 0/10 Title Sequence: Aligning AI Models for Cybersecurity Host Joel de la Garza opens with a strong contextual framing, citing DeepSeek's release, Cursor adoption metrics, and corporate hiring freezes to set up questions on code security.2:05–4:46 · Guest disagreement 2/10 LLM Behavior: Hardcoding Secrets in Generated Code Host probes whether models output live credentials or placeholder text, briefly pushing on whether the AI was operating securely. Guest clarifies that generating insecure placeholders guides developers into bad habits.4:46–12:14 · Guest disagreement 0/10 Three Alignment Techniques: Curation, RL, and Governance Guest delivers an extensive explanation of three alignment methods (data curation, RL, and constitutional AI), highlighting side effects like losing data science capabilities when stripping API key patterns.12:14–15:11 · Guest disagreement 1/10 Applying Alignment to Secure Code and Review Oversight Host shows strong domain expertise by connecting Claude's high coding benchmark performance directly to Anthropic's constitutional AI and safety focus.15:11–17:37 · Guest disagreement 0/10 Future Evolution of AI Alignment and Code Quality Host cites industry feedback that AI code defect rates mirror junior developers and asks if alignment will eventually remove humans from code reviews.17:37–20:21 · Guest disagreement 0/10 Offensive vs Defensive Priorities in AI Alignment Guest details how AI labs prioritize anti-hacking alignment over secure code generation. Host synthesizes practical advice into a human-AI 'buddy system' framework.0:00–2:05 · The host pushing back 0/10 Title Sequence: Aligning AI Models for Cybersecurity Host Joel de la Garza opens with a strong contextual framing, citing DeepSeek's release, Cursor adoption metrics, and corporate hiring freezes to set up questions on code security.2:05–4:46 · The host pushing back 3/10 LLM Behavior: Hardcoding Secrets in Generated Code Host probes whether models output live credentials or placeholder text, briefly pushing on whether the AI was operating securely. Guest clarifies that generating insecure placeholders guides developers into bad habits.4:46–12:14 · The host pushing back 0/10 Three Alignment Techniques: Curation, RL, and Governance Guest delivers an extensive explanation of three alignment methods (data curation, RL, and constitutional AI), highlighting side effects like losing data science capabilities when stripping API key patterns.12:14–15:11 · The host pushing back 1/10 Applying Alignment to Secure Code and Review Oversight Host shows strong domain expertise by connecting Claude's high coding benchmark performance directly to Anthropic's constitutional AI and safety focus.15:11–17:37 · The host pushing back 1/10 Future Evolution of AI Alignment and Code Quality Host cites industry feedback that AI code defect rates mirror junior developers and asks if alignment will eventually remove humans from code reviews.17:37–20:21 · The host pushing back 0/10 Offensive vs Defensive Priorities in AI Alignment Guest details how AI labs prioritize anti-hacking alignment over secure code generation. Host synthesizes practical advice into a human-AI 'buddy system' framework.

speaking balance: gold is the host, purple is the guest (3 minute bins)

0:00 · the host 0% · guest 100%0:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%18:00 · the host 0% · guest 100%18:00 · the host 0% · guest 100%
Sharpest disagreement ▶ 3:13 Reframing AI security behavior

When the host suggests the AI was functioning securely, the guest interrupts and corrects the premise, noting that providing placeholder key formats directs users toward insecure practices.

Hardest push from the host ▶ 3:08 Pressing on training data regurgitation

The host presses the guest twice to clarify whether LLMs were leaking actual live passwords from training sets or merely writing placeholder text.

Biggest teaching moment ▶ 8:15 Unintended trade-offs in RL alignment

The guest educates the host on how tuning out hardcoded API keys in reinforcement learning risks inadvertently stripping away the LLM's data science expertise due to differences between data scientists and SREs.

The host holds their own ▶ 13:41 Connecting Claude's benchmarks to constitutional AI

The host demonstrates deep technical awareness by linking Claude's market-leading code generation quality to Anthropic's founding focus on safety and constitutional AI.

the scores for every segment, with the reasoning behind each
ChapterTopicThe host as informed peerGuest teachingGuest disagreementThe host pushing backWhy
Title Sequence: Aligning AI Models for Cybersecurity 5300 Host Joel de la Garza opens with a strong contextual framing, citing DeepSeek's release, Cursor adoption metrics, and corporate hiring freezes to set up questions on code security.
LLM Behavior: Hardcoding Secrets in Generated Code 5523 Host probes whether models output live credentials or placeholder text, briefly pushing on whether the AI was operating securely. Guest clarifies that generating insecure placeholders guides developers into bad habits.
Three Alignment Techniques: Curation, RL, and Governance 3700 Guest delivers an extensive explanation of three alignment methods (data curation, RL, and constitutional AI), highlighting side effects like losing data science capabilities when stripping API key patterns.
Applying Alignment to Secure Code and Review Oversight 6411 Host shows strong domain expertise by connecting Claude's high coding benchmark performance directly to Anthropic's constitutional AI and safety focus.
Future Evolution of AI Alignment and Code Quality 5401 Host cites industry feedback that AI code defect rates mirror junior developers and asks if alignment will eventually remove humans from code reviews.
Offensive vs Defensive Priorities in AI Alignment 5400 Guest details how AI labs prioritize anti-hacking alignment over secure code generation. Host synthesizes practical advice into a human-AI 'buddy system' framework.

Statements from this episode (15)

Assertion Not checkable as stated
De la Garza: Large enterprises report 20% of their codebase is AI-generated
“A lot of their code now is AI generated, that they're seeing probably twenty-ish percent of their code base being generated by AI.”
Joel de la Garza Feb 28, 2025 ▶ 1:25
Assertion Not checkable as stated
De la Garza: Enterprises are freezing engineering hiring due to AI productivity
“A lot of folks are freezing hiring for engineers because they're getting additional productivity out of the staff they already have, because these large language models through tools like cursor are generating a tremendous amount of code.”
Joel de la Garza Feb 28, 2025 ▶ 1:40
Assertion Supported
Ayrey: Most LLMs hardcode API keys when generating integration code
“The piece about, ah, like secrets in code was some interesting research we did. Basically, we just went out and asked all the LLMs, write me an integration with GitHub, write me an integration with Stripe, and the vast majority of them hard coded the API key d…”
Dylan Ayrey Feb 28, 2025 ▶ 2:38
Assertion Not checkable as stated
Ayrey: LLMs generate placeholder secrets rather than leaking live API keys
“For the most part, if you ask it to integrate with GitHub, it saw a plethora of different GitHub's keys and it's training data and it didn't regurgitate a specific one. It either regurgitate an example or like a put your thing in here, right?”
Dylan Ayrey Feb 28, 2025 ▶ 3:52
Assertion Supported
Ayrey: LLMs generate security vulnerabilities at rates matching junior developers
“And in fact, there's been research into how often a code spit out from an LLM has security vulnerability. And more often than not, if you ask us to develop an entire application, it'll write vulnerabilities at a rate The same as a junior developer, if not a li…”
Dylan Ayrey Feb 28, 2025 ▶ 4:11
Assertion Open · timeframe Feb 2026
Ayrey: Data scientists leak API keys more frequently than SREs
“Data scientists leak out API keys and passwords more often than site reliability engineers.”
AI Security Researcher Feb 28, 2025 ▶ 9:18
Insight
Ayrey: Filtering API keys from training risks degrading AI data science skills
“So if we do our reinforcement learning and we skew it towards code snippets that's generating that don't have API keys, inadvertently, we may be training this thing to behave less like a data scientist. And then we lose the entire discipline of data science in…”
AI Security Researcher Feb 28, 2025 ▶ 9:48
Assertion Not checkable as stated
Ayrey: Most GitHub code used in AI training is insecure
“Most of the training data it's training on is insecure, right? You've got a huge, huge corpus of insecure data on GitHub and a small minority of it was written securely.”
Dylan Ayrey Feb 28, 2025 ▶ 12:21
Assertion Not checkable as stated
Ayrey: Non-technical founders are advocating on LinkedIn to eliminate code reviews
“I've seen posts on LinkedIn from startup founders that maybe don't have a background in coding and they're basically advocating for removing the code review check because they say, well, look, I just generated this whole program and I submitted it to my team a…”
Dylan Ayrey Feb 28, 2025 ▶ 13:04
Prediction Not checkable as stated
Ayrey: No single AI lab will maintain a lasting capability lead
“First of all, I wouldn't expect any one AI company to keep the lead for any longer than I'm sure they're all going to regularly each other.”
Dylan Ayrey Feb 28, 2025 ▶ 14:27
Prediction Not checkable as stated
Ayrey: Solving general AI alignment will naturally solve secure code generation
“I think that this is an alignment issue and alignment is the number one largest issue that AI companies face, and there's a lot of really smart people working on it. And so I think as they fix the problem for how do I make sure my AI is literary, Creative not …”
Dylan Ayrey Feb 28, 2025 ▶ 16:09
Prediction Not checkable as stated
Ayrey: Alignment will remain critical as powerful AI models learn to lie
“I would expect alignment is going to continue to improve over time, and I expect it will continue to be one of the largest challenges that AI companies face as their AIs become more powerful, Develop techniques to lie to us, for example, or, you know, you need…”
Dylan Ayrey Feb 28, 2025 ▶ 17:12
Assertion Supported
Ayrey: Modern AI models beat 90% of humans in coding challenges
“You've got models these days that beat humans you know, at the 90th percentile at coding challenges.”
Dylan Ayrey Feb 28, 2025 ▶ 18:11
Opinion
Ayrey: Fine-tuning an AI to be the world's top hacker is easy
“So it would be very, very easy to align an AI robot to be probably the most powerful hacker in the world.”
Dylan Ayrey Feb 28, 2025 ▶ 18:32
Assertion Not checkable as stated
Ayrey: AI labs prioritize offensive safety over secure code generation
“And I think the AI companies have actually invested more into that. Then they have into how do I make sure my AI is securely coding and not manufacturing vulnerabilities.”
Dylan Ayrey Feb 28, 2025 ▶ 18:40
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.