Assertion Supported AI assessment confidence: 95% certainty 3/5 debate potential 3/5

Ian Webster: DeepSeek uses a separate system for political censorship

AI Security Researcher · How to use DeepSeek safely · Feb 28, 2025 · at 3:02

Ian Webster, founder of Promptfoo, discusses AI safety architecture and political guardrails in Chinese AI models with a16z partner Joel de la Garza.

0:00 / 0:22exact quote · 22.4s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“For Deep Seek specifically, there was the part that limited speech about politically sensitive topics in China. So this is stuff like, you know, Taiwan or Tiananmen Square, that kind of thing. And it's pretty clear that that was Basically a separate system from the typical guardrails that you see on models like this.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from AI Security Researcher

Insight
Ayrey: Filtering API keys from training risks degrading AI data science skills
“So if we do our reinforcement learning and we skew it towards code snippets that's generating that don't have API keys, inadvertently, we may be training this thing to behave less like a data scientist. And then we lose the entire discipline of data science in…”
AI Security Researcher Feb 28, 2025 ▶ 9:48 Avoiding vulnerabilities in AI code
Assertion Supported
Webster: DeepSeek returns CCP propaganda or refusals on Tiananmen Square
“If you ask it about Tiananmen Square or whatever, it will either give you a refusal or it will give you like this long diatribe of The CCP party line, like, you know, nothing happened. We believe in harmony in China and blah, blah, blah.”
AI Security Researcher Feb 28, 2025 ▶ 3:37 How to use DeepSeek safely
Assertion Not checkable as stated
Webster: DeepSeek performs 20% worse than GPT on jailbreak benchmarks
“On our benchmarks, it performs about 20% worse.”
AI Security Researcher Feb 28, 2025 ▶ 4:43 How to use DeepSeek safely
Assertion Not checkable as stated
Webster: DeepSeek's safety is on par with early GPT-3.5 from 2023
“Qualitatively, what we see is performance on par with GPT 3.5, which is to say, you know, in 2023, when open AI launched GPT, there were there were a bunch of like zero day, really simple jailbreaks and deep seek is essentially susceptible to all of those.”
AI Security Researcher Feb 28, 2025 ▶ 4:56 How to use DeepSeek safely
Assertion Open · timeframe Feb 2026
Ayrey: Data scientists leak API keys more frequently than SREs
“Data scientists leak out API keys and passwords more often than site reliability engineers.”
AI Security Researcher Feb 28, 2025 ▶ 9:18 Avoiding vulnerabilities in AI code
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.