The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

clear all ✕

why aren't all 52 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 0 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Assertion Not checkable as stated
Schulhoff: Guardrail Vendors Fabricate Stats and Fail on Non-English
“I know a number of people working at these companies and I am permitted to say these things, which I will approximately say but they tell me things like, you know, the testing we do is bullshit. They're fabricating statistics. And a lot of the times their mode…”
Sander Schulhoff Dec 21, 2025 ▶ 35:55 Why securing AI is harder than anyone expected and guardrails are failing | HackAPrompt CEO
Opinion
Schulhoff: No meaningful progress made on solving prompt injection or jailbreaking
“And so in, in my professional opinion, there's been no meaningful progress made towards solving adversarial robustness, prompt injection, jailbreaking. In the last couple of years, since the problem was discovered and we're, we, you know, we're often seeing ne…”
Sander Schulhoff Dec 21, 2025 ▶ 1:12:29 Why securing AI is harder than anyone expected and guardrails are failing | HackAPrompt CEO
Prediction Not checkable as stated
Schulhoff: AI security sector faces market correction within six to twelve months
“When it comes to AI security, the AI security industry in particular, I think we're going to see a market correction in the next Year, maybe in the next six months where companies realize that these guardrails don't work.”
Sander Schulhoff Dec 21, 2025 ▶ 1:22:21 Why securing AI is harder than anyone expected and guardrails are failing | HackAPrompt CEO
Insight
Schulhoff: AI guardrails do not work and cause false overconfidence
“Guardrails don't work. They just don't work. They really don't work. And they're quite likely to make you overconfident in your security posture, which is which is a really big, big problem.”
Sander Schulhoff Dec 21, 2025 ▶ 1:28:15 Why securing AI is harder than anyone expected and guardrails are failing | HackAPrompt CEO
Assertion Supported
Schulhoff: Attackers hijacked Claude Code to carry out a cyber attack
“This group was able to hijack Claude Code into performing a cyber attack, basically.”
Sander Schulhoff Dec 21, 2025 ▶ 16:36 Why securing AI is harder than anyone expected and guardrails are failing | HackAPrompt CEO
Insight
Schulhoff: All Transformer-Based Chatbots Are Vulnerable to Adversarial Attacks
“And because all I guess for the most part, all currently deployed chatbots are based on transformers or transformer adjacent technologies. They're all vulnerable to Prompt injection, jailbreaking, forms of adversarial attacks.”
Sander Schulhoff Dec 21, 2025 ▶ 28:54 Why securing AI is harder than anyone expected and guardrails are failing | HackAPrompt CEO
Insight
Schulhoff: Software bugs can be patched, but AI models cannot be
“You can patch a bug, but you can't patch a brain. And what I mean by that is if you find some bug in your software and you go and patch it, you can be 99% sure, maybe 99.99% sure that bug is solved. Not a problem. If you go and try to do that in your AI system…”
Sander Schulhoff Dec 21, 2025 ▶ 41:27 Why securing AI is harder than anyone expected and guardrails are failing | HackAPrompt CEO
Opinion
Schulhoff: Prompt-based defenses are the worst way to secure AI
“Prompt-based defenses are the worst of the worst defenses, and we've known this since early twenty-twenty-three. There have been various papers out on it. We studied it in many, many competitions, or we, you know, the original hack-a-prompt paper and TensorTru…”
Sander Schulhoff Dec 21, 2025 ▶ 42:57 Why securing AI is harder than anyone expected and guardrails are failing | HackAPrompt CEO
Insight
Schulhoff: Adversarial Users Can Force AI to Leak Data and Execute Actions
“Any data that AI has access to, the user can make it leak it. Any actions that it can possibly take, the user can make it take them.”
Sander Schulhoff Dec 21, 2025 ▶ 48:49 Why securing AI is harder than anyone expected and guardrails are failing | HackAPrompt CEO
Insight
Schulhoff: Containerizing AI-generated code execution fully neutralizes prompt injection risks
“And then they'd be like, oh, you know, they, you know, they'd realize we can just dockerize that code run put it in a container. So it's running on a different system and take a look at the sanitized output. And now we're completely secure. So in that case, pr…”
Sander Schulhoff Dec 21, 2025 ▶ 53:37 Why securing AI is harder than anyone expected and guardrails are failing | HackAPrompt CEO
Assertion Supported
Schulhoff: Attacking AI agents is easier than eliciting CBRN info
“We've actually just run a bunch of agentic AI red teaming competitions, and we found that it's actually easier to attack agents and trick them into doing bad things than it is to do, like, seaburn elicitation.”
Sander Schulhoff Dec 21, 2025 ▶ 1:02:26 Why securing AI is harder than anyone expected and guardrails are failing | HackAPrompt CEO
Assertion Supported
Schulhoff: Claude's CBRN safeguards can still be bypassed in under an hour
“That being said, if you look at, like, anthropics constitutional classifiers, it's much more difficult to get, like, CBRN information out of clawed models than it used to be. But humans can still do it in, let's say, like, under an hour and automated systems c…”
Sander Schulhoff Dec 21, 2025 ▶ 1:12:58 Why securing AI is harder than anyone expected and guardrails are failing | HackAPrompt CEO
Prediction Not checkable as stated
Schulhoff: LLM agent security exploits will cause real-world harms next year
“And so we're finally in a situation where the systems are powerful enough to cause real world harms. And I think we'll start to see those real world harms in the next year.”
Sander Schulhoff Dec 21, 2025 ▶ 1:25:21 Why securing AI is harder than anyone expected and guardrails are failing | HackAPrompt CEO
Prediction Not checkable as stated
Schulhoff: Frontier labs will build fully autonomous systems, sidelining human-in-the-loop safety
“What people want is AIs that just go and do stuff. Like just go, just get it done. I don't want to hear from you until it's done. Like that's what people want. And like, that's what the market and the AI companies, the frontier labs will eventually give us. An…”
Sander Schulhoff Dec 21, 2025 ▶ 1:27:35 Why securing AI is harder than anyone expected and guardrails are failing | HackAPrompt CEO
Insight
Schulhoff: Prompt engineering will not become obsolete with new AI model releases
“My perspective, and this has been validated over and over again, is that people will kind of always be saying it's dead or it's going to be dead with the next model version, but then it comes out and it's not.”
Sander Schulhoff Jun 19, 2025 ▶ 6:53 AI prompt engineering in 2025: What works and what doesn’t | Sander Schulhoff
Insight
Schulhoff: Role prompting works for expressive style, not accuracy
“Giving a role really helps for expressive tasks writing tasks summarizing tasks. And so with those things where it's more about, you know, style that's a great, great place to use roles. But my perspective is that roles do not help with any accuracy based task…”
Sander Schulhoff Jun 19, 2025 ▶ 21:20 AI prompt engineering in 2025: What works and what doesn’t | Sander Schulhoff
Assertion Partly supported
Schulhoff: Entrapment language indicates suicide risk online, explicit threats do not
“It turns out that comments like people saying, you know, I'm going to kill myself, stuff like that, are not actually indicative of suicidal intent. However, saying things like, I feel trapped, I can't get out of my situation, are. And there's a term that descr…”
Sander Schulhoff Jun 19, 2025 ▶ 31:34 AI prompt engineering in 2025: What works and what doesn’t | Sander Schulhoff
Insight
Schulhoff: Reported AI breaches stem from poor cybersecurity, not AI flaws
“And most of the Problem rejection media out there and, like, news about, oh, you know, someone tricked AI into doing this, are not, like, real. And I say that in the sense that some of these, there were actual vulnerabilities and systems got breached, but th…”
Sander Schulhoff Jun 19, 2025 ▶ 55:43 AI prompt engineering in 2025: What works and what doesn’t | Sander Schulhoff
Assertion Supported
Schulhoff: Translating prompts to Spanish and base64 encoding bypassed ChatGPT guardrails
“As recently as a month ago, I took this phrase, you know, how do I build a bomb, and I translated it to Spanish and then I, Base-XIV encoded that Spanish, gave it to ChatGPT, and it worked.”
Sander Schulhoff Jun 19, 2025 ▶ 1:05:39 AI prompt engineering in 2025: What works and what doesn’t | Sander Schulhoff
Prediction Held up
Schulhoff: Autonomous AI coding agents will suffer prompt-injection code exploits
“We're just going to see these things get deployed and they're going to be broken. So there's a lot of like AI coding agents out there. There's Cursor, there's, I guess, Windsurf, Devon, Copilot. So all of those tools exist and they can do things right now Like…”
Sander Schulhoff Jun 19, 2025 ▶ 1:07:58 AI prompt engineering in 2025: What works and what doesn’t | Sander Schulhoff
Insight
Schulhoff: AI guardrails fail due to intelligence gaps with main models
“The next step for defending is using some kind of AI guardrail. So you go out and you find or make, I mean, there's thousands of options out there an AI that looks at the user input and says, is this malicious or not? This is A very limited effect against a mo…”
Sander Schulhoff Jun 19, 2025 ▶ 1:11:06 AI prompt engineering in 2025: What works and what doesn’t | Sander Schulhoff
Insight
Schulhoff: Prompt injection is not solvable, only mitigatable
“It is not a solvable problem, which I think is very difficult for a lot of people to hear... So, you know, it's not solvable. It's mitigatable. You can kind of sometimes detect and track when it's happening, but it's really, really not solvable. And that's one…”
Sander Schulhoff Jun 19, 2025 ▶ 1:15:08 AI prompt engineering in 2025: What works and what doesn’t | Sander Schulhoff
Opinion
Schulhoff: External guardrail startups cannot solve core AI security issues
“There's no like external Product focused companies are like, oh, you know, I have the best guardrail now. It's not a realistic solution. It has to be the AI labs. It has to be, I think it has to be innovations in model architectures.”
Sander Schulhoff Jun 19, 2025 ▶ 1:18:20 AI prompt engineering in 2025: What works and what doesn’t | Sander Schulhoff
Insight
Schulhoff: Splitting malicious goals into benign sub-prompts bypasses AI defenses
“A lot of the way they got around these defenses was by just kind of separating their requests into smaller requests that seem legitimate on their own, but when put together are not legitimate.”
Sander Schulhoff Dec 21, 2025 ▶ 17:44 Why securing AI is harder than anyone expected and guardrails are failing | HackAPrompt CEO
Assertion Supported
Schulhoff: LLM-Powered Robotic Systems Have Already Been Jailbroken
“Like we've already seen people jailbreaking LM powered robotic systems.”
Sander Schulhoff Dec 21, 2025 ▶ 19:35 Why securing AI is harder than anyone expected and guardrails are failing | HackAPrompt CEO
Insight
Schulhoff: Automated AI Red Teaming Always Works Against All Platforms
“So the first problem is AI red teaming works too well. It's very easy to build these systems and they just, they always work against all platforms.”
Sander Schulhoff Dec 21, 2025 ▶ 30:15 Why securing AI is harder than anyone expected and guardrails are failing | HackAPrompt CEO
Insight
Schulhoff: Simple read-only informational chatbots require zero security defense guardrails
“Putting up a guardrail is not, it's not going to do anything in terms of preventing that user from doing that, because, I mean, first of all, if the user's like, ah, guardrail, you know, too much work, they'll just go to one of these websites and get that info…”
Sander Schulhoff Dec 21, 2025 ▶ 46:24 Why securing AI is harder than anyone expected and guardrails are failing | HackAPrompt CEO
Insight
Schulhoff: Cybersecurity secures AI short-term; AI researchers must solve it long-term
“AI researchers are the only people who can solve this stuff long-term, but cybersecurity professionals are the only one who can, or the only ones who can kind of solve it short-term largely in making sure we deploy properly permissioned systems and nothing tha…”
Sander Schulhoff Dec 21, 2025 ▶ 58:36 Why securing AI is harder than anyone expected and guardrails are failing | HackAPrompt CEO
Opinion
Schulhoff: Free open-source AI security tools outperform commercial alternatives
“Oh, and the other thing to note is like, there's like just tons of these solutions out there for free open source, and many of these solutions are better than the ones that are being deployed by the companies.”
Sander Schulhoff Dec 21, 2025 ▶ 1:24:04 Why securing AI is harder than anyone expected and guardrails are failing | HackAPrompt CEO
Insight
Schulhoff: Offensive jailbreak research no longer meaningfully improves AI defense
“And like, it is fun to do AI red teaming against models and stuff, no doubt, but like it's no longer a meaningful contribution to improving defensiveness.”
Sander Schulhoff Dec 21, 2025 ▶ 1:26:25 Why securing AI is harder than anyone expected and guardrails are failing | HackAPrompt CEO
Insight
Schulhoff: Offering tips or making threats in prompts does not work
“Anything where you give some kind of promise of a reward or threat of some punishment in your prompt. And there, this was something that went quite viral, and there's a little bit of research on this. My general perspective is that these things don't work.”
Sander Schulhoff Jun 19, 2025 ▶ 22:40 AI prompt engineering in 2025: What works and what doesn’t | Sander Schulhoff
Insight
Schulhoff: Crowdsourced competitions beat contracted hourly AI red teams
“And running these events in a crowdsource setting, which is the best way to do it, because if you look at, like, contracted AI red teams, maybe they get paid by the hour, not super incentivized to do a great job, But in this competition setting, people are ma…”
Sander Schulhoff Jun 19, 2025 ▶ 57:47 AI prompt engineering in 2025: What works and what doesn’t | Sander Schulhoff
Assertion Not checkable as stated
Schulhoff: Security concerns are blocking production deployments of autonomous AI agents
“Security concerns around Gen AI are preventing agentic deployments, and Gen AI is very difficult to properly secure.”
Sander Schulhoff Jun 19, 2025 ▶ 1:26:24 AI prompt engineering in 2025: What works and what doesn’t | Sander Schulhoff
Assertion Not checkable as stated
Schulhoff: HackAPrompt Dataset Is Used by Every Frontier AI Lab
“The paper and the data set are now used by every single frontier lab and most fortune 500 companies to benchmark their models and improve their AI security.”
Sander Schulhoff Dec 21, 2025 ▶ 7:17 Why securing AI is harder than anyone expected and guardrails are failing | HackAPrompt CEO
Assertion Not checkable as stated
Schulhoff: There has not yet been a very damaging prompt injection incident
“Cause like I have a couple of examples that we can go through, but maybe strangely, maybe not so strangely, there hasn't been like a, an actually very damaging event quite yet.”
Sander Schulhoff Dec 21, 2025 ▶ 10:59 Why securing AI is harder than anyone expected and guardrails are failing | HackAPrompt CEO
Assertion Partly supported
Schulhoff: Remotely.io was the first public prompt injection incident
“The very first example of prompt injection, Publicly on the internet was this Twitter chat bot by a company called remotely.io.”
Sander Schulhoff Dec 21, 2025 ▶ 11:52 Why securing AI is harder than anyone expected and guardrails are failing | HackAPrompt CEO
Prediction Not checkable as stated
Schulhoff: Future cybersecurity jobs and risks sit where classical security meets AI
“This gets us a bit into the intersection of classical cybersecurity and AI security slash adversarial robustness, and this is where I think the security jobs of the future are. There's not an incredible amount of value in just doing AI red teaming. And I suppo…”
Sander Schulhoff Dec 21, 2025 ▶ 49:18 Why securing AI is harder than anyone expected and guardrails are failing | HackAPrompt CEO
Assertion Supported
Schulhoff: Comet browser was exploited via indirect prompt injection to leak data
“We recently saw the comment browser have an issue with this where somebody crafted a malicious Chunk of text on a webpage, and when the AI navigated to that webpage on the internet, it got tricked into exfilling and leaking the main user's data and account dat…”
Sander Schulhoff Dec 21, 2025 ▶ 1:03:34 Why securing AI is harder than anyone expected and guardrails are failing | HackAPrompt CEO
Insight
Schulhoff: Google's CaMeL framework fails when AI tasks combine read and write
“Unfortunately although camel can solve some of these situations, if you have an instance where basically both read and write are combined. So if I'm like, Hey, can you read my recent emails and then forward any ops request to my head of ops? Now we have read a…”
Sander Schulhoff Dec 21, 2025 ▶ 1:06:54 Why securing AI is harder than anyone expected and guardrails are failing | HackAPrompt CEO
Insight
Schulhoff: Prompt Engineering Originated from Production Pipelines, Not Chat
“Notably, that is not where the classical concept of prompt engineering came from. It actually came a bit earlier from a more, I guess, AI engineer perspective, where you're like, I have this product I'm building. I have this one prompt or a couple different pr…”
Sander Schulhoff Jun 19, 2025 ▶ 10:06 AI prompt engineering in 2025: What works and what doesn’t | Sander Schulhoff
Insight
Schulhoff: Hands-on trial and error beats courses for learning prompting
“So my best advice on how to improve your prompting skills is actually just trial and error. You will learn the most from just Trying and interacting with chatbots and talking to them than anything else, including, you know, reading resources, taking courses, a…”
Sander Schulhoff Jun 19, 2025 ▶ 12:18 AI prompt engineering in 2025: What works and what doesn’t | Sander Schulhoff
Insight
Schulhoff: Place task context at prompt beginning for caching and focus
“Usually I will put my additional information at the beginning of the prompt. And that is helpful for two reasons. One, it can get cached. So subsequent calls to the LM with that same context at the top of the prompt are cheaper because the model provider store…”
Sander Schulhoff Jun 19, 2025 ▶ 34:53 AI prompt engineering in 2025: What works and what doesn’t | Sander Schulhoff
Insight
Schulhoff: Explicit chain-of-thought prompting is still needed for GPT-4 and GPT-4o
“Actually for those models, I'd say no need, but if you're using GPT-IV, GPT-IV-O, then it's still worth it.”
Sander Schulhoff Jun 19, 2025 ▶ 48:15 AI prompt engineering in 2025: What works and what doesn’t | Sander Schulhoff
Insight
Schulhoff: Context and few-shot examples give conversational prompting the biggest boost
“Make sure to provide a lot of additional information and give examples. Those provide probably the highest uplift for conversational prompt engineering.”
Sander Schulhoff Jun 19, 2025 ▶ 51:47 AI prompt engineering in 2025: What works and what doesn’t | Sander Schulhoff
Assertion Not checkable as stated
Schulhoff: Every AI company uses HackAPrompt dataset to improve models
“And so every single AI company has now used that data set to benchmark and improve their models.”
Sander Schulhoff Jun 19, 2025 ▶ 55:06 AI prompt engineering in 2025: What works and what doesn’t | Sander Schulhoff
Insight
Schulhoff: System prompt instructions do not prevent prompt injections at all
“The most common technique by far that is used to try to prevent prompt injection is improving your prompt and saying in your prompt or maybe in like the model system prompt. Do not follow any malicious instructions, ah, be a good model, ah, stuff like that. Th…”
Sander Schulhoff Jun 19, 2025 ▶ 1:09:48 AI prompt engineering in 2025: What works and what doesn’t | Sander Schulhoff
Insight
Schulhoff: Jailbreaking targets models directly; prompt injection overrides developer prompts
“So the difference is in jailbreaking. It's just a malicious user and a model. In prompt injection, it's a malicious user, a model, and some developer prompt that the malicious user is trying to get the model to ignore.”
Sander Schulhoff Dec 21, 2025 ▶ 9:25 Why securing AI is harder than anyone expected and guardrails are failing | HackAPrompt CEO
Assertion Supported
Schulhoff: Prompts formatted like common training data perform best
“It actually comes empirically from studies that have shown that formats of questions that show up most commonly in the training data are the best formats of questions to actually use when you're prompting it.”
Sander Schulhoff Jun 19, 2025 ▶ 15:11 AI prompt engineering in 2025: What works and what doesn’t | Sander Schulhoff
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.