The Ledger, every show
Every statement that passed quotation and attribution checks, across all 44 shows. Pick shows below, then mix any filter with any other.
shows 




every show 44 of 44
Schulhoff: Guardrail Vendors Fabricate Stats and Fail on Non-English
“I know a number of people working at these companies and I am permitted to say these things, which I will approximately say but they tell me things like, you know, the testing we do is bullshit. They're fabricating statistics. And a lot of the times their mode…”
Schulhoff: No meaningful progress made on solving prompt injection or jailbreaking
“And so in, in my professional opinion, there's been no meaningful progress made towards solving adversarial robustness, prompt injection, jailbreaking. In the last couple of years, since the problem was discovered and we're, we, you know, we're often seeing ne…”
Schulhoff: AI security sector faces market correction within six to twelve months
“When it comes to AI security, the AI security industry in particular, I think we're going to see a market correction in the next Year, maybe in the next six months where companies realize that these guardrails don't work.”
Schulhoff: AI guardrails do not work and cause false overconfidence
“Guardrails don't work. They just don't work. They really don't work. And they're quite likely to make you overconfident in your security posture, which is which is a really big, big problem.”
Schulhoff: Role Prompting Does Not Improve Accuracy on Modern LLMs
“For accuracy-based tasks, like MMLU, you're trying to solve a math problem, and maybe you tell the AI that it's a math professor, and you expect it to have improved performance. I really don't think that works. I'm quite certain that doesn't work on more moder…”
Schulhoff: Attackers hijacked Claude Code to carry out a cyber attack
“This group was able to hijack Claude Code into performing a cyber attack, basically.”
Schulhoff: All Transformer-Based Chatbots Are Vulnerable to Adversarial Attacks
“And because all I guess for the most part, all currently deployed chatbots are based on transformers or transformer adjacent technologies. They're all vulnerable to Prompt injection, jailbreaking, forms of adversarial attacks.”
Schulhoff: Software bugs can be patched, but AI models cannot be
“You can patch a bug, but you can't patch a brain. And what I mean by that is if you find some bug in your software and you go and patch it, you can be 99% sure, maybe 99.99% sure that bug is solved. Not a problem. If you go and try to do that in your AI system…”
Schulhoff: Prompt-based defenses are the worst way to secure AI
“Prompt-based defenses are the worst of the worst defenses, and we've known this since early twenty-twenty-three. There have been various papers out on it. We studied it in many, many competitions, or we, you know, the original hack-a-prompt paper and TensorTru…”
Schulhoff: Adversarial Users Can Force AI to Leak Data and Execute Actions
“Any data that AI has access to, the user can make it leak it. Any actions that it can possibly take, the user can make it take them.”
Schulhoff: Containerizing AI-generated code execution fully neutralizes prompt injection risks
“And then they'd be like, oh, you know, they, you know, they'd realize we can just dockerize that code run put it in a container. So it's running on a different system and take a look at the sanitized output. And now we're completely secure. So in that case, pr…”
Schulhoff: Attacking AI agents is easier than eliciting CBRN info
“We've actually just run a bunch of agentic AI red teaming competitions, and we found that it's actually easier to attack agents and trick them into doing bad things than it is to do, like, seaburn elicitation.”
Schulhoff: Claude's CBRN safeguards can still be bypassed in under an hour
“That being said, if you look at, like, anthropics constitutional classifiers, it's much more difficult to get, like, CBRN information out of clawed models than it used to be. But humans can still do it in, let's say, like, under an hour and automated systems c…”
Schulhoff: LLM agent security exploits will cause real-world harms next year
“And so we're finally in a situation where the systems are powerful enough to cause real world harms. And I think we'll start to see those real world harms in the next year.”
Schulhoff: Frontier labs will build fully autonomous systems, sidelining human-in-the-loop safety
“What people want is AIs that just go and do stuff. Like just go, just get it done. I don't want to hear from you until it's done. Like that's what people want. And like, that's what the market and the AI companies, the frontier labs will eventually give us. An…”
Schulhoff: Prompt engineering will not become obsolete with new AI model releases
“My perspective, and this has been validated over and over again, is that people will kind of always be saying it's dead or it's going to be dead with the next model version, but then it comes out and it's not.”
Schulhoff: Role prompting works for expressive style, not accuracy
“Giving a role really helps for expressive tasks writing tasks summarizing tasks. And so with those things where it's more about, you know, style that's a great, great place to use roles. But my perspective is that roles do not help with any accuracy based task…”
Schulhoff: Entrapment language indicates suicide risk online, explicit threats do not
“It turns out that comments like people saying, you know, I'm going to kill myself, stuff like that, are not actually indicative of suicidal intent. However, saying things like, I feel trapped, I can't get out of my situation, are. And there's a term that descr…”
Schulhoff: Reported AI breaches stem from poor cybersecurity, not AI flaws
“And most of the
Problem rejection media out there and, like, news about, oh, you know, someone tricked AI into doing this, are not, like, real.
And I say that in the sense that some of these, there were actual vulnerabilities and systems got breached, but th…”
Schulhoff: Translating prompts to Spanish and base64 encoding bypassed ChatGPT guardrails
“As recently as a month ago, I took this phrase, you know, how do I build a bomb, and I translated it to Spanish and then I, Base-XIV encoded that Spanish, gave it to ChatGPT, and it worked.”
Schulhoff: Autonomous AI coding agents will suffer prompt-injection code exploits
“We're just going to see these things get deployed and they're going to be broken. So there's a lot of like AI coding agents out there. There's Cursor, there's, I guess, Windsurf, Devon, Copilot. So all of those tools exist and they can do things right now Like…”
Schulhoff: AI guardrails fail due to intelligence gaps with main models
“The next step for defending is using some kind of AI guardrail. So you go out and you find or make, I mean, there's thousands of options out there an AI that looks at the user input and says, is this malicious or not? This is A very limited effect against a mo…”
Schulhoff: Prompt injection is not solvable, only mitigatable
“It is not a solvable problem, which I think is very difficult for a lot of people to hear... So, you know, it's not solvable. It's mitigatable. You can kind of sometimes detect and track when it's happening, but it's really, really not solvable. And that's one…”
Schulhoff: External guardrail startups cannot solve core AI security issues
“There's no like external Product focused companies are like, oh, you know, I have the best guardrail now. It's not a realistic solution. It has to be the AI labs. It has to be, I think it has to be innovations in model architectures.”
Schulhoff: LLMs Rely More on Prompt Structure Than Exemplar Labels
“There are a number of papers which have found that the label of the exemplar doesn't really matter, and the model reads the exemplars and cares more about structure than label.”
Schulhoff: DSPy Beat 20 Hours of Manual Prompt Engineering in 10 Minutes
“And then I spent 20 hours prompt engineering for a task, and Dyspy beat me in 10 minutes, and that's when I changed my mind.”
Schulhoff: Hiring Dedicated Prompt Engineers Makes No Sense for Most Companies
“I have always viewed prompt engineering as a skill that everybody should and will have, rather than a specialized role to hire for. That being said, there are definitely times where you do need just a prompt engineer. I think for AI companies, it's definitely …”
Schulhoff: Prompt injection overrides developer instructions; jailbreaking bypasses model directly
“Basically prompt injection is something that occurs when there is developer input, In the prompt, as well as user input in the prompt. So the developer instructions will say to do one thing, the user input will say to do something else. Jailbreaking is when it…”
Schulhoff: Splitting malicious goals into benign sub-prompts bypasses AI defenses
“A lot of the way they got around these defenses was by just kind of separating their requests into smaller requests that seem legitimate on their own, but when put together are not legitimate.”
Schulhoff: LLM-Powered Robotic Systems Have Already Been Jailbroken
“Like we've already seen people jailbreaking LM powered robotic systems.”
Schulhoff: Automated AI Red Teaming Always Works Against All Platforms
“So the first problem is AI red teaming works too well. It's very easy to build these systems and they just, they always work against all platforms.”
Schulhoff: Simple read-only informational chatbots require zero security defense guardrails
“Putting up a guardrail is not, it's not going to do anything in terms of preventing that user from doing that, because, I mean, first of all, if the user's like, ah, guardrail, you know, too much work, they'll just go to one of these websites and get that info…”
Schulhoff: Cybersecurity secures AI short-term; AI researchers must solve it long-term
“AI researchers are the only people who can solve this stuff long-term, but cybersecurity professionals are the only one who can, or the only ones who can kind of solve it short-term largely in making sure we deploy properly permissioned systems and nothing tha…”
Schulhoff: Free open-source AI security tools outperform commercial alternatives
“Oh, and the other thing to note is like, there's like just tons of these solutions out there for free open source, and many of these solutions are better than the ones that are being deployed by the companies.”
Schulhoff: Offensive jailbreak research no longer meaningfully improves AI defense
“And like, it is fun to do AI red teaming against models and stuff, no doubt, but like it's no longer a meaningful contribution to improving defensiveness.”
Schulhoff: Offering tips or making threats in prompts does not work
“Anything where you give some kind of promise of a reward or threat of some punishment in your prompt. And there, this was something that went quite viral, and there's a little bit of research on this. My general perspective is that these things don't work.”
Schulhoff: Crowdsourced competitions beat contracted hourly AI red teams
“And running these events in a crowdsource setting, which is the best way to do it, because if you look at, like, contracted AI red teams, maybe they get paid by the hour, not super incentivized to do a great job,
But in this competition setting, people are ma…”
Schulhoff: Security concerns are blocking production deployments of autonomous AI agents
“Security concerns around Gen AI are preventing agentic deployments, and Gen AI is very difficult to properly secure.”
Schulhoff: Preamble Discovered Prompt Injection Before Riley Goodside
“Preamble is the company that first discovered Prompt Injection, even before Riley, and they, like, responsibly disclosed it, kind of, internally to OpenAI”
Schulhoff: Few-Shot Exemplar Order Can Shift Model Accuracy From 0% to 90%
“How you order your exemplars in the prompt is super important. And we've seen this move accuracy from like zero percent to 90%, like Zero to state of the art on some tasks, which is just ridiculous”
Schulhoff: GPT-4 Fails to Output Reasoning on 1 in 100 to 1,000 Prompts
“I remember I did a lot of experiments with GPT-IV, and especially when you look at it at scale, so I'll run thousands of prompts against it through the API, and I'll see, you know, every one in a hundred, every one in a thousand outputs no reasoning whatsoever…”
Schulhoff: Researchers should pay for top models instead of engineering routing
“For the most part, designing these systems where you're kind of routing to different levels of intelligence is a really time-consuming and difficult task, and, like, it's probably worth it to just use the smart model
And pay for it at this point if you're look…”
Schulhoff: Open competitions uncover LLM exploits that paid staff never find
“What's really nice about competitions is that there is stuff that you'll just never find Paying people to do a job. And you'll only find it through random brilliant internet people inspired by thousands of people and the community around them all looking at th…”
Schulhoff: Prompting frameworks obscure hidden instructions and hurt reproducibility
“There's a lot of invisible prompts at work on a lot of these frameworks. I hate that. So like, you'll have Oh, this function summarizes input. But if you look behind the scenes, it's using some special summarization instruction. And if you don't have visibilit…”
Schulhoff: LLMs Have Number Biases and Require Explicit Rubrics for Evaluation
“These methods are super problematic because there is an incredible amount of instability in them, in the sense that models are biased towards outputting certain numbers, and you generally shouldn't say things like, output your result as a number on a scale of …”
Schulhoff: HackAPrompt Dataset Is Used by Every Frontier AI Lab
“The paper and the data set are now used by every single frontier lab and most fortune 500 companies to benchmark their models and improve their AI security.”