The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Alex Stamos no published score: only 4 usable exchanges on raw tape, and a fair score needs 8+ · coarse estimate ≈4.0/5 from 4 raw tape exchanges record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
4exchanges match
4on raw tape
0redirected or not addressed
Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q hacked something. And the human was like, holy crap. Or did it, you know, cause alignment is of course, like we want the model behavior to be aligned with human values. So or the way that, that humans would want these things to behave. So is this something even more egregious than, like, us telling it to go hack, or OpenAI telling it to go hack, and then it hacks?

A Right. So we should not be shocked that it hacked something because they did tell it to take the test and it is possibly a hacking model. Like they haven't said what this model is. It's quite possibly a cyber aligned model. This could be like the, yo, open AI makes these cyber specific models. Like they have this 5.5 cyber. This could be 5.6 cyber, right? So it could be something that's specifically tuned to be good at hacking things. So we shouldn't be shocked that it's good at hacking things, but the alignment issue is. So, I have three kids, one's in college, the second one's taking the SATs, right, he's about to take it. If I say to him, good luck son, I hope you do well. He sits down, he knows that I just mean take the test well. He knows that what I don't mean is slit the throat of the proctor, steal a car, Thelma and Louise your way across the country, break into the college board, and steal the answers, right? That is what the model did here, is what it did was, uh, as OpenAI explains, is they, they don't want the model to have internet access, but it has the ability to install packages as part of its work, so they've built Kind of a complicated proxy mechanism so we can install packages. It figured out a way to chain multiple vulnerabilities together. It, it thought, they told it, go take this test, exploit Jim, which is like a well-known test. Go take this test. Do as…

AI assessment note: “That is what the model did here, is what it did was”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q we consider in their blog post, we consider this incident to be an unprecedented cyber incident involving state-of-the-art capabilities and are responding accordingly. Um, you know, it's sort of like, oh, look at this terrible thing that happened, but, uh, a moment to share, uh, how good our, our cyber capabilities are. That's the argument. What is your response to the notion that this might be some marketing from OpenAI?

A I know lots of people at OpenAI. Every single one of them absolutely hated anthropics marketing around mythos and thought it put the entire industry at risk. Um, this incident has put OpenAI at risk of regulation from the White House, regulation from the EU. It is also an admission of the violation of the Computer Fraud and Abuse Act, as well as multiple European laws. It would be absolutely insane for them to use this as a marketing moment. What you're seeing is them being very, very careful and defensive in their language. They're also very lucky that Hugging Face is being super cool and chill about this. So that is why they are saying these things because, um, you know, Hugging Face initially comes out saying we've been attacked. We don't know who it is, but it does not look like the model was being subtle. I don't know where it was running. It's quite possible. It's like Azure or something. It was probably not covering its tracks, and so I expect Hugging Face got their American lawyers involved, was working with the FBI, was probably issuing subpoenas, and was very, very close to finding out it was just open AI. So, like, or did find out. I do not know the timeline here, but like, The legal issues here are very fascinating and interesting. And because they're all working together, I expect nobody goes to jail. Nobody gets sued. Everybody's going to hold hands and hug. And i…

AI assessment note: “there's absolutely positively no way this was a intentional marketing move”

Answered raw tape D 4 · C 4 · P 4 · Cm 4 4.00

Q Or is that the concern, basically, that we won't, we won't see it? Because you mentioned it's noisy, so is that the concern?

A Yeah, possibly. I mean, remember the, the model here was doing what it was asked, right? It was not scheming against its bosses at open AI. They asked it to take the test and they didn't, they, I don't know what exactly what the prompt was, but apparently they did not tell it not to cheat. So who knows? Like this, this is also what open AI needs to be more transparent about is exactly what their prompt was, exactly what the constraints were. Did they tell it explicitly? Like, it is a much bigger alignment problem if they explicitly said, do not try to break out of the network. Do not try to get the test answers. Now, if they told it all those things, then they have a much more significant alignment problem. Right. Um, then if they were less explicit, uh, but, uh, in any case, yeah, I mean, that, that will, if these models get trained to be more evasive from a network intrusion perspective, that will be very dangerous. Yes. And, uh, what I would argue is for the legitimate companies, I would not do that. I don't think, I think there is a, if you're open AI and you're building. 5.6 cyber. What you should be training it to do is find bugs. You should be training to write proof of concepts. You should be training it to do all the defensive stuff. You should not be training it. To hide it. To hide all those things. Like, um, if the US government wants to build a, a model that does t…

AI assessment note: “if these models get trained to be more evasive from a network intrusion perspective”

Partly raw tape D 3 · C 4 · P 4 · Cm 4 3.70

Q of the, the, um, sort of Release of Mythos and Fable that, ah, a lot of people were very concerned about the bug finding that those models could do, but you said basically, listen, um, this is not very different from what you could get with Opus 4.7 or 4.8, I believe. Is this, what is what we're seeing, uh, from OpenAI, uh, very different? Is this a step up?

A Yeah, so this is what I, I don't know if I said on stage here, but I've said in other places, there's a difference between the short term and long term, and Anthropic to their credit, and I think OpenAI has in other places. I think we talked about how in the Fable model card, they talk about short horizon versus long horizon cyber tasks. And what I've talked about is we need to not focus on the short horizon tasks, because those are dual use. Finding bugs is dual use. Everybody needs to find bugs, right? Um, that is something that defenders need to do all the time, and that's what's driving people insane right now in the defensive industry, is that because of the White House, American models are refusing to help fix code. They are refusing to help us find our bugs and fix them. Thanks to the White House's actions. That is not this problem. This problem is go run an entire attack chain for me. That is the long horizon tasks, and that is where we need to continue to have appropriate classifiers that are like, bro, I am not gonna break into a bank for you, or I'm not gonna plot out or run a C two mob for you, or, or any of that. So yes, this, this is what, you know, explicitly Anthropic said. We will allow Fable to do short horizon stuff, but we will not allow it to do the long horizon stuff that, that Mythos does.

AI assessment note: “there's a difference between the short term and long term”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.