Nicholas Carlini

Research Scientist, Anthropic · 1 appearance on the record.

computed by AI from the episodes · how this works → · full disclaimer →

scientistauthornicholas.carlini.com ↗Wikipedia ↗

Nicholas Carlini is a researcher in adversarial machine learning and AI cybersecurity known for co-creating the Carlini & Wagner attacks. He holds a Ph.D. from UC Berkeley and previously conducted research at Google Brain and DeepMind before joining Anthropic.

31statements → 10claims → 8claims resolved → 100%fully supported → 4.03/5average certainty → 2.03/5average debate potential → 12said about them ↓

8 supported 0 partly supported 0 contradicted 1 not yet assessed 1 not checkable as stated how the 10 claims stand · each chip opens the sources

10 assertions · 3 opinions · 15 insights · 3 disclosures · every statement was checked. The predictions and assertions are the 10 claims: statements the public record can support or contradict. 8 are resolved, 1 is not yet assessed, and 1 names no date, number or outcome precise enough to check. Everything else (opinions, insights, what ifs, disclosures) can never be settled by the record, so it carries no assessment.

The record, in short

What the tape says about how Nicholas argues and how the claims held up. Everything they said, and everything said about them, is in the tabs below.

Their most notable supported claim

Assertion Supported
Attackers can poison the LAION dataset simply by purchasing expired domains
“Here's this new dataset. It is being distributed in such a way that anyone in the world can buy domains that let you then inject arbitrary images in the dataset.”
Nicholas Carlini Aug 28, 2024 ▶ 55:14 Personal benchmarks vs HumanEval - with Nicholas Carlini of DeepMind

Expressed certainty vs assessment result

none yet certainty 1
none yet certainty 2
none yet certainty 3
100% certainty 4
100% certainty 5

weighted support: a fully supported claim counts one, a partly supported claim counts half. Each filled bar is clickable and opens exactly those claims; "none yet" means nothing said at that certainty level has resolved yet

Everything Nicholas Carlini said on Latent Space that made the record, most notable first. Filter by type, assessment or year in the ledger →

Insight
Carlini: If LLMs always give desired answers, questions aren't hard enough
“When you're using these models, if you're getting the answer you want, always, it means you're not asking them hard enough questions.”
Nicholas Carlini Aug 28, 2024 ▶ 15:30 Personal benchmarks vs HumanEval - with Nicholas Carlini of DeepMind
Insight
Carlini: 90% of scientific research is routine work that AI can automate
“90% of this is not doing something new. Like, 90% of this is like doing things a million people have done before, and then a little bit of something that was new. There's a reason why we say we stand on the shoulders of giants. It's true. Almost everything tha…”
Nicholas Carlini Aug 28, 2024 ▶ 19:50 Personal benchmarks vs HumanEval - with Nicholas Carlini of DeepMind
Insight
Carlini: Users should build personalized AI benchmarks instead of relying on public leaderboards
“The argument that I tried to lay out in this post is that more people should make benchmarks that are tailored to them.”
Nicholas Carlini Aug 28, 2024 ▶ 39:10 Personal benchmarks vs HumanEval - with Nicholas Carlini of DeepMind
Insight
Carlini: GPT-4 would exist identically without adversarial machine learning research
“Nothing about GPT-IV would be at all different if the field of, like the entire field of Everson machine learning disappeared. Like everything to do with Everson examples, like all of the, like for the most part, like GPT-IV would exist identically.”
Nicholas Carlini Aug 28, 2024 ▶ 58:56 Personal benchmarks vs HumanEval - with Nicholas Carlini of DeepMind
Opinion
Carlini: Most AI commentators spin arguments based on ideology rather than reality
“I feel like most people who write about language models being good or bad, some underlying message of like, you know, they have their camp and their camp is like, AI is bad or AI is good or whatever. And they like, they spin whatever they're gonna say accordin…”
Nicholas Carlini Aug 28, 2024 ▶ 5:40 Personal benchmarks vs HumanEval - with Nicholas Carlini of DeepMind
Insight
Carlini: Copying and pasting error messages effectively creates a coding agent
“Currently though, make a model into an agent by just copying and pasting error messages for the most part. And that's what I do is, you know, you run it and it gives you some code that doesn't work and either I'll fix the code or it will give me buggy code and…”
Nicholas Carlini Aug 28, 2024 ▶ 10:17 Personal benchmarks vs HumanEval - with Nicholas Carlini of DeepMind
Insight
Carlini: Language models should not be trusted in adversarial situations
“My research says is entirely on this. Like you probably shouldn't trust these models to do the things in adversarial situations.”
Nicholas Carlini Aug 28, 2024 ▶ 23:21 Personal benchmarks vs HumanEval - with Nicholas Carlini of DeepMind
Insight
Carlini: Always qualify claims about AI with 'for current models'
“Whenever someone says X is true about language models, you should always append the suffix for current models, because I'll be the first to admit I was one of the people who was very much on the opinion that these language models are fun toys and are going to …”
Nicholas Carlini Aug 28, 2024 ▶ 24:36 Personal benchmarks vs HumanEval - with Nicholas Carlini of DeepMind
Insight
Carlini: Unpopular AI benchmarks protect against model contamination and overfitting
“And by having a benchmark that is not very popular, you can be relatively certain that no one has tried to optimize their model for your benchmark.”
Nicholas Carlini Aug 28, 2024 ▶ 40:31 Personal benchmarks vs HumanEval - with Nicholas Carlini of DeepMind
Insight
Carlini: LLM-as-a-judge is almost always accurate when prompted correctly
“I've inspected the outputs of these and like, they're almost always correct. If you sort of, if you ask the model to judge these things in the right way, they're very good at being able to tell this.”
Nicholas Carlini Aug 28, 2024 ▶ 42:52 Personal benchmarks vs HumanEval - with Nicholas Carlini of DeepMind
Opinion
Carlini: ML security research failed to adapt to real-world systems
“And then machine learning started to work. And the thing that bothered me is it seems like the other machine learning community didn't then try and adapt and try and actually start studying real problems.”
Nicholas Carlini Aug 28, 2024 ▶ 54:13 Personal benchmarks vs HumanEval - with Nicholas Carlini of DeepMind
Assertion Supported
Attackers can poison the LAION dataset simply by purchasing expired domains
“Here's this new dataset. It is being distributed in such a way that anyone in the world can buy domains that let you then inject arbitrary images in the dataset.”
Nicholas Carlini Aug 28, 2024 ▶ 55:14 Personal benchmarks vs HumanEval - with Nicholas Carlini of DeepMind
Assertion Supported
Carlini extracted production models from Google and OpenAI with legal permission
“We ran the attack that let us, yeah, stole several of OpenAI's models. With their permission... We notified everyone who was vulnerable to this attack. Some Google models were vulnerable. Some open AM models were vulnerable. There were one or two other people …”
Nicholas Carlini Aug 28, 2024 ▶ 57:21 Personal benchmarks vs HumanEval - with Nicholas Carlini of DeepMind
Assertion Supported
Prompting ChatGPT to repeat a word indefinitely leaks verbatim training data
“One of my co-authors, Milad was working on some other random experiments, and he figured out that if you prompt ChatGPT to repeat a word forever, then it will repeat the word many, many, many times in a row, and then like explode and like just start doing rand…”
Nicholas Carlini Aug 28, 2024 ▶ 1:02:16 Personal benchmarks vs HumanEval - with Nicholas Carlini of DeepMind
Assertion Supported
Carlini: ChatGPT emitted verbatim 50+ word sequences from internet training data
“And what I can say is that the output of the model was a verbatim, at least 50 word in a row match. To some other document that appeared on the internet previously.”
Nicholas Carlini Aug 28, 2024 ▶ 1:03:03 Personal benchmarks vs HumanEval - with Nicholas Carlini of DeepMind
Insight
Carlini: Imperfect LLMs remain useful because users already distrust internet content
“You can't trust these things blindly, but I feel like most people on the internet already understand that things on the internet you can't trust blindly. And so there's not like, this is not like a big mental shift you have to go through to understand that it …”
Nicholas Carlini Aug 28, 2024 ▶ 10:36 Personal benchmarks vs HumanEval - with Nicholas Carlini of DeepMind
Insight
Carlini: LLMs reduce onboarding to unfamiliar tools from hours to 10 minutes
“It would have taken me. You know, several hours to figure out some things that take 10 minutes if you could just ask exactly the question you want the answer to.”
Nicholas Carlini Aug 28, 2024 ▶ 13:00 Personal benchmarks vs HumanEval - with Nicholas Carlini of DeepMind
Insight
Carlini: AI helper functions preserve programmer mental state on complex problems
“One of the ways we currently don't think about being distracted is you're solving some hard problem and you realize you need a helper function that does X where X is like, it's a known algorithm... Instead of using my mental capacity and solving that problem, …”
Nicholas Carlini Aug 28, 2024 ▶ 20:50 Personal benchmarks vs HumanEval - with Nicholas Carlini of DeepMind
Assertion Not checkable as stated
Carlini: LLMs decompile obscure binaries into readable Python code
“It can turn the compiled source code, which is impossible for any human to understand into the Python code that is entirely reasonable to understand. And, you know, it doesn't run. It has a bunch of problems, but like, it's so much nicer that it's immediately …”
Nicholas Carlini Aug 28, 2024 ▶ 29:52 Personal benchmarks vs HumanEval - with Nicholas Carlini of DeepMind
Insight
Carlini: Anyone claiming 0% or 100% certainty on 5-year AI capabilities is probably wrong
“If you would say there's a zero percent chance that something, you know, the models will get very, very good in the next five years, you're probably wrong. If you're going to say there's a hundred percent chance that in the next five years, some, then you're p…”
Nicholas Carlini Aug 28, 2024 ▶ 31:27 Personal benchmarks vs HumanEval - with Nicholas Carlini of DeepMind
Insight
Carlini: AI security research depends on whether models act autonomously or keep humans in loop
“The way in which security intersects with these things depends a lot in exactly how people use these tools. You know, if it turns out to be the case that these models get to be truly amazing and can solve, you know, tasks completely autonomously, that's a very…”
Nicholas Carlini Aug 28, 2024 ▶ 32:17 Personal benchmarks vs HumanEval - with Nicholas Carlini of DeepMind
Insight
Carlini: If prompt engineering takes longer than manual work, LLMs save no time
“If I have to spend so much time thinking about how I want to frame the question that it would have been faster for me just to get the answer. Didn't save me any time. And so oftentimes, you know, what I do is like, I just dump in whatever current thought that …”
Nicholas Carlini Aug 28, 2024 ▶ 44:11 Personal benchmarks vs HumanEval - with Nicholas Carlini of DeepMind
Disclosure
Carlini never uses few-shot prompting for personal language model queries
“I don't because usually when I want the answer, I just, I want to get the answer.”
Nicholas Carlini Aug 28, 2024 ▶ 45:57 Personal benchmarks vs HumanEval - with Nicholas Carlini of DeepMind
Opinion
Carlini: AI community should do more multi-turn evaluations
“This is the thing that I think many people should be doing more of. I would like more multi-turn evals. I might be writing a paper on this at some point if I get around to it.”
Nicholas Carlini Aug 28, 2024 ▶ 48:27 Personal benchmarks vs HumanEval - with Nicholas Carlini of DeepMind

Show 7statements(7 left)

The other half of the tape: Nicholas Carlini's own voice is left out of every number here. Other people bring the name up 12 times in 4 episodes on Latent Space. every mention, with the transcript →

Who brings them up most Mike Merrill 7John V 2Shawn Wang 1Alex Shaw 1Alessio Fanelli 1

Every mention by year

tap a year for its mentions
008215320242025episodesmentions
02320242025episodes it came up in
0021.54320242025episodesmentions per episode

Appearances (1)

EpisodeDateSpeaking time
Personal benchmarks vs HumanEval - with Nicholas Carlini of DeepMind Aug 28, 2024 48m
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.