Insight certainty 4/5 debate potential 2/5

Carlini: AI helper functions preserve programmer mental state on complex problems

Nicholas Carlini · Personal benchmarks vs HumanEval - with Nicholas Carlini of DeepMind · Aug 28, 2024 · at 20:50

Nicholas Carlini of DeepMind discusses the underrated productivity benefits of using LLMs for mundane coding subtasks.

0:00 / 0:45
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“One of the ways we currently don't think about being distracted is you're solving some hard problem and you realize you need a helper function that does X where X is like, it's a known algorithm... Instead of using my mental capacity and solving that problem, and then coming back to the problem I was originally trying to solve, you could just ask model, please solve this problem for me. It gives you the answer. You run it. You can check that it works very, very quickly. And now you go back to solving the problem without having lost all the mental state.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Nicholas Carlini

Insight
Carlini: If LLMs always give desired answers, questions aren't hard enough
“When you're using these models, if you're getting the answer you want, always, it means you're not asking them hard enough questions.”
Nicholas Carlini Aug 28, 2024 ▶ 15:30 Personal benchmarks vs HumanEval - with Nicholas Carlini of DeepMind
Insight
Carlini: 90% of scientific research is routine work that AI can automate
“90% of this is not doing something new. Like, 90% of this is like doing things a million people have done before, and then a little bit of something that was new. There's a reason why we say we stand on the shoulders of giants. It's true. Almost everything tha…”
Nicholas Carlini Aug 28, 2024 ▶ 19:50 Personal benchmarks vs HumanEval - with Nicholas Carlini of DeepMind
Insight
Carlini: Users should build personalized AI benchmarks instead of relying on public leaderboards
“The argument that I tried to lay out in this post is that more people should make benchmarks that are tailored to them.”
Nicholas Carlini Aug 28, 2024 ▶ 39:10 Personal benchmarks vs HumanEval - with Nicholas Carlini of DeepMind
Insight
Carlini: GPT-4 would exist identically without adversarial machine learning research
“Nothing about GPT-IV would be at all different if the field of, like the entire field of Everson machine learning disappeared. Like everything to do with Everson examples, like all of the, like for the most part, like GPT-IV would exist identically.”
Nicholas Carlini Aug 28, 2024 ▶ 58:56 Personal benchmarks vs HumanEval - with Nicholas Carlini of DeepMind
Opinion
Carlini: Most AI commentators spin arguments based on ideology rather than reality
“I feel like most people who write about language models being good or bad, some underlying message of like, you know, they have their camp and their camp is like, AI is bad or AI is good or whatever. And they like, they spin whatever they're gonna say accordin…”
Nicholas Carlini Aug 28, 2024 ▶ 5:40 Personal benchmarks vs HumanEval - with Nicholas Carlini of DeepMind
Insight
Carlini: Copying and pasting error messages effectively creates a coding agent
“Currently though, make a model into an agent by just copying and pasting error messages for the most part. And that's what I do is, you know, you run it and it gives you some code that doesn't work and either I'll fix the code or it will give me buggy code and…”
Nicholas Carlini Aug 28, 2024 ▶ 10:17 Personal benchmarks vs HumanEval - with Nicholas Carlini of DeepMind
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.