Claude 3.5, every mention
108 scenes, the whole family · ← back to Claude 3.5
tap a year for its mentions
every year anyone Shawn Wang 28Alessio Fanelli 15Erik Schluntz 7David Hershey 7Matthew Berman 6Mike Merrill 5Mike Krieger 5Dex Horthy 5Lukas Petersson 4Florent Crivello 4
Verbatim, from the transcripts: the passages where Claude 3.5 comes up
Podcast Crossover: AIE, AGI, frontier lab strategy with @matthew_berman and @swyxtv
- ▶ 4:38 Matthew Berman Well, all right, I want to talk about something new, and you've been, obviously, quite busy this week, but Fable V is now back. 6 times in the scene
- ▶ 6:27 Shawn Wang That's why they're not rolling out Mythos and Fable. 2 times in the scene
- ▶ 19:07 Shawn Wang Like, like, I genuinely do expect Fable to be the end of this era of LLMs because you can't, like, I already told you about the slowness. 4 times in the scene
⚡️Every product of the future will be a living system — Ronak Malde, Trajectory.ai
- ▶ 0:44 Ronak Malde I, I still remember when we first launched Windsurf, it was like, Sonnet 3.5 had just come out and we're playing around with the capabilities and everything.
When AI Agents Run Businesses — Lukas Petersson and Axel Backlund of Andon Labs
- ▶ 15:10 Lukas Petersson Yeah, so what happened, was it Claude, yeah, three, 3.5 sonnet, um, ages ago.
- ▶ 19:39 Lukas Petersson And this was also like Sonnet 3.5, right?
- ▶ 42:53 unnamed speaker You took off Opus four six here, though. 5 times in the scene
- ▶ 52:02 Lukas Petersson Like, it's been Claude, 4.6 Opus, Sonnet 4.6 Mythos, and Opus 4.7.
- ▶ 1:04:50 Lukas Petersson But to be clear, I think one, one thing that is, is important to pin on here, like this was Sonnet 3.5.
Notion’s Sarah Sachs & Simon Last on Custom Agents, Evals, and the Future of Work
- ▶ 4:24 Simon Last Probably like Sonic 3.6 or seven, uh, early last year.
- ▶ 23:42 Alessio Fanelli They were just 3.5. 2 times in the scene
The Stove Guy: Sam D'Amico Shows New AI Cooking Features on America's Most Powerful Stove at Impulse
- ▶ 19:19 Sam D'Amico So like I was, I, I, when Sonnet three, five came out, I was like, okay, this is like, we're actually, this is good enough to start.
Why Your AI Agents Don’t Work with Dex Horthy of HumanLayer | In-Context Cooking
- ▶ 13:41 Dex Horthy When Opus 4.5 came out, they're like, oh, 5 times in the scene
⚡️ Polsia: Solo Founder Tiny Team from 0 to 1m ARR in 1 month & the future of Self-Running Companies
- ▶ 32:40 unnamed speaker And so it's, it's Opus 4.6 extra thinking.
Measuring Exponential Trends Rising (in AI) — Joel Becker, METR
- ▶ 16:51 Alessio Fanelli If you were to redo it with Opus 4.5, would you expect the results to be dramatically different?
Dylan Patel Explains the AI War While Cooking | In-Context Cooking
- ▶ 21:03 Dylan Patel Um, you know, obviously Claude 4.5 came out, and Claude Code came out, ah, last year.
Claude Code for Finance + The Global Memory Shortage: Doug O'Laughlin, SemiAnalysis
- ▶ 24:48 Doug O'Laughlin Uh, Opus 4.5 did.
Inside AI’s $10B+ Capital Flywheel — Martin Casado & Sarah Wang of a16z
- ▶ 35:17 Martin Casado But it seems to me that, like, listen, Codex, in my experience, is for sure better than Opus 4.5 for coding. 2 times in the scene
[State of RL/Reasoning] IMO/IOI Gold, OpenAI o3/GPT-5, and Cursor Composer — Ashvin Nair, Cursor
- ▶ 32:08 Ashvin Nair The anthropic, uh, models, like the, uh, Opus two, 4.5, it has this kind of like, uh, there's this like RKGI two plot that looks exactly like the API ones, right?
Anthropic, Glean & OpenRouter: How AI Moats Are Built with Deedy Das of Menlo Ventures
- ▶ 29:20 Shawn Wang I mean, uh, Opus 4.1, Sonnet 4.5, Haiku 4.5, all released in the last few months.
- ▶ 29:20 Shawn Wang I mean, uh, Opus 4.1, Sonnet 4.5, Haiku 4.5, all released in the last few months.
- ▶ 29:20 Shawn Wang I mean, uh, Opus 4.1, Sonnet 4.5, Haiku 4.5, all released in the last few months.
Terminal-Bench 2.0: the most impt coding agent benchmark of 2025 gets a v2! Launch + Q&A w/ founders
- ▶ 1:56 Mike Merrill So only three days after we announced it, it made it onto the Cloud Four model card to set a new state of the art. 2 times in the scene
How Zyphra went all-in on AMD + Why Devs feel faster with AI but are slower — with Quentin Anthony
- ▶ 37:27 Quentin Anthony So Claude, 3.5 sonnet kind of, um, really, it finally took all of the low level thinking of unit tests and skeletons and documentation, all that stuff.
Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limits
- ▶ 3:58 Mike Merrill I think, I think one of the really key moments for us was getting onto the Claude IV model card.
- ▶ 8:45 Alex Shaw You know, GPT-Five or Claude Sonnet decides to write, but, but somehow they work.
- ▶ 23:18 Mike Merrill So if your agent does really well on terminal bench and you're using cloud four or five sonnet, that's great. 2 times in the scene
- ▶ 32:28 unnamed speaker You know, like today the Sonic 4.5 release, it was like, oh, it ran for like 30 hours.
⚡️Claude Sonnet 4.5 and Anthropic's roadmap for Agents and Developers — Mike Krieger, Anthropic
- ▶ 1:15 Mike Krieger We have more traffic on Sonnet 4.5 than we had on Sonnet four.
- ▶ 3:36 Mike Krieger I have three, uh, and, and the, the team got really tired of me by the end of this, uh, Sonnet four foot five process. 2 times in the scene
- ▶ 7:26 Mike Krieger There was definitely the, like, white background, purple tint on the top of rounded rectangle or something like Claude Forrest on it. 2 times in the scene
- ▶ 10:58 Alessio Fanelli For example, the Cloud AI example that you tweeted, the final one with Sonnet 4.5, it has the chat history right on the side, versus on the actual Cloud AI, you have to click through to go to the history.
- ▶ 16:08 Alessio Fanelli So if I look at the Sonnet 4.5 benchmarks, you have retail airline and telecom at the three agentic tool use top bench, uh, categories that you mentioned.
- ▶ 18:33 unnamed speaker Uh, but 4.5 was so good that, like, they got all the engineers excited and, uh, you know, it's basically fast forwarded a rewrite that was already sort of, kind of in the works. 2 times in the scene
Amp: The Emperor Has No Clothes
- ▶ 0:59 Thorsten Ball I mean, I'll start, you can jump in, but basically I came back to SoftSquare February, and then this was when cloud three five, three seven happened too.
- ▶ 23:27 Alessio Fanelli The cost of like Sonnet four versus Sonnet 3.5, it's kind of like minimal compared to like a 152 103 hundred K once you do taxes and benefits and all of that, that you pay to employees.
- ▶ 27:39 Quinn Slack So first, it took people eight or nine months to figure out what Three Five Sonnet was capable of.
- ▶ 29:57 Alessio Fanelli Like when cursors switch from Sonnet to GPT-V as like the default model that was like, you know, 2 times in the scene
Context Engineering for Agents - Lance Martin, LangChain
- ▶ 55:08 Lance Martin Cloud three, five hits, and then boom, it kind of unlocks the product.
A Technical History of Generative Media
- ▶ 28:45 Batuhan Taskaya Uh, well, one thing that I, I, I keep thinking about this is, like, is this true for LOMs, though, you know, like, you, like, would you, like, I, I always default to cloud 4.1 opus, right?
⚡️Launching Ona: Coding Agent with Fully Sandboxed Cloud Environment
- ▶ 5:26 Johannes (Johannes Landgraf) So when the 3.5, so Sonnet 3.5 came out, it became clear that it will not be very far away.
⚡️OpenCode: Claude Code but Open Source, with Any Model, and frontier TUI - with Dax Reed (@thdxr)
⚡️Mercury: Ultra-Fast Diffusion LLMs — Estefano Ermon, CEO Inception Labs
- ▶ 17:03 Stefano Ermon So it's, it's comparable to GPT, 4.1 nano, cloud haiku, kind of like
⚡️Composio: 10,000+ tools that evolve for Agents — Karan Vaidya and Soham Ganatra
- ▶ 20:51 Karan Vaidya So kind of like, I think I really love Claude Sonet, for example, because of the same, and Gropfor had like some of a mix of Opus and Sonet, which I really liked. 2 times in the scene
🕰️ The Oral History of Windsurf (ft. Varun Mohan, Scott Wu, Jeff Wang, Kevin Hou, Anshul R)
Cline: The Collaborative AI Coder
- ▶ 5:38 Saoud (Saud) Rizwan Yeah, when I first started working on Klein, this was, I think, 10 days after Cloud through Five Sonic came out. 3 times in the scene
- ▶ 45:04 unnamed speaker Like in our own internal benchmarks, Claude Sonnet four recently hit a sub five percent or like around actually four percent diff edit failure rate.
⚡️Ranking Agentic LLMs — Pratik Bhavsar, Galileo
- ▶ 10:48 Pratik Bhavsar I would say that when we started the leaderboard, Claude 3.5 was somewhere in the middle. 2 times in the scene
⚡️Anthropic vs Cognition on Multi-Agents: A Breakdown with Dylan Davis
- ▶ 7:19 Dylan Davis So we're going to have the large model at the beginning, so it's going to be a big, beefy model, like O three, Gemini two dot five pro, Cloud Sonnet Opus, Sonnet, or yeah, Cloud Sonnet four, Cloud Sonnet four Opus, et cetera.
⚡️Warp 2.0: the Agentic Development Environment - Zach Lloyd and Ben Holmes
- ▶ 21:44 Zach Lloyd It's probably better than Windsurf since they aren't, they don't have Sonnet for.
⚡️Launching AI Diplomacy: the hardest LLM Game Benchmark yet - Alex Duffy
- ▶ 16:19 Alex Duffy I think one of the things that we'll do next is have three O four minis versus cloud two five pros and allow them to, you know, create an alliance, essentially like, you know, change their system prompts so that they, they know if any of…
The Utility of Interpretability — Emmanuel Amiesen
- ▶ 3:14 Emmanuel Ameisen So to give you an example, like one of the things in the paper is this sort of like multi-hop reasoning where, you know, we ask, you know, like, uh, cloud three, four, five haiku, like, oh, the capital of the state where Dallas is, is…
- ▶ 45:04 unnamed speaker And then the other thing is, um, can you just, like, up, write good code, don't write bad code, and make Sondra 3.5? 2 times in the scene
- ▶ 1:18:15 unnamed speaker So in these random examples here where, where, like, you have this poem, the silver moon cast a gentle light, and then Claude III.V haiku would, like, rhyme with illuminating the peaceful night.
The AI Coding Factory
- ▶ 23:13 Matan Grinberg And then we, let's say when we upgraded from Sonnet, 3.5 to 3.7, we suddenly had a lot of developers being like, Hey, wait, it now does this less, or it does this more what's happening.
⚡️Multi-Turn RL for Multi-Hour Agents — with Will Brown, Prime Intellect
- ▶ 3:31 Will Brown Sometimes even like cloud 3.5 would do like a tiny bit of thinking, and it was really just like deciding which tool to use for the most part.
Claude Code: Anthropic's CLI Agent
- ▶ 56:34 Shawn Wang Uh, he actually misses the common sense of 3.5 because 3.7 being so persistent, 3.5 actually had some entertaining stories where apparently it like gave up on tasks and just 2.7 doesn't. 2 times in the scene
- ▶ 59:48 Alessio Fanelli I know that 3.5 haiku was the number four model on Adr when it came out.
Zed Agents — with Zed Cofounders Nathan Sobo & Antonio Scandurra
- ▶ 13:07 Antonio Scandurra And now they've trained it so that like, yeah, just cloud 3.5 sonnet has no problem just going, you know, you know, loop over and over, you know, call this tool and then come back, call this other tool like that.
⚡️Factorio Learning Environment: the ultimate Game Agent Eval — Jack Hopkins
- ▶ 16:27 Jack Hopkins Claude, for instance, the Sonnet 3.5 was very much fire and forget.
GPT 4.1: The New OpenAI Workhorse
- ▶ 23:12 Shawn Wang There's been criticisms of Claude Sonnet trying to rewrite too many files at once when I just wanted to make one thing.
Claude Plays Pokémon Hackathon: Escape from Mt. Moon!
- ▶ 1:13 David Hershey So June when Sonnet, or, uh, 3.5 Sonnet came out, uh, I just kind of, like, wanted to build agents. 3 times in the scene
- ▶ 16:33 David Hershey Um, and those were comparably performant on, like, 3.5 Sonnet back then. 2 times in the scene
The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
- ▶ 4:03 Shawn Wang Like, you know, you know, Claude 3.5 Opus, like if it, if it does exist, still not like, you know, the, the thing that we actually use is Sonnet, right?
- ▶ 4:03 Shawn Wang Like, you know, you know, Claude 3.5 Opus, like if it, if it does exist, still not like, you know, the, the thing that we actually use is Sonnet, right?
Fullstack-Bench: The Eval for Coding Agents — with Sujay Jayakar, Chief Scientist, Convex
- ▶ 3:21 Sujay Jayakar And we don't have, we're not gonna have to watch through all of this, but it's pretty remarkable to just see that when it has the right feedback in cursor composer, and this was even on cloud three, five, it can just autonomously code for…
- ▶ 17:23 Sujay Jayakar Like, for example, we just tried clod three seven and it performs worse than clod three five on convex evals with the same prompting. 2 times in the scene
How Claude Plays Pokémon was made
- ▶ 2:32 David Hershey This was, like, Sonnet III.V came out in June of last year, which is when I started, kicked it around. 2 times in the scene
- ▶ 22:05 unnamed speaker I'm curious, um, as you switched from 3.5 to 3.7 and sort of reasoning models, were there any degradations there?
Gemini 2.0 Flash and Flash Thinking: the new SOTA models for the agentic era
smol agents are all you need
- ▶ 14:10 unnamed speaker Is it just because of like O-one is better than Sonnet or, um, yeah. 2 times in the scene
Beating OpenAI and Anthropic by Looking At Data: the new #1 on SWE-Bench w/ W&B CTO Shawn Lewis
- ▶ 15:23 Shawn Lewis It's all done for first principles, and I really would like to, I, I didn't get a chance to actually run, um, Sonnet through this all the way, so I don't know what the result would be if I just dropped Sonnet in here. 3 times in the scene
DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
- ▶ 15:11 unnamed speaker Like the, the failures of 3.5 Opus, the failures of GPC five, like, you know, it's a, it's a thing that people are sort of rumoring.
OpenAI o1 isn’t a chat model (and that’s the point)
- ▶ 6:32 Dan McAteer I'm typically using LLMs, mostly for, for coding use cases, and, like, using Sonnet, 3.5, and GPT four out.
- ▶ 16:58 Alessio Fanelli Or, uh, I mean, for coding, I use Sonnet 3.6, unofficial name. 2 times in the scene
- ▶ 20:45 Ben Hillock I think for some amount of time, you know, when three, 3.5 sonnet is like both pretty fast and pretty intelligent, and I think I could do a lot of things in one, one sort of prompt.
- ▶ 26:42 Alessio Fanelli So 3.6 is, is pretty good, but maybe it's just. 2 times in the scene
2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents
- ▶ 37:15 Shawn Wang That Claude Sonnet so far is beating O-one on coding tasks without, uh,
- ▶ 40:13 Shawn Wang As far as everyone is concerned, Claude, you know, Opus 3.5 is not coming out.
- ▶ 1:19:19 Shawn Wang Um, in November, we had 3.5 Haiku Niu. 2 times in the scene
- ▶ 1:19:23 Shawn Wang Um, and obviously we had Sonnet as well, uh, Sonnet as, uh, as not, I don't know where there's Sonnet on this chart, but, um, Haiku New, uh, basically, uh, was four X the price of old Haiku, or the, sorry, 3.5 Haiku was four X the price of…
- ▶ 1:25:27 Shawn Wang And obviously now we know that Sonnet is, is kind of the workhorse, um, just like four O is the workhorse of, of OpenAI.
- ▶ 1:34:07 Shawn Wang Immediately said it was shit because I'm still using Sonnet or whatever, but like still very good.
0 to over $8M ARR in 2 months as a Claude Wrapper (Bolt.new, Qodo)
- ▶ 40:14 Itamar Friedman Although there is Gemini or even Sonet, I think is available on GCP, just an example. 2 times in the scene
- ▶ 1:00:48 Eric Simons And so I think there's, there's been an incredible amount of improvement to the product, to the agent, also to like the underlying models too, like Sonnet, uh, you know, they just happened to do an update on, you know, uh, with their, with… 2 times in the scene
The new Claude 3.5 Sonnet, Computer Use, and Building SOTA Agents — with Erik Schluntz, Anthropic
- ▶ 3:43 Shawn Wang I feel like there's just been a series of releases related with cloud 3.5 sonnet around about two, three months ago, 3.5 sonnet came out. 2 times in the scene
- ▶ 19:38 Erik Schluntz I think especially, uh, the new Sonnet 3.5 is very, very good at self-correction. 3 times in the scene
- ▶ 48:02 Erik Schluntz And then, uh, you know, Sonnet reads just those, and you save 4 times in the scene
- ▶ 48:33 Shawn Wang It turns out, I think you did do 3.5 haiku with your tools and it scored a 40.6. 2 times in the scene
Why Compound AI + Open Source will beat Closed AI — with Lin Qiao, CEO of Fireworks AI
- ▶ 37:53 Shawn Wang So much so that they, they made Claude's on it, and for, oh, just, like, they, they made the previous state of the art look bad.
- ▶ 42:45 Shawn Wang Like I, I actually use a lot of Sonnet. 2 times in the scene
Agents @ Work: Lindy.ai (with live demo!)
- ▶ 32:04 Florent Crivello 3.5 sonnet. 3 times in the scene
- ▶ 35:01 Florent Crivello It's, I'm seeing some tweets that say that the new 3.5 sonnet is as good as O-one, but with none of all the crazy, uh, it beats O-one on some measures.
Agents @ Work: Dust.tt — with Stanislas Polu
- ▶ 42:17 Stanislas Polu And Clouds, 3.5 Sonnet is great as well.
Building the Silicon Brain - Drew Houston of Dropbox
- ▶ 12:53 Drew Houston I mean, Sonnet three five is probably the best all around, but then these things are like pretty limited if you don't give them the right context.
[Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz
- ▶ 4:23 Vibhu Sapra They do augment it a little bit, but point being, they can get better than Gemini, 1.5, better than Claude, 3.5 sonnet, better than GPT for V at a much smaller size with about a million samples of data, which is very impressive, right?
- ▶ 30:52 Vibhu Sapra but yeah, I mean, across the board, the seven B, the one B, the 72 B very on par with four O four V 1.5 thoughts on it and whatnot. 2 times in the scene
The Ultimate Guide to Prompting - with Sander Schulhoff from LearnPrompting.org
- ▶ 32:34 Alessio Fanelli I found when I use Sonnet, a lot of times it does Chain of Thought on its own without having to ask to think step by step.
The Winds of AI Winter (Q2 Four Wars of the AI Stack Recap)
- ▶ 4:59 unnamed speaker with 3.5 on it, at least in like the, some of the hard, 4 times in the scene
page 1 of 2 · 100 scenes per page · newest episode first next →