Claude 3.5 Sonnet, every mention
89 scenes · ← back to Claude 3.5 Sonnet
tap a year for its mentions
every year anyone Shawn Wang 18Alessio Fanelli 14Erik Schluntz 7David Hershey 7Matthew Berman 6Mike Merrill 5Mike Krieger 5Lukas Petersson 4Florent Crivello 4Vibhu Sapra 3
Verbatim, from the transcripts: the passages where Claude 3.5 Sonnet comes up
Podcast Crossover: AIE, AGI, frontier lab strategy with @matthew_berman and @swyxtv
- ▶ 4:38 Matthew Berman Well, all right, I want to talk about something new, and you've been, obviously, quite busy this week, but Fable V is now back. 6 times in the scene
- ▶ 19:07 Shawn Wang Like, like, I genuinely do expect Fable to be the end of this era of LLMs because you can't, like, I already told you about the slowness. 4 times in the scene
⚡️Every product of the future will be a living system — Ronak Malde, Trajectory.ai
- ▶ 0:44 Ronak Malde I, I still remember when we first launched Windsurf, it was like, Sonnet 3.5 had just come out and we're playing around with the capabilities and everything.
When AI Agents Run Businesses — Lukas Petersson and Axel Backlund of Andon Labs
- ▶ 15:10 Lukas Petersson Yeah, so what happened, was it Claude, yeah, three, 3.5 sonnet, um, ages ago.
- ▶ 19:39 Lukas Petersson And this was also like Sonnet 3.5, right?
- ▶ 52:02 Lukas Petersson Like, it's been Claude, 4.6 Opus, Sonnet 4.6 Mythos, and Opus 4.7.
- ▶ 1:04:50 Lukas Petersson But to be clear, I think one, one thing that is, is important to pin on here, like this was Sonnet 3.5.
Notion’s Sarah Sachs & Simon Last on Custom Agents, Evals, and the Future of Work
- ▶ 4:24 Simon Last Probably like Sonic 3.6 or seven, uh, early last year.
- ▶ 23:42 Alessio Fanelli They were just 3.5. 2 times in the scene
The Stove Guy: Sam D'Amico Shows New AI Cooking Features on America's Most Powerful Stove at Impulse
- ▶ 19:19 Sam D'Amico So like I was, I, I, when Sonnet three, five came out, I was like, okay, this is like, we're actually, this is good enough to start.
⚡️ Polsia: Solo Founder Tiny Team from 0 to 1m ARR in 1 month & the future of Self-Running Companies
- ▶ 32:40 unnamed speaker And so it's, it's Opus 4.6 extra thinking.
Measuring Exponential Trends Rising (in AI) — Joel Becker, METR
- ▶ 16:51 Alessio Fanelli If you were to redo it with Opus 4.5, would you expect the results to be dramatically different?
Dylan Patel Explains the AI War While Cooking | In-Context Cooking
- ▶ 21:03 Dylan Patel Um, you know, obviously Claude 4.5 came out, and Claude Code came out, ah, last year.
Inside AI’s $10B+ Capital Flywheel — Martin Casado & Sarah Wang of a16z
- ▶ 35:17 Martin Casado But it seems to me that, like, listen, Codex, in my experience, is for sure better than Opus 4.5 for coding. 2 times in the scene
Anthropic, Glean & OpenRouter: How AI Moats Are Built with Deedy Das of Menlo Ventures
- ▶ 29:20 Shawn Wang I mean, uh, Opus 4.1, Sonnet 4.5, Haiku 4.5, all released in the last few months.
Terminal-Bench 2.0: the most impt coding agent benchmark of 2025 gets a v2! Launch + Q&A w/ founders
- ▶ 1:56 Mike Merrill So only three days after we announced it, it made it onto the Cloud Four model card to set a new state of the art. 2 times in the scene
How Zyphra went all-in on AMD + Why Devs feel faster with AI but are slower — with Quentin Anthony
- ▶ 37:27 Quentin Anthony So Claude, 3.5 sonnet kind of, um, really, it finally took all of the low level thinking of unit tests and skeletons and documentation, all that stuff.
Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limits
- ▶ 3:58 Mike Merrill I think, I think one of the really key moments for us was getting onto the Claude IV model card.
- ▶ 8:45 Alex Shaw You know, GPT-Five or Claude Sonnet decides to write, but, but somehow they work.
- ▶ 23:18 Mike Merrill So if your agent does really well on terminal bench and you're using cloud four or five sonnet, that's great. 2 times in the scene
- ▶ 32:28 unnamed speaker You know, like today the Sonic 4.5 release, it was like, oh, it ran for like 30 hours.
⚡️Claude Sonnet 4.5 and Anthropic's roadmap for Agents and Developers — Mike Krieger, Anthropic
- ▶ 1:15 Mike Krieger We have more traffic on Sonnet 4.5 than we had on Sonnet four.
- ▶ 3:36 Mike Krieger I have three, uh, and, and the, the team got really tired of me by the end of this, uh, Sonnet four foot five process. 2 times in the scene
- ▶ 7:26 Mike Krieger There was definitely the, like, white background, purple tint on the top of rounded rectangle or something like Claude Forrest on it. 2 times in the scene
- ▶ 10:58 Alessio Fanelli For example, the Cloud AI example that you tweeted, the final one with Sonnet 4.5, it has the chat history right on the side, versus on the actual Cloud AI, you have to click through to go to the history.
- ▶ 16:08 Alessio Fanelli So if I look at the Sonnet 4.5 benchmarks, you have retail airline and telecom at the three agentic tool use top bench, uh, categories that you mentioned.
- ▶ 18:33 unnamed speaker Uh, but 4.5 was so good that, like, they got all the engineers excited and, uh, you know, it's basically fast forwarded a rewrite that was already sort of, kind of in the works. 2 times in the scene
Amp: The Emperor Has No Clothes
- ▶ 0:59 Thorsten Ball I mean, I'll start, you can jump in, but basically I came back to SoftSquare February, and then this was when cloud three five, three seven happened too.
- ▶ 23:27 Alessio Fanelli The cost of like Sonnet four versus Sonnet 3.5, it's kind of like minimal compared to like a 152 103 hundred K once you do taxes and benefits and all of that, that you pay to employees.
- ▶ 27:39 Quinn Slack So first, it took people eight or nine months to figure out what Three Five Sonnet was capable of.
- ▶ 29:57 Alessio Fanelli Like when cursors switch from Sonnet to GPT-V as like the default model that was like, you know, 2 times in the scene
Context Engineering for Agents - Lance Martin, LangChain
- ▶ 55:08 Lance Martin Cloud three, five hits, and then boom, it kind of unlocks the product.
A Technical History of Generative Media
- ▶ 28:45 Batuhan Taskaya Uh, well, one thing that I, I, I keep thinking about this is, like, is this true for LOMs, though, you know, like, you, like, would you, like, I, I always default to cloud 4.1 opus, right?
⚡️Launching Ona: Coding Agent with Fully Sandboxed Cloud Environment
- ▶ 5:26 Johannes (Johannes Landgraf) So when the 3.5, so Sonnet 3.5 came out, it became clear that it will not be very far away.
⚡️OpenCode: Claude Code but Open Source, with Any Model, and frontier TUI - with Dax Reed (@thdxr)
⚡️Composio: 10,000+ tools that evolve for Agents — Karan Vaidya and Soham Ganatra
- ▶ 20:51 Karan Vaidya So kind of like, I think I really love Claude Sonet, for example, because of the same, and Gropfor had like some of a mix of Opus and Sonet, which I really liked. 2 times in the scene
🕰️ The Oral History of Windsurf (ft. Varun Mohan, Scott Wu, Jeff Wang, Kevin Hou, Anshul R)
Cline: The Collaborative AI Coder
- ▶ 5:38 Saoud (Saud) Rizwan Yeah, when I first started working on Klein, this was, I think, 10 days after Cloud through Five Sonic came out. 3 times in the scene
- ▶ 45:04 unnamed speaker Like in our own internal benchmarks, Claude Sonnet four recently hit a sub five percent or like around actually four percent diff edit failure rate.
⚡️Anthropic vs Cognition on Multi-Agents: A Breakdown with Dylan Davis
- ▶ 7:19 Dylan Davis So we're going to have the large model at the beginning, so it's going to be a big, beefy model, like O three, Gemini two dot five pro, Cloud Sonnet Opus, Sonnet, or yeah, Cloud Sonnet four, Cloud Sonnet four Opus, et cetera.
⚡️Warp 2.0: the Agentic Development Environment - Zach Lloyd and Ben Holmes
- ▶ 21:44 Zach Lloyd It's probably better than Windsurf since they aren't, they don't have Sonnet for.
⚡️Launching AI Diplomacy: the hardest LLM Game Benchmark yet - Alex Duffy
- ▶ 16:19 Alex Duffy I think one of the things that we'll do next is have three O four minis versus cloud two five pros and allow them to, you know, create an alliance, essentially like, you know, change their system prompts so that they, they know if any of…
The Utility of Interpretability — Emmanuel Amiesen
- ▶ 45:04 unnamed speaker And then the other thing is, um, can you just, like, up, write good code, don't write bad code, and make Sondra 3.5? 2 times in the scene
The AI Coding Factory
- ▶ 23:13 Matan Grinberg And then we, let's say when we upgraded from Sonnet, 3.5 to 3.7, we suddenly had a lot of developers being like, Hey, wait, it now does this less, or it does this more what's happening.
Claude Code: Anthropic's CLI Agent
- ▶ 56:34 Shawn Wang Uh, he actually misses the common sense of 3.5 because 3.7 being so persistent, 3.5 actually had some entertaining stories where apparently it like gave up on tasks and just 2.7 doesn't. 2 times in the scene
Zed Agents — with Zed Cofounders Nathan Sobo & Antonio Scandurra
- ▶ 13:07 Antonio Scandurra And now they've trained it so that like, yeah, just cloud 3.5 sonnet has no problem just going, you know, you know, loop over and over, you know, call this tool and then come back, call this other tool like that.
⚡️Factorio Learning Environment: the ultimate Game Agent Eval — Jack Hopkins
- ▶ 16:27 Jack Hopkins Claude, for instance, the Sonnet 3.5 was very much fire and forget.
GPT 4.1: The New OpenAI Workhorse
- ▶ 23:12 Shawn Wang There's been criticisms of Claude Sonnet trying to rewrite too many files at once when I just wanted to make one thing.
Claude Plays Pokémon Hackathon: Escape from Mt. Moon!
- ▶ 1:13 David Hershey So June when Sonnet, or, uh, 3.5 Sonnet came out, uh, I just kind of, like, wanted to build agents. 3 times in the scene
- ▶ 16:33 David Hershey Um, and those were comparably performant on, like, 3.5 Sonnet back then. 2 times in the scene
The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
- ▶ 4:03 Shawn Wang Like, you know, you know, Claude 3.5 Opus, like if it, if it does exist, still not like, you know, the, the thing that we actually use is Sonnet, right?
Fullstack-Bench: The Eval for Coding Agents — with Sujay Jayakar, Chief Scientist, Convex
- ▶ 3:21 Sujay Jayakar And we don't have, we're not gonna have to watch through all of this, but it's pretty remarkable to just see that when it has the right feedback in cursor composer, and this was even on cloud three, five, it can just autonomously code for…
- ▶ 17:23 Sujay Jayakar Like, for example, we just tried clod three seven and it performs worse than clod three five on convex evals with the same prompting. 2 times in the scene
How Claude Plays Pokémon was made
- ▶ 2:32 David Hershey This was, like, Sonnet III.V came out in June of last year, which is when I started, kicked it around. 2 times in the scene
- ▶ 22:05 unnamed speaker I'm curious, um, as you switched from 3.5 to 3.7 and sort of reasoning models, were there any degradations there?
Gemini 2.0 Flash and Flash Thinking: the new SOTA models for the agentic era
- ▶ 4:29 unnamed speaker Like, you know, the rumor is that Anthropic has clawed in 3.5 Opus or, and it's distilling for, for Sonnet, right? 2 times in the scene
smol agents are all you need
- ▶ 14:10 unnamed speaker Is it just because of like O-one is better than Sonnet or, um, yeah. 2 times in the scene
Beating OpenAI and Anthropic by Looking At Data: the new #1 on SWE-Bench w/ W&B CTO Shawn Lewis
- ▶ 15:23 Shawn Lewis It's all done for first principles, and I really would like to, I, I didn't get a chance to actually run, um, Sonnet through this all the way, so I don't know what the result would be if I just dropped Sonnet in here. 3 times in the scene
OpenAI o1 isn’t a chat model (and that’s the point)
- ▶ 6:32 Dan McAteer I'm typically using LLMs, mostly for, for coding use cases, and, like, using Sonnet, 3.5, and GPT four out.
- ▶ 16:58 Alessio Fanelli Or, uh, I mean, for coding, I use Sonnet 3.6, unofficial name. 2 times in the scene
- ▶ 20:45 Ben Hillock I think for some amount of time, you know, when three, 3.5 sonnet is like both pretty fast and pretty intelligent, and I think I could do a lot of things in one, one sort of prompt.
- ▶ 26:42 Alessio Fanelli So 3.6 is, is pretty good, but maybe it's just. 2 times in the scene
2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents
- ▶ 37:15 Shawn Wang That Claude Sonnet so far is beating O-one on coding tasks without, uh,
- ▶ 1:19:23 Shawn Wang Um, and obviously we had Sonnet as well, uh, Sonnet as, uh, as not, I don't know where there's Sonnet on this chart, but, um, Haiku New, uh, basically, uh, was four X the price of old Haiku, or the, sorry, 3.5 Haiku was four X the price of…
- ▶ 1:25:27 Shawn Wang And obviously now we know that Sonnet is, is kind of the workhorse, um, just like four O is the workhorse of, of OpenAI.
- ▶ 1:34:07 Shawn Wang Immediately said it was shit because I'm still using Sonnet or whatever, but like still very good.
0 to over $8M ARR in 2 months as a Claude Wrapper (Bolt.new, Qodo)
- ▶ 40:14 Itamar Friedman Although there is Gemini or even Sonet, I think is available on GCP, just an example. 2 times in the scene
- ▶ 1:00:48 Eric Simons And so I think there's, there's been an incredible amount of improvement to the product, to the agent, also to like the underlying models too, like Sonnet, uh, you know, they just happened to do an update on, you know, uh, with their, with… 2 times in the scene
The new Claude 3.5 Sonnet, Computer Use, and Building SOTA Agents — with Erik Schluntz, Anthropic
- ▶ 3:43 Shawn Wang I feel like there's just been a series of releases related with cloud 3.5 sonnet around about two, three months ago, 3.5 sonnet came out. 2 times in the scene
- ▶ 19:38 Erik Schluntz I think especially, uh, the new Sonnet 3.5 is very, very good at self-correction. 3 times in the scene
- ▶ 48:02 Erik Schluntz And then, uh, you know, Sonnet reads just those, and you save 4 times in the scene
Why Compound AI + Open Source will beat Closed AI — with Lin Qiao, CEO of Fireworks AI
- ▶ 37:53 Shawn Wang So much so that they, they made Claude's on it, and for, oh, just, like, they, they made the previous state of the art look bad.
- ▶ 42:45 Shawn Wang Like I, I actually use a lot of Sonnet. 2 times in the scene
Agents @ Work: Lindy.ai (with live demo!)
- ▶ 32:04 Florent Crivello 3.5 sonnet. 3 times in the scene
- ▶ 35:01 Florent Crivello It's, I'm seeing some tweets that say that the new 3.5 sonnet is as good as O-one, but with none of all the crazy, uh, it beats O-one on some measures.
Agents @ Work: Dust.tt — with Stanislas Polu
- ▶ 42:17 Stanislas Polu And Clouds, 3.5 Sonnet is great as well.
Building the Silicon Brain - Drew Houston of Dropbox
- ▶ 12:53 Drew Houston I mean, Sonnet three five is probably the best all around, but then these things are like pretty limited if you don't give them the right context.
[Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz
- ▶ 4:23 Vibhu Sapra They do augment it a little bit, but point being, they can get better than Gemini, 1.5, better than Claude, 3.5 sonnet, better than GPT for V at a much smaller size with about a million samples of data, which is very impressive, right?
- ▶ 30:52 Vibhu Sapra but yeah, I mean, across the board, the seven B, the one B, the 72 B very on par with four O four V 1.5 thoughts on it and whatnot. 2 times in the scene
The Ultimate Guide to Prompting - with Sander Schulhoff from LearnPrompting.org
- ▶ 32:34 Alessio Fanelli I found when I use Sonnet, a lot of times it does Chain of Thought on its own without having to ask to think step by step.
The Winds of AI Winter (Q2 Four Wars of the AI Stack Recap)
- ▶ 4:59 unnamed speaker with 3.5 on it, at least in like the, some of the hard, 4 times in the scene
- ▶ 28:56 Alessio Fanelli You know, right now, for the large models, we're still pretty aware of like, oh, is this Sonic Troop M five?
- ▶ 43:01 unnamed speaker And then Claude also did it with Opus and then with three fives on it, right?
- ▶ 1:12:10 unnamed speaker That's definitely one of the emerging stories of the year that has happened is efficiency matters for, for all, for all mini and three fives on it in a way that in January we would, nobody was talking about.
[LLM Paper Club] Llama 3.1 Paper: The Llama Family of Models
Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
- ▶ 37:51 Thomas Scialom But, by far, compared to the version originally released, uh, even now, I think there's maybe the last cloud Sonya FF-V and GPT-IV-Zero that are performing it.
How To Hire AI Engineers (ft. James Brady and Adam Wiggins of Elicit)
- ▶ 34:27 unnamed speaker Uh, and I was wondering, you know, maybe this would be a good, good exercise is how do people have curiosity, enthusiasm for capabilities and language models when, for example, the research paper for cloud 3.5 is four pages.