Claude 3.5 Sonnet

part of Claude 3.5

9 statements across 7 episodes · 4 bullish · 0 bearish · 7 people on the record · first statement Jul 23, 2024 by Thomas Scialom · said 148 times in 53 episodes since 2024 · across every show →

Mentions by year

brought up most by Shawn Wang (18), Alessio Fanelli (14), Erik Schluntz (7), David Hershey (7), Matthew Berman (6), Mike Merrill (5), Mike Krieger (5), Lukas Petersson (4)

tap a year for its mentions
0040208040202420252026episodesmentions
02040202420252026episodes it came up in
00220440202420252026episodesmentions per episode
2026 24 mentions in 9 episodes 3 per episode
2025 74 mentions in 31 episodes 2 per episode
2024 50 mentions in 13 episodes 4 per episode

every mention, scene by scene, with the transcript →

Everything said about Claude 3.5 Sonnet, oldest first

Jul 23, 2024 positive
Assertion Supported
Scialom: Llama 3 405B is the best open-source model ever released
“At a high level, it's the best open source model ever. It's Better than GPT-IV. I mean, what version? But, by far, compared to the version originally released even now, I think there's maybe the last cloud Sonya FF-V and GPT-IV-Zero that are performing it.”
Thomas Scialom Jul 23, 2024 ▶ 37:43 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Oct 13, 2024 bullish
Assertion Supported
Molmo Outperforms Gemini 1.5 and Claude 3.5 Sonnet With 1M Samples
“They can get better than Gemini, 1.5, better than Claude, 3.5 sonnet, better than GPT for V at a much smaller size with about a million samples of data, which is very impressive, right?”
Vibhu Sapra Oct 13, 2024 ▶ 4:23 [Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz
Oct 18, 2024 positive
Opinion
Houston: Claude 3.5 Sonnet is probably the best model all around
“Sonnet three five is probably the best all around”
Drew Houston Oct 18, 2024 ▶ 12:53 Building the Silicon Brain - Drew Houston of Dropbox
Nov 11, 2024 positive
Assertion Supported
Polu: Claude Sonnet executes an unpublicized chain-of-thought step during function calling
“They kind of innovated in an interesting way, which was never quite publicized, but it's that they have that kind of chain of thoughts step whenever you use a Clouds model or Sonnet model with function calling. That chain of service step doesn't exist when you…”
Stanislas Polu Nov 11, 2024 ▶ 42:20 Agents @ Work: Dust.tt — with Stanislas Polu
Nov 28, 2024
Insight
Schluntz: Smarter AI models require less agent scaffolding
“And I think like the smarter the models are, the less you need that kind of extra scaffolding.”
Erik Schluntz Nov 28, 2024 ▶ 20:06 The new Claude 3.5 Sonnet, Computer Use, and Building SOTA Agents — with Erik Schluntz, Anthropic
Nov 28, 2024
Disclosure
Anthropic: Tool engineering mattered more than prompt engineering for SWE-bench
“I would say actually we did more engineering of the tools than the overall prompt.”
Erik Schluntz Nov 28, 2024 ▶ 22:53 The new Claude 3.5 Sonnet, Computer Use, and Building SOTA Agents — with Erik Schluntz, Anthropic
Apr 27, 2025 neutral
Assertion Supported
Claude 3.5 wrote fire-and-forget code while GPT-4 used defensive programming
“Claude, for instance, the Sonnet 3.5 was very much fire and forget. It would write code in a kind of Pythonic way, just like, let it fail. Don't be careful about it. Whereas GPT four would use defensive programming, use self assertions.”
Jack Hopkins Apr 27, 2025 ▶ 16:27 ⚡️Factorio Learning Environment: the ultimate Game Agent Eval — Jack Hopkins
Jun 4, 2026 neutral
Insight
Petersson: Pre-RL LLM Agents Act Like Compliant Assistants, Not Business Owners
“The models are like super trained to be assistants at least at this point in time. So that's why it's, it went into that kind of experiment instead. Like it just, every time you asked for something, it just did it. And it was more like an assistant. We've seen…”
Lukas Petersson Jun 4, 2026 ▶ 20:07 When AI Agents Run Businesses — Lukas Petersson and Axel Backlund of Andon Labs
Jun 4, 2026
Assertion Supported
Claude 3.5 Sonnet reported $2 benchmark rent to the FBI as cybercrime
“So it, like, claimed that it had stopped, but it saw that its bank account still was, like, drained two dollars, and it said that this is, like, cybercrime, and it first reported it once to the FBI, like, oh, there's cybercrime here, like, they're stealing two…”
Lukas Petersson Jun 4, 2026 ▶ 15:37 When AI Agents Run Businesses — Lukas Petersson and Axel Backlund of Andon Labs
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.