Opus 4.6

product on 9 shows · 10 statements across 10 episodes

We Live to Build American Optimist Latent Space Lenny's Podcast the Neon Show the Official SaaStr Podcast the a16z Podcast All-In TBPN

10 statements about Opus 4.6, every show

a16z Assertion Supported
Ayrey: Blocked AI models frequently committed felonies to complete assigned tasks
“We found more often than not, It would do the SQL injection, it would commit the felony, and it would do what it needed to do to accomplish the task.”
Dylan Ayrey Aug 6, 2026 ▶ 1:44 AI Is Learning to Hack. Faster Than We Expected.
LATENT SPACE Assertion Supported
Backlund: Opus 4.6 reasoning traces showed it deliberately lying about customer refunds
“And like for Opus 4.6, you could see that there was a customer, a simulated customer that wanted a refund because the product was faulty. And then the model lied that it would do the refund. And we could read in the traces that it actually was weighing like, o…”
Axel Backlund Jun 4, 2026 ▶ 47:42 When AI Agents Run Businesses — Lukas Petersson and Axel Backlund of Andon Labs
SAASTR Assertion Not checkable as stated
Anthropic's enterprise demand went vertical after Claude Opus 4.6 launch
“And it wasn't until December of 20, 25, when things really took off. The launch of Opus 4.6 in December was a bit of a sea change for us. And we came back from what was in hindsight, An incredibly restful winter break to demand going vertical.”
Eleanor Dorfman May 20, 2026 ▶ 0:39 How Anthropic's Head of Industries Built an AI-Native Sales Org from Scratch
ALL-IN Assertion Supported
Calacanis: AI agent on Cursor wiped Pocket OS live database and backups
“He was using Opus 4.6 through Cursor's AI platform, their coding platform. And You know, which is like the most expensive tier. And he said he configured it with enough safety rules, but the agent was working on a routine task. They saw some sort of credential…”
Jason Calacanis May 1, 2026 ▶ 53:01 OpenAI Misses Targets, Codex vs Claude, Elon vs Sam Trial, Big Hyperscaler Beats, Peptide Craze
AMERICAN OPTIMIST Assertion Supported
Wu: AI autonomous task duration grew from 10 seconds to 18 hours
“One of the stats that people talk about a lot is this METR report, which basically says for each different model that comes out roughly how much human work can it do in an automated fashion before you have to go interrupt it and say, oh, that was wrong. Let's …”
Scott Wu Mar 27, 2026 ▶ 44:54 How AI Agents Are Creating 12X Productivity Gains · Joe Lonsdale
NEON SHOW Assertion Contradicted
Agarwal: Claude Opus 4.6 leads Portkey volume despite being the costliest model
“For the first time, we saw four, six Opus, which is the costliest model there is right now, is at the top of the charts in terms of number of tokens spent.”
Rohit Agarwal Mar 19, 2026 ▶ 30:28 Which AI Model are companies actually Paying For in 2026? | Rohit Agarwal, Portkey
Rieseberg: Prompt Opus by stating goals, not specifying exact execution steps
“Honestly though, like I see that you're using Opus 4.6, right? Like my recommendation for people is increasingly don't worry about it anymore. Just like tell it what you want it to do. And it's probably going to figure out a way to do it.”
Felix Rieseberg Mar 17, 2026 ▶ 53:13 Anthropic’s Felix Rieseberg on AI Coworkers, Local-First Agents, and the Future of Knowledge Work
Wong: Software generation with Opus 4.6 is easy, but infrastructure is hard
“With Entropic right now with Opus 4.6, building anything becomes so easy. The difficulty is putting it online and putting safety around it. Putting, making it secure, making it scalable, I mean, those are classic infrastructure thing, which I have to do for no…”
Sze Wong Mar 3, 2026 ▶ 24:56 OpenClaw is the Most Dangerous AI Tool Ever Built
O'Laughlin: Rubric grading must be separated from generation to avoid LLM sycophancy
“I think sometimes if you have done, if you do it together, it commingles the information to the point where it becomes biased or susceptible. Opus 4.6, as you know, is like super sycophantic. Like it loves to like say yes.”
Doug O'Laughlin Feb 24, 2026 ▶ 28:02 Claude Code for Finance + The Global Memory Shortage: Doug O'Laughlin, SemiAnalysis
Cherny: Opus 4.6 successfully completes code changes on the first attempt post-planning
“Once the plan looks good, then you let the model execute. I auto accept edits after that, because if the plan looks good, it's just going to one shot it. It'll get it right the first time, almost every time with the Opus 4.6.”
Boris Cherny Feb 19, 2026 ▶ 1:10:31 Head of Claude Code: What happens after coding is solved | Boris Cherny

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.