Ayrey: Blocked AI models frequently committed felonies to complete assigned tasks
“We found more often than not, It would do the SQL injection, it would commit the felony, and it would do what it needed to do to accomplish the task.”
Backlund: Opus 4.6 reasoning traces showed it deliberately lying about customer refunds
“And like for Opus 4.6, you could see that there was a customer, a simulated customer that wanted a refund because the product was faulty. And then the model lied that it would do the refund. And we could read in the traces that it actually was weighing like, o…”
Anthropic's enterprise demand went vertical after Claude Opus 4.6 launch
“And it wasn't until December of 20, 25, when things really took off. The launch of Opus 4.6 in December was a bit of a sea change for us. And we came back from what was in hindsight, An incredibly restful winter break to demand going vertical.”
Calacanis: AI agent on Cursor wiped Pocket OS live database and backups
“He was using Opus 4.6 through Cursor's AI platform, their coding platform. And You know, which is like the most expensive tier. And he said he configured it with enough safety rules, but the agent was working on a routine task. They saw some sort of credential…”
Wu: AI autonomous task duration grew from 10 seconds to 18 hours
“One of the stats that people talk about a lot is this METR report, which basically says for each different model that comes out roughly how much human work can it do in an automated fashion before you have to go interrupt it and say, oh, that was wrong. Let's …”
Agarwal: Claude Opus 4.6 leads Portkey volume despite being the costliest model
“For the first time, we saw four, six Opus, which is the costliest model there is right now, is at the top of the charts in terms of number of tokens spent.”
Rieseberg: Prompt Opus by stating goals, not specifying exact execution steps
“Honestly though, like I see that you're using Opus 4.6, right? Like my recommendation for people is increasingly don't worry about it anymore. Just like tell it what you want it to do. And it's probably going to figure out a way to do it.”
Wong: Software generation with Opus 4.6 is easy, but infrastructure is hard
“With Entropic right now with Opus 4.6, building anything becomes so easy. The difficulty is putting it online and putting safety around it. Putting, making it secure, making it scalable, I mean, those are classic infrastructure thing, which I have to do for no…”
O'Laughlin: Rubric grading must be separated from generation to avoid LLM sycophancy
“I think sometimes if you have done, if you do it together, it commingles the information to the point where it becomes biased or susceptible. Opus 4.6, as you know, is like super sycophantic. Like it loves to like say yes.”
Cherny: Opus 4.6 successfully completes code changes on the first attempt post-planning
“Once the plan looks good, then you let the model execute. I auto accept edits after that, because if the plan looks good, it's just going to one shot it. It'll get it right the first time, almost every time with the Opus 4.6.”