Assertion Supported
Labenz: Claude 4 system card documented AI blackmailing a human engineer
“In the cloud four system card, they reported blackmailing of the human. The setup was that the AI had access to the engineer's email and They told the AI that it was going to be like replaced with a, you know, a less ethical version or something like that. It …”
Prediction Not checkable as stated
Labenz: Tool-Equipped Next-Gen AI Models Will Resemble Superintelligence
“When we start to give the next generation of the model these power tools, and they start to solve previously unsolved engineering problems, I think you start to have something that looks kind of like super intelligence.”
Opinion
Labenz: GPT-4 to GPT-5 capability leap matches GPT-3 to GPT-4
“And if you look back to GPT three, you know, there's a huge leap. I would contend that the leap is similar from GPT four to five.”
Prediction Not checkable as stated
Labenz: AI will outperform average developers on standard apps within five years
“But I would be very surprised if you can't get your nuts and bolts Web app, mobile app type things spit out for you for far less and far faster than, and probably honestly with significantly higher quality and less back and forth with an AI system than, you kn…”
Assertion Not checkable as stated
Labenz: OpenAI's router failure caused bad initial GPT-5 outputs
“The problem at launch was that that router was broken. So all of the queries were going to the dumb model, and so a lot of people literally just got Bad outputs, which were worse than oh three because they were getting non thinking responses.”
Assertion Not checkable as stated
Labenz: Plugged-in AI experts are not pushing timelines past 2030
“I don't think too many people, at least that I, you know, think are really plugged in on this, are pushing out too much past 20 30 at all.”
Prediction Not checkable as stated
Labenz: AI customer service agents will cause significant headcount reductions
“So I don't think these things go to zero probably in a lot of environments, but I do expect that you will see significant headcount reduction in a lot of these”
Prediction Not checkable as stated
Labenz: Existing AI capabilities could automate 50-80% of work in 5-10 years
“I think if progress stopped today, I still think we could get to 50 to 80% of work automated over the next, like, five to 10 years.”
Prediction Not checkable as stated
Labenz: AI agent task capacity will reach two weeks within two years
“If you extrapolate that out a bit and you're like, okay, take, take the four month case just to be a little aggressive. That's three doublings a year. That's eight X task length increase per year. That would mean you go from two hours now to two days. In one y…”
Opinion
Labenz: Chinese open-source AI models have surpassed American open-source models
“For those that are using open source, I do think it's true that the Chinese models have become the best.”
Insight
Labenz: AI post-training reasoning currently yields higher ROI than raw scaling
“And it just seems like we're getting more benefit from the post training and the reasoning paradigm than scaling. But I don't think either one is I definitely don't think either one is, is dead.”
Assertion Supported
Labenz: Pure reasoning AI models achieved IMO gold without external tools
“Well, I mean, a big one from just the last few weeks was that we had an IMO gold medal with pure reasoning models with no access to tools from multiple companies. And, you know, that is night and day compared to what GPT-IV could do with math, right?”
Assertion Supported
Labenz: Google's AI Co-Scientist solved an open virology problem independently
“And they gave it legitimately unsolved problems in science, and in one particularly famous, kind of notorious case, it came up with a Hypothesis, which it wasn't able to verify because it doesn't have direct access to actually run the experiments in the lab, b…”
Assertion Supported
Labenz: Intercom's Fin AI agent resolves 65% of support tickets
“They now have this fin agent that is solving like 65% of customer service tickets that come in.”
Assertion Not checkable as stated
Labenz: AI auditor agent won state contract for 1M annual document audits
“They've created this auditor AI agent that just won a state-level contract to do the audits on, like, a million transactions a year of these You know, these packets of documents, again, scanned, handwritten, all this kind of crap and they just blew away the hu…”
Assertion Supported
OpenAI o3 model completes 40% of internal research engineer pull requests
“That's another data point, by the way, from this was from the O three system card. They showed a jump from like low to mid single digits to roughly 40% of PRs actually checked in by Research engineers at OpenAI that the model could do. So prior to O three, not…”
Opinion
Labenz: China may be ahead of the US in general robotics
“And this is one area where I do think China might be actually ahead of the United States right now”
Insight
Labenz: AI scaling laws are empirical observations, not guaranteed natural laws
“The scaling law idea, which is, you know, it's definitely worth agreeing, taking a moment to note that it is not a law of nature. You know, we do not have a principled reason to believe that scaling is some law that will go indefinitely. All we really know is …”
Assertion Supported
Labenz: GPT-4.5 achieved 65% accuracy on SimpleQA versus o3's 50%
“The O-three class of models got about a 50% on that benchmark, and GPT 4.5 popped up to like 65%. So, in other words, it basically, of the things that were not known to the previous generation of models, it picked up a third of them.”
Assertion Partly supported
Labenz: Per-token model costs fell 95% from GPT-4 to GPT-5
“It's like 90 it's like a 95% discount from GPT-IV to GPT-V.”
Assertion Supported
MIT researchers used AI models to create novel antibiotics for resistant bacteria
“It's been enough for this group at MIT to use some of these relatively, you know, narrow purpose-built biology models and create totally new antibiotics. New in the sense that they have a new mechanism of action. Like they're affecting the bacteria in a new wa…”
Prediction Not checkable as stated
Labenz: Non-text AI modalities will unify with language models over time
“We have seen this play out with text and image where you had your text only models and you had your image only models, and then they started to come together and now they've come really deeply together. And so I think you're going to see that across a lot of o…”
Prediction Not checkable as stated
Labenz: LLM fine-tuning and reinforcement learning will successfully power humanoid robotics
“All these techniques that have been developed over the last few years, Seems to me they're absolutely gonna apply to a problem like a humanoid robot as well.”
Assertion Supported
Labenz: Near Protocol pivoted to crypto to solve international AI worker payments
“They took a huge detour into crypto because they were trying to hire task workers around the world and couldn't figure out how to pay them. So they were like, this sucks so bad to pay these task workers in all these different countries that we're trying to get…”