Assertion Supported AI assessment confidence: 95% certainty 4/5 debate potential 2/5

Trojanowski: DeepSeek-R1 succeeded by scaling outcome supervision over process supervision

Mitch Trojanowski · How to Build Autonomous, Long-Horizon AI Agents | Basis · Aug 5, 2026 · at 20:03

Mitch Trojanowski, co-founder of Basis, analyzes how reasoning models like DeepSeek-R1 use RL with verifiable rewards (RLVR) based on final outcomes rather than intermediate step supervision.

0:00 / 0:27exact quote · 27.9s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“If you look, you know, if you fast forward a bit and you look at the, like the DeepSeq R-one paper where they effectively laid out, you know, what I think all the labs were doing at that time, or at least OpenAI was doing in terms of you know, RLVR reasoning from reinforcement learning from verifiable rewards that effectively had very little process supervision. And instead was essentially just saying, hey, did you get the outcome right? Yes. Okay. Let me reward you. And then scaling that up which, you know, obviously worked well.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Mitch Trojanowski

Opinion
Trojanowski: Applying the Bitter Lesson to Complex Tax Workflows Is a Mistake
“I think it is a mistake to throw out those learnings and say, you know, bitter lesson, throw out those learnings. We're just going to have the agents at runtime develop an entirely new way to do a tax return that is, you know, because bitter lesson, yada, yada…”
Mitch Trojanowski Aug 5, 2026 ▶ 34:40 How to Build Autonomous, Long-Horizon AI Agents | Basis
Insight
Trojanowski: Technical AI moats are not real moats
“Technical moats are not real moats. Like, I, there's no portion of basis's long-term terminal value that stems from some, you know, secret RL trick we found that nobody else found.”
Mitch Trojanowski Aug 5, 2026 ▶ 1:19:08 How to Build Autonomous, Long-Horizon AI Agents | Basis
Assertion Not checkable as stated
Trojanowski: AI coding agents remain worse than junior engineers over two-week projects
“Even now with coding, like, the agents are not yet they're not human level at being coherent over long periods of time. That's obvious because they can't code like a junior engineer on a project for two weeks. So they can't, they're, that's worse than a human …”
Mitch Trojanowski Aug 5, 2026 ▶ 26:35 How to Build Autonomous, Long-Horizon AI Agents | Basis
Opinion
Trojanowski: Enterprise agents need process reliability, not 'Move 37' breakthroughs
“And so the thing that you know, someone is buying from us is not, this will be the best ever tax return. They're buying that, you know, the confidence... They're buying that it's going to be consistent and reliable and something that they can trust that actual…”
Mitch Trojanowski Aug 5, 2026 ▶ 49:24 How to Build Autonomous, Long-Horizon AI Agents | Basis
Insight
Trojanowski: 10-Hour Agent Runs Are Not Black Boxes
“Not thinking that an agent operating over 10 hours is a black box. It's not. It has a lot of data, and you're probably doing a disservice to your customers if you don't understand, like, how it's going about the work.”
Mitch Trojanowski Aug 5, 2026 ▶ 58:07 How to Build Autonomous, Long-Horizon AI Agents | Basis
Prediction Not checkable as stated
Trojanowski: Closing the agent self-improvement loop will be close by year-end
“I think that closing the loop is gonna happen pretty fast. I think you'll have, like, I don't know about the entire loop being closed, but I think you'll be Relatively close by end of year.”
Mitch Trojanowski Aug 5, 2026 ▶ 1:12:02 How to Build Autonomous, Long-Horizon AI Agents | Basis
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.