Ronak Malde, CEO of Trajectory.ai, contrasts the error tolerance of software engineering agents with mission-critical legal AI workflows.
Assertion Not checkable as stated
Malde: SWE-ONE beat frontier models via user-signal post-training
“And this was the kind of major unlock for the company as well, is we had all this massive data. We were able to post train on all of that user signal and now beat the frontier.”
Prediction Not checkable as stated
Malde: Continual learning will be AI's next major unlock
“And we realized continual learning is kind of the ultimate, like, paradigm to do that. Is like, how do you have humans in the loop? How do you build this intelligence around them that is constantly learning and growing on its own? And I think that's going to b…”
Opinion
Malde: Western open-source AI lags Chinese models at trillion-parameter scale
“I think America or the Western world has some work to do still. Like obviously having one trillion parameter models like Kimi or like an amazing models like GLM and DeepSeq, I don't think we're quite there yet for that size of model.”
Insight
Malde: Direct user modifications, not binary feedback, drive effective AI continual learning
“I think when people think about, like, online learning, continual learning, they'll first think of, like, accept, reject, or thumbs up, thumbs down, like some of those, like, kind of binary signals. It turns out that that's actually, like, very noisy. You can …”
Opinion
Malde: Standard reinforcement learning is broken for continual learning
“RL, it's still taking all of this kind of Useful information from the real world, like I mentioned, all the corrections and everything, and putting it into just one number. Which is really broken.”
Assertion Not checkable as stated
Malde: Nobody Scaled SDPO to Real-World Cases Before Trajectory
“It's been done in a lot of academic cases, but no one's actually been able to scale it up to real world use cases.”