Jun 21, 2026 · 33m · latent-space

⚡️Every product of the future will be a living system — Ronak Malde, Trajectory.ai

Ronak Malde · 26m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this Latent Space podcast episode, Trajectory.ai co-founder Ronak Malde explains how his experience at Windsurf and Google DeepMind inspired him to build continual learning infrastructure that transforms static enterprise foundation models into self-improving, living systems.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 4.0 Guest teaching 3.0 Guest disagreement 0.6 The hosts pushing back 0.7
05100:0010:0020:0030:000:03–6:36 · The hosts as informed peer 5/10 Welcome to Latent Space and Early Windsurf Breakthroughs The host exhibits strong insider context regarding the Windsurf team, DeepMind acquihire, and recent deals like Cursor and xAI. The tone is highly collaborative and celebratory as Ronak shares insider anecdotes from the Google acquisition.6:37–9:17 · The hosts as informed peer 2/10 Founding Trajectory.ai to Solve Continual Learning The host poses open-ended conversational prompts about Ronak leaving DeepMind to start Trajectory.ai. Ronak details the founding thesis around continual learning and his co-founders' backgrounds across robotics and vision.9:18–13:37 · The hosts as informed peer 3/10 Powering Legal AI with Harvey and NVIDIA Nemotron The host asks about different schools of thought in continual learning. Ronak walks through practical case studies with Harvey and NVIDIA Nemotron, illustrating why legal domains have zero tolerance for errors compared to coding.13:38–17:50 · The hosts as informed peer 5/10 Open Source Frontier Models and Real-World Data Curation The host demonstrates domain familiarity by noting base model trends and SWE-ONE's lineage before probing on data curation mechanisms. Ronak explains why binary user feedback like thumbs up/down is useless noise compared to active human edit trajectories.17:51–23:08 · The hosts as informed peer 3/10 Scaling Self-Distillation Policy Optimization (SDPO) Ronak explains the mechanics of Self-Distillation Policy Optimization (SDPO) and why standard RL scalar rewards fail for real-world continual learning. When the host assumes Trajectory is just a three-person team, Ronak politely corrects him, noting they have scaled to eleven researchers and engineers.23:10–28:06 · The hosts as informed peer 6/10 Open Sourcing Continuous LoRA for Concurrent Training The host actively inspects the benchmark results, questioning why latency increases at eight concurrent runs, and bridges Continuous LoRA to classical operating system concepts like preemptive scheduling and starvation. Ronak affirms the systems-level comparison.28:08–33:52 · The hosts as informed peer 4/10 Future Roadmap: Autonomous Harnesses, Observability, and Enterprise Scale The host tests whether focusing exclusively on verifiable domains like coding would be safer than general enterprise domains. Ronak defends the domain-agnostic approach, and the host closes with praise for the launch execution.0:03–6:36 · Guest teaching 2/10 Welcome to Latent Space and Early Windsurf Breakthroughs The host exhibits strong insider context regarding the Windsurf team, DeepMind acquihire, and recent deals like Cursor and xAI. The tone is highly collaborative and celebratory as Ronak shares insider anecdotes from the Google acquisition.6:37–9:17 · Guest teaching 2/10 Founding Trajectory.ai to Solve Continual Learning The host poses open-ended conversational prompts about Ronak leaving DeepMind to start Trajectory.ai. Ronak details the founding thesis around continual learning and his co-founders' backgrounds across robotics and vision.9:18–13:37 · Guest teaching 4/10 Powering Legal AI with Harvey and NVIDIA Nemotron The host asks about different schools of thought in continual learning. Ronak walks through practical case studies with Harvey and NVIDIA Nemotron, illustrating why legal domains have zero tolerance for errors compared to coding.13:38–17:50 · Guest teaching 3/10 Open Source Frontier Models and Real-World Data Curation The host demonstrates domain familiarity by noting base model trends and SWE-ONE's lineage before probing on data curation mechanisms. Ronak explains why binary user feedback like thumbs up/down is useless noise compared to active human edit trajectories.17:51–23:08 · Guest teaching 5/10 Scaling Self-Distillation Policy Optimization (SDPO) Ronak explains the mechanics of Self-Distillation Policy Optimization (SDPO) and why standard RL scalar rewards fail for real-world continual learning. When the host assumes Trajectory is just a three-person team, Ronak politely corrects him, noting they have scaled to eleven researchers and engineers.23:10–28:06 · Guest teaching 3/10 Open Sourcing Continuous LoRA for Concurrent Training The host actively inspects the benchmark results, questioning why latency increases at eight concurrent runs, and bridges Continuous LoRA to classical operating system concepts like preemptive scheduling and starvation. Ronak affirms the systems-level comparison.28:08–33:52 · Guest teaching 2/10 Future Roadmap: Autonomous Harnesses, Observability, and Enterprise Scale The host tests whether focusing exclusively on verifiable domains like coding would be safer than general enterprise domains. Ronak defends the domain-agnostic approach, and the host closes with praise for the launch execution.0:03–6:36 · Guest disagreement 0/10 Welcome to Latent Space and Early Windsurf Breakthroughs The host exhibits strong insider context regarding the Windsurf team, DeepMind acquihire, and recent deals like Cursor and xAI. The tone is highly collaborative and celebratory as Ronak shares insider anecdotes from the Google acquisition.6:37–9:17 · Guest disagreement 0/10 Founding Trajectory.ai to Solve Continual Learning The host poses open-ended conversational prompts about Ronak leaving DeepMind to start Trajectory.ai. Ronak details the founding thesis around continual learning and his co-founders' backgrounds across robotics and vision.9:18–13:37 · Guest disagreement 1/10 Powering Legal AI with Harvey and NVIDIA Nemotron The host asks about different schools of thought in continual learning. Ronak walks through practical case studies with Harvey and NVIDIA Nemotron, illustrating why legal domains have zero tolerance for errors compared to coding.13:38–17:50 · Guest disagreement 1/10 Open Source Frontier Models and Real-World Data Curation The host demonstrates domain familiarity by noting base model trends and SWE-ONE's lineage before probing on data curation mechanisms. Ronak explains why binary user feedback like thumbs up/down is useless noise compared to active human edit trajectories.17:51–23:08 · Guest disagreement 1/10 Scaling Self-Distillation Policy Optimization (SDPO) Ronak explains the mechanics of Self-Distillation Policy Optimization (SDPO) and why standard RL scalar rewards fail for real-world continual learning. When the host assumes Trajectory is just a three-person team, Ronak politely corrects him, noting they have scaled to eleven researchers and engineers.23:10–28:06 · Guest disagreement 1/10 Open Sourcing Continuous LoRA for Concurrent Training The host actively inspects the benchmark results, questioning why latency increases at eight concurrent runs, and bridges Continuous LoRA to classical operating system concepts like preemptive scheduling and starvation. Ronak affirms the systems-level comparison.28:08–33:52 · Guest disagreement 0/10 Future Roadmap: Autonomous Harnesses, Observability, and Enterprise Scale The host tests whether focusing exclusively on verifiable domains like coding would be safer than general enterprise domains. Ronak defends the domain-agnostic approach, and the host closes with praise for the launch execution.0:03–6:36 · The hosts pushing back 0/10 Welcome to Latent Space and Early Windsurf Breakthroughs The host exhibits strong insider context regarding the Windsurf team, DeepMind acquihire, and recent deals like Cursor and xAI. The tone is highly collaborative and celebratory as Ronak shares insider anecdotes from the Google acquisition.6:37–9:17 · The hosts pushing back 0/10 Founding Trajectory.ai to Solve Continual Learning The host poses open-ended conversational prompts about Ronak leaving DeepMind to start Trajectory.ai. Ronak details the founding thesis around continual learning and his co-founders' backgrounds across robotics and vision.9:18–13:37 · The hosts pushing back 0/10 Powering Legal AI with Harvey and NVIDIA Nemotron The host asks about different schools of thought in continual learning. Ronak walks through practical case studies with Harvey and NVIDIA Nemotron, illustrating why legal domains have zero tolerance for errors compared to coding.13:38–17:50 · The hosts pushing back 1/10 Open Source Frontier Models and Real-World Data Curation The host demonstrates domain familiarity by noting base model trends and SWE-ONE's lineage before probing on data curation mechanisms. Ronak explains why binary user feedback like thumbs up/down is useless noise compared to active human edit trajectories.17:51–23:08 · The hosts pushing back 0/10 Scaling Self-Distillation Policy Optimization (SDPO) Ronak explains the mechanics of Self-Distillation Policy Optimization (SDPO) and why standard RL scalar rewards fail for real-world continual learning. When the host assumes Trajectory is just a three-person team, Ronak politely corrects him, noting they have scaled to eleven researchers and engineers.23:10–28:06 · The hosts pushing back 3/10 Open Sourcing Continuous LoRA for Concurrent Training The host actively inspects the benchmark results, questioning why latency increases at eight concurrent runs, and bridges Continuous LoRA to classical operating system concepts like preemptive scheduling and starvation. Ronak affirms the systems-level comparison.28:08–33:52 · The hosts pushing back 1/10 Future Roadmap: Autonomous Harnesses, Observability, and Enterprise Scale The host tests whether focusing exclusively on verifiable domains like coding would be safer than general enterprise domains. Ronak defends the domain-agnostic approach, and the host closes with praise for the launch execution.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 21:50 Correcting host on team headcount

Ronak gently interrupts and corrects the host's assumption that Trajectory is operating with only three founders, explaining they have scaled to eleven people.

Hardest push from the hosts ▶ 25:33 Interrogating concurrency degradation on chart

The host spots a trend anomaly on the benchmark slide and immediately presses Ronak to explain why eight concurrent training runs cause the curve to increase.

Biggest teaching moment ▶ 19:10 Breaking down RL scalar limitations versus SDPO

Ronak educates the audience and host on why traditional RL policy updates collapse real-world textual feedback into a single scalar number and how privileged teacher hints solve this.

The host holds their own ▶ 27:10 Mapping continuous LoRA to operating system schedulers

The host showcases deep systems engineering knowledge by recontextualizing the guest's distributed ML architecture in terms of classic OS preemptive scheduling and resource starvation.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Welcome to Latent Space and Early Windsurf Breakthroughs 5200 The host exhibits strong insider context regarding the Windsurf team, DeepMind acquihire, and recent deals like Cursor and xAI. The tone is highly collaborative and celebratory as Ronak shares insider anecdotes from the Google acquisition.
Founding Trajectory.ai to Solve Continual Learning 2200 The host poses open-ended conversational prompts about Ronak leaving DeepMind to start Trajectory.ai. Ronak details the founding thesis around continual learning and his co-founders' backgrounds across robotics and vision.
Powering Legal AI with Harvey and NVIDIA Nemotron 3410 The host asks about different schools of thought in continual learning. Ronak walks through practical case studies with Harvey and NVIDIA Nemotron, illustrating why legal domains have zero tolerance for errors compared to coding.
Open Source Frontier Models and Real-World Data Curation 5311 The host demonstrates domain familiarity by noting base model trends and SWE-ONE's lineage before probing on data curation mechanisms. Ronak explains why binary user feedback like thumbs up/down is useless noise compared to active human edit trajectories.
Scaling Self-Distillation Policy Optimization (SDPO) 3510 Ronak explains the mechanics of Self-Distillation Policy Optimization (SDPO) and why standard RL scalar rewards fail for real-world continual learning. When the host assumes Trajectory is just a three-person team, Ronak politely corrects him, noting they have scaled to eleven researchers and engineers.
Open Sourcing Continuous LoRA for Concurrent Training 6313 The host actively inspects the benchmark results, questioning why latency increases at eight concurrent runs, and bridges Continuous LoRA to classical operating system concepts like preemptive scheduling and starvation. Ronak affirms the systems-level comparison.
Future Roadmap: Autonomous Harnesses, Observability, and Enterprise Scale 4201 The host tests whether focusing exclusively on verifiable domains like coding would be safer than general enterprise domains. Ronak defends the domain-agnostic approach, and the host closes with praise for the launch execution.

Statements from this episode (12)

Assertion Not checkable as stated
Malde: SWE-ONE beat frontier models via user-signal post-training
“And this was the kind of major unlock for the company as well, is we had all this massive data. We were able to post train on all of that user signal and now beat the frontier.”
Ronak Malde Jun 21, 2026 ▶ 3:06
Assertion Not checkable as stated
Malde walked away from $2B DeepMind acquisition to found Trajectory.ai
“Obviously the acquisition was for two billion dollars and went over to DeepBind, and then I decided to give up all the acquisition money to start trajectory.”
Ronak Malde Jun 21, 2026 ▶ 6:53
Prediction Not checkable as stated
Malde: Continual learning will be AI's next major unlock
“And we realized continual learning is kind of the ultimate, like, paradigm to do that. Is like, how do you have humans in the loop? How do you build this intelligence around them that is constantly learning and growing on its own? And I think that's going to b…”
Ronak Malde Jun 21, 2026 ▶ 8:47
Insight
Malde: In legal AI, getting 80% of the way there is zero
“For a field like legal, like getting 80% of the way there is the same thing as zero.”
Ronak Malde Jun 21, 2026 ▶ 10:10
Assertion Supported
Malde: NVIDIA Nemotron is drastically cheaper and faster than frontier models
“The really cool part is Nemetron is a drastically cheaper and faster model than the frontier.”
Ronak Malde Jun 21, 2026 ▶ 12:17
Disclosure
Malde reveals Trajectory.ai partners: Clay, Harvey, Rogo, Decagon, and Mercor
“Our current partners are Clay, Harvey, Rogo, Dakagon, Mercore.”
Ronak Malde Jun 21, 2026 ▶ 13:02
Opinion
Malde: Western open-source AI lags Chinese models at trillion-parameter scale
“I think America or the Western world has some work to do still. Like obviously having one trillion parameter models like Kimi or like an amazing models like GLM and DeepSeq, I don't think we're quite there yet for that size of model.”
Ronak Malde Jun 21, 2026 ▶ 14:41
Insight
Malde: Direct user modifications, not binary feedback, drive effective AI continual learning
“I think when people think about, like, online learning, continual learning, they'll first think of, like, accept, reject, or thumbs up, thumbs down, like some of those, like, kind of binary signals. It turns out that that's actually, like, very noisy. You can …”
Ronak Malde Jun 21, 2026 ▶ 16:54
Opinion
Malde: Standard reinforcement learning is broken for continual learning
“RL, it's still taking all of this kind of Useful information from the real world, like I mentioned, all the corrections and everything, and putting it into just one number. Which is really broken.”
Ronak Malde Jun 21, 2026 ▶ 19:13
Assertion Not checkable as stated
Malde: Nobody Scaled SDPO to Real-World Cases Before Trajectory
“It's been done in a lot of academic cases, but no one's actually been able to scale it up to real world use cases.”
Ronak Malde Jun 21, 2026 ▶ 20:55
Disclosure
Trajectory.ai open-sources continual learning training stack with Berkeley and Anyscale
“One of the other exciting things we did is, is open sourcing a training stack for continual learning. And so this is in conjunction with Sky RL Berkeley's Sky RL lab in any scale as well.”
Ronak Malde Jun 21, 2026 ▶ 23:38
Assertion Supported
Trajectory: Continuous LoRA cuts training wall-clock time in half for two concurrent jobs
“We ran this on several different scale up of experiments, and you can see that across the board, even as you scale up experiments, so like two concurrent jobs were able to cut the wall clock time in half.”
Ronak Malde Jun 21, 2026 ▶ 25:15
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.