Inference Engine

topic on 2 shows · 3 statements across 3 episodes

Latent Space the Neon Show

3 statements about Inference Engine, every show

NEON SHOW Disclosure
Goel: Cartesia Built Custom Inference Engine for Real-Time Models
“So as an example, we've built our own inference engine. It's pretty important because we think that these interactive real-time models are going to need to run in a pretty different way. So having the ability to have your own engine that is not designed for LL…”
Karan Goel Aug 7, 2026 ▶ 21:08 The Billion Dollar AI Lab Founder Who Sees The Future First | Karan Goel, Founder & CEO of Cartesia
Speculative decoding creates hardware resource contention and engine orchestration complexity
“The thing with speculators is one of the practical constraints on using them is that you do have to run a small model on the same hardware that you're running the big model on. There is a orchestration and resource competition problem inherent in that. And tha…”
Philip Kiely Aug 3, 2026 ▶ 47:02 Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
LATENT SPACE Assertion Not checkable as stated
Ermon: Inception Labs built proprietary engine for production inference traffic
“So just like you would normally serve an LLM using a VLLM or SGLang or a Tensor or TLLM, we have built our own inference engine. And so we are supporting production traffic already with our own inference engine. We support continuous batch and quantization.”
Stefano Ermon Aug 4, 2025 ▶ 21:33 ⚡️Mercury: Ultra-Fast Diffusion LLMs — Estefano Ermon, CEO Inception Labs

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.