inference engine
2 statements across 2 episodes · 0 bullish · 0 bearish · 2 people on the record · first statement Aug 4, 2025 by Stefano Ermon · across every show →
Everything said about inference engine, oldest first
Aug 4, 2025 neutral
Ermon: Inception Labs built proprietary engine for production inference traffic
“So just like you would normally serve an LLM using a VLLM or SGLang or a Tensor or TLLM, we have built our own inference engine. And so we are supporting production traffic already with our own inference engine. We support continuous batch and quantization.”
Aug 3, 2026
Speculative decoding creates hardware resource contention and engine orchestration complexity
“The thing with speculators is one of the practical constraints on using them is that you do have to run a small model on the same hardware that you're running the big model on. There is a orchestration and resource competition problem inherent in that. And tha…”