Mar 6, 2026 · 1h 9m · neon-show
Where SMALL models will Win | Sudarshan kamath, Smallest ai
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
Sudarshan Kamath, founder of Smallest AI, breaks down how his company built ultra-low-latency real-time voice models using specialized small-model architectures. He shares their rapid progression from a viral bedroom prototype in Bangalore to securing Silicon Valley backing and landing multimillion-dollar enterprise contracts.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →
speaking balance: gold is Siddhartha, purple is the guest (3 minute bins)
Sudarshan flatly rejects the host's classification of voice application startups like Bolna as competitors, clarifying that Smallest AI powers them as infrastructure.
Hardest push from Siddhartha ▶ 42:00 Challenging the bespoke support cost of on-premise deploymentsSiddhartha directly pushes back against Sudarshan's preference for on-prem deployments, arguing it imposes heavy, non-scalable bespoke support overhead.
Biggest teaching moment ▶ 59:40 Explaining hardware memory-bandwidth bottlenecks in voice inferenceSudarshan delivers a technical breakdown of why NVIDIA GPUs bottleneck real-time voice due to physical distance between compute cores and memory.
Siddhartha holds their own ▶ 58:30 Highlighting OpenAI margin vulnerability and AI bubble risksSiddhartha challenges AI growth assumptions by highlighting OpenAI's unknown gross margins and the systemic risk if frontier capital drying up bursts the ecosystem.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | Siddhartha as informed peer | Guest teaching | Guest disagreement | Siddhartha pushing back | Why |
|---|---|---|---|---|---|---|
| Trailer: The Multi-Trillion Dollar Vision of Smallest AI | 4 | 5 | 1 | 1 | The host opens warmly by expressing regret over missing out on investing in Smallest AI and sets up the narrative of breaking Indian ecosystem ceilings. Sudarshan delivers an extensive monologue recounting his background in physics, computer vision, self-driving cars with Arjun Jain, and the power of model distillation. | |
| Domain Inception, Legal Tech Experiments, and Going Full-Time | 3 | 4 | 1 | 1 | Siddhartha asks clarifying chronological questions about the founding timeline. Sudarshan explains his experiments with legal tech in India, early vector database adoption with Pinecone, context engineering, and the conviction to commit personal capital full-time. | |
| Discovering Real-Time Voice AI and the Viral LinkedIn Demo | 4 | 4 | 2 | 1 | The host tracks the fundraising timeline and seed VC metrics. Sudarshan humorously details the viral LinkedIn demo running off a single GPU and pitching VCs in Bangalore with a torn bag and shirt. | |
| Flipping to Delaware: Building a Global Deep-Tech Enterprise | 3 | 5 | 2 | 1 | The host inquires about banking and money transfers. Sudarshan articulates a decisive stance on flipping immediately to a Delaware C-corp to access global capital, talent, and more ambitious frontier AI problems. | |
| Unconventional Hiring and Moving to Silicon Valley | 3 | 4 | 1 | 1 | Siddhartha asks about round extensions and relocation. Sudarshan outlines his non-traditional hiring philosophy via Twitter/LinkedIn and the density of ambition in Silicon Valley compared to other global tech hubs. | |
| Building the Voice AI Infrastructure Portfolio and Pre-Revenue Seed Pitch | 4 | 6 | 2 | 2 | The host explores traction and initial validation metrics. Sudarshan details the full voice AI infrastructure stack needed to support sub-100ms TTS agents, including synthetic data generation and separating compute from infinite memory. | |
| Securing Seed Funding with Sierra Ventures and Enterprise Access | 4 | 4 | 1 | 1 | Siddhartha shares peer anecdotes about Sierra Ventures' dinner-closing style. Sudarshan describes securing their seed lead Tim Gullery and the immediate high-caliber enterprise introductions that followed. | |
| Closing the First Seven-Figure Multi-Year Telephony Contract | 4 | 5 | 2 | 2 | The host probes the contract timeline and conversion to a seven-figure deal. Sudarshan outlines navigating technical grilling by an enterprise CTO and moving from pilot to a multi-year enterprise contract. | |
| Preemptive Series A with Seligman Ventures and Passing on M&A | 4 | 5 | 1 | 1 | The host asks about acquisition interest and secondary sales. Sudarshan explains passing on a $150M+ acquisition offer to close a preemptive Series A with Seligman Ventures and deferring personal secondaries until sales repeatability is proven. | |
| Competitive Landscape: Differentiating Real-Time Voice from ElevenLabs and Cartesia | 4 | 6 | 3 | 2 | Siddhartha pushes on competitors like ElevenLabs and Bolna. Sudarshan firmly distinguishes applications from model providers and clarifies that ElevenLabs focuses on offline content generation while Smallest AI focuses on real-time conversational low-latency models. | |
| Architectural Deep Dive: State Space Models vs. Transformers and JEPA | 5 | 7 | 2 | 2 | The host asks for a layman explanation of architectural approaches. Sudarshan breaks down quadratic vs. subquadratic computation in SSMs and why Smallest AI favors predictive world-model architectures inspired by Yann LeCun's JEPA. | |
| Virtual Private Cloud Deployments and Scaling to $2M ARR | 5 | 6 | 2 | 2 | Siddhartha challenges the operational burden of bespoke on-prem deployments. Sudarshan clarifies that deliveries are containerized in client VPCs, offering 100% software margins and eliminating the GPU autoscaling forecasting headache. | |
| Horizontal Platform Positioning and System Integrator Partnerships | 4 | 5 | 1 | 1 | Siddhartha explores enterprise application builds and partner enablement. Sudarshan explains their horizontal platform strategy and why systems integrators like Accenture and Capgemini are better suited for verticalized enterprise implementations. | |
| Core Research Frontiers: Asynchronous Thinking and Infinite Memory | 4 | 7 | 2 | 1 | The host asks what the doubled research team will focus on. Sudarshan gives an educational breakdown of asynchronous thinking in voice interactions, continuous learning, and separating compute intelligence from infinite memory layers. | |
| The Divergence of LLM Context Scaling and Conversational Intelligence | 5 | 6 | 2 | 2 | Siddhartha presses on why major frontier labs like OpenAI and Anthropic wouldn't capture the entire voice market. Sudarshan uses an airplane-versus-bird analogy to explain that giant context-scaling LLMs represent a different branch of intelligence than agile, low-latency conversational models. | |
| Mathematical Playbook for Scaling from $1M to $100M ARR | 4 | 5 | 1 | 2 | The host asks about capital and timeline requirements to hit $100M ARR. Sudarshan breaks down a sales capacity mathematical model (reps doing $1M-$30M ARR combined with cloud and SI channel partnerships). | |
| Bottom-Up Developer Adoption and Horizontal Voice Use Cases | 4 | 5 | 2 | 2 | The host questions whether enterprise AI adoption is getting clogged by top-down vendor motions. Sudarshan argues that bottom-up developer self-serve adoption creates organic internal pull, citing Cursor as an example. | |
| Evaluating AI Bubble Risks and Hardware Bottlenecks in Voice Inference | 5 | 7 | 3 | 3 | Siddhartha presses on AI bubble risks and OpenAI's unverified gross margins. Sudarshan argues that voice modernizes existing utility markets rather than inventing speculative ones, and explains the physical memory bottleneck in standard NVIDIA GPUs for real-time voice inference. | |
| The Palantir-Style Sales Strategy and Deep Custom Kernel Engineering | 4 | 6 | 1 | 1 | The host asks how large accounts were cracked in detail. Sudarshan shares his Palantir-style account targeting playbook, winning over skeptical public-company CTOs through rapid two-day feature turnarounds, and writing custom NVIDIA kernels. |