Mar 6, 2026 · 1h 9m · neon-show

Where SMALL models will Win | Sudarshan kamath, Smallest ai

Sudarshan Kamath · 58m spoken Siddhartha Ahluwalia · 5m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Sudarshan Kamath, founder of Smallest AI, breaks down how his company built ultra-low-latency real-time voice models using specialized small-model architectures. He shares their rapid progression from a viral bedroom prototype in Bangalore to securing Silicon Valley backing and landing multimillion-dollar enterprise contracts.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

Siddhartha as informed peer 4.0 Guest teaching 5.4 Guest disagreement 1.7 Siddhartha pushing back 1.5
05100:0015:0030:0045:001:00:000:00–6:38 · Siddhartha as informed peer 4/10 Trailer: The Multi-Trillion Dollar Vision of Smallest AI The host opens warmly by expressing regret over missing out on investing in Smallest AI and sets up the narrative of breaking Indian ecosystem ceilings. Sudarshan delivers an extensive monologue recounting his background in physics, computer vision, self-driving cars with Arjun Jain, and the power of model distillation.6:38–9:04 · Siddhartha as informed peer 3/10 Domain Inception, Legal Tech Experiments, and Going Full-Time Siddhartha asks clarifying chronological questions about the founding timeline. Sudarshan explains his experiments with legal tech in India, early vector database adoption with Pinecone, context engineering, and the conviction to commit personal capital full-time.9:04–15:07 · Siddhartha as informed peer 4/10 Discovering Real-Time Voice AI and the Viral LinkedIn Demo The host tracks the fundraising timeline and seed VC metrics. Sudarshan humorously details the viral LinkedIn demo running off a single GPU and pitching VCs in Bangalore with a torn bag and shirt.15:07–17:23 · Siddhartha as informed peer 3/10 Flipping to Delaware: Building a Global Deep-Tech Enterprise The host inquires about banking and money transfers. Sudarshan articulates a decisive stance on flipping immediately to a Delaware C-corp to access global capital, talent, and more ambitious frontier AI problems.17:23–21:09 · Siddhartha as informed peer 3/10 Unconventional Hiring and Moving to Silicon Valley Siddhartha asks about round extensions and relocation. Sudarshan outlines his non-traditional hiring philosophy via Twitter/LinkedIn and the density of ambition in Silicon Valley compared to other global tech hubs.21:09–23:11 · Siddhartha as informed peer 4/10 Building the Voice AI Infrastructure Portfolio and Pre-Revenue Seed Pitch The host explores traction and initial validation metrics. Sudarshan details the full voice AI infrastructure stack needed to support sub-100ms TTS agents, including synthetic data generation and separating compute from infinite memory.23:11–27:18 · Siddhartha as informed peer 4/10 Securing Seed Funding with Sierra Ventures and Enterprise Access Siddhartha shares peer anecdotes about Sierra Ventures' dinner-closing style. Sudarshan describes securing their seed lead Tim Gullery and the immediate high-caliber enterprise introductions that followed.27:18–31:42 · Siddhartha as informed peer 4/10 Closing the First Seven-Figure Multi-Year Telephony Contract The host probes the contract timeline and conversion to a seven-figure deal. Sudarshan outlines navigating technical grilling by an enterprise CTO and moving from pilot to a multi-year enterprise contract.31:43–35:22 · Siddhartha as informed peer 4/10 Preemptive Series A with Seligman Ventures and Passing on M&A The host asks about acquisition interest and secondary sales. Sudarshan explains passing on a $150M+ acquisition offer to close a preemptive Series A with Seligman Ventures and deferring personal secondaries until sales repeatability is proven.35:23–38:27 · Siddhartha as informed peer 4/10 Competitive Landscape: Differentiating Real-Time Voice from ElevenLabs and Cartesia Siddhartha pushes on competitors like ElevenLabs and Bolna. Sudarshan firmly distinguishes applications from model providers and clarifies that ElevenLabs focuses on offline content generation while Smallest AI focuses on real-time conversational low-latency models.38:27–41:22 · Siddhartha as informed peer 5/10 Architectural Deep Dive: State Space Models vs. Transformers and JEPA The host asks for a layman explanation of architectural approaches. Sudarshan breaks down quadratic vs. subquadratic computation in SSMs and why Smallest AI favors predictive world-model architectures inspired by Yann LeCun's JEPA.41:22–43:52 · Siddhartha as informed peer 5/10 Virtual Private Cloud Deployments and Scaling to $2M ARR Siddhartha challenges the operational burden of bespoke on-prem deployments. Sudarshan clarifies that deliveries are containerized in client VPCs, offering 100% software margins and eliminating the GPU autoscaling forecasting headache.43:52–45:54 · Siddhartha as informed peer 4/10 Horizontal Platform Positioning and System Integrator Partnerships Siddhartha explores enterprise application builds and partner enablement. Sudarshan explains their horizontal platform strategy and why systems integrators like Accenture and Capgemini are better suited for verticalized enterprise implementations.45:54–49:04 · Siddhartha as informed peer 4/10 Core Research Frontiers: Asynchronous Thinking and Infinite Memory The host asks what the doubled research team will focus on. Sudarshan gives an educational breakdown of asynchronous thinking in voice interactions, continuous learning, and separating compute intelligence from infinite memory layers.49:04–51:24 · Siddhartha as informed peer 5/10 The Divergence of LLM Context Scaling and Conversational Intelligence Siddhartha presses on why major frontier labs like OpenAI and Anthropic wouldn't capture the entire voice market. Sudarshan uses an airplane-versus-bird analogy to explain that giant context-scaling LLMs represent a different branch of intelligence than agile, low-latency conversational models.51:24–54:11 · Siddhartha as informed peer 4/10 Mathematical Playbook for Scaling from $1M to $100M ARR The host asks about capital and timeline requirements to hit $100M ARR. Sudarshan breaks down a sales capacity mathematical model (reps doing $1M-$30M ARR combined with cloud and SI channel partnerships).54:11–57:09 · Siddhartha as informed peer 4/10 Bottom-Up Developer Adoption and Horizontal Voice Use Cases The host questions whether enterprise AI adoption is getting clogged by top-down vendor motions. Sudarshan argues that bottom-up developer self-serve adoption creates organic internal pull, citing Cursor as an example.57:10–1:00:53 · Siddhartha as informed peer 5/10 Evaluating AI Bubble Risks and Hardware Bottlenecks in Voice Inference Siddhartha presses on AI bubble risks and OpenAI's unverified gross margins. Sudarshan argues that voice modernizes existing utility markets rather than inventing speculative ones, and explains the physical memory bottleneck in standard NVIDIA GPUs for real-time voice inference.1:00:53–1:09:00 · Siddhartha as informed peer 4/10 The Palantir-Style Sales Strategy and Deep Custom Kernel Engineering The host asks how large accounts were cracked in detail. Sudarshan shares his Palantir-style account targeting playbook, winning over skeptical public-company CTOs through rapid two-day feature turnarounds, and writing custom NVIDIA kernels.0:00–6:38 · Guest teaching 5/10 Trailer: The Multi-Trillion Dollar Vision of Smallest AI The host opens warmly by expressing regret over missing out on investing in Smallest AI and sets up the narrative of breaking Indian ecosystem ceilings. Sudarshan delivers an extensive monologue recounting his background in physics, computer vision, self-driving cars with Arjun Jain, and the power of model distillation.6:38–9:04 · Guest teaching 4/10 Domain Inception, Legal Tech Experiments, and Going Full-Time Siddhartha asks clarifying chronological questions about the founding timeline. Sudarshan explains his experiments with legal tech in India, early vector database adoption with Pinecone, context engineering, and the conviction to commit personal capital full-time.9:04–15:07 · Guest teaching 4/10 Discovering Real-Time Voice AI and the Viral LinkedIn Demo The host tracks the fundraising timeline and seed VC metrics. Sudarshan humorously details the viral LinkedIn demo running off a single GPU and pitching VCs in Bangalore with a torn bag and shirt.15:07–17:23 · Guest teaching 5/10 Flipping to Delaware: Building a Global Deep-Tech Enterprise The host inquires about banking and money transfers. Sudarshan articulates a decisive stance on flipping immediately to a Delaware C-corp to access global capital, talent, and more ambitious frontier AI problems.17:23–21:09 · Guest teaching 4/10 Unconventional Hiring and Moving to Silicon Valley Siddhartha asks about round extensions and relocation. Sudarshan outlines his non-traditional hiring philosophy via Twitter/LinkedIn and the density of ambition in Silicon Valley compared to other global tech hubs.21:09–23:11 · Guest teaching 6/10 Building the Voice AI Infrastructure Portfolio and Pre-Revenue Seed Pitch The host explores traction and initial validation metrics. Sudarshan details the full voice AI infrastructure stack needed to support sub-100ms TTS agents, including synthetic data generation and separating compute from infinite memory.23:11–27:18 · Guest teaching 4/10 Securing Seed Funding with Sierra Ventures and Enterprise Access Siddhartha shares peer anecdotes about Sierra Ventures' dinner-closing style. Sudarshan describes securing their seed lead Tim Gullery and the immediate high-caliber enterprise introductions that followed.27:18–31:42 · Guest teaching 5/10 Closing the First Seven-Figure Multi-Year Telephony Contract The host probes the contract timeline and conversion to a seven-figure deal. Sudarshan outlines navigating technical grilling by an enterprise CTO and moving from pilot to a multi-year enterprise contract.31:43–35:22 · Guest teaching 5/10 Preemptive Series A with Seligman Ventures and Passing on M&A The host asks about acquisition interest and secondary sales. Sudarshan explains passing on a $150M+ acquisition offer to close a preemptive Series A with Seligman Ventures and deferring personal secondaries until sales repeatability is proven.35:23–38:27 · Guest teaching 6/10 Competitive Landscape: Differentiating Real-Time Voice from ElevenLabs and Cartesia Siddhartha pushes on competitors like ElevenLabs and Bolna. Sudarshan firmly distinguishes applications from model providers and clarifies that ElevenLabs focuses on offline content generation while Smallest AI focuses on real-time conversational low-latency models.38:27–41:22 · Guest teaching 7/10 Architectural Deep Dive: State Space Models vs. Transformers and JEPA The host asks for a layman explanation of architectural approaches. Sudarshan breaks down quadratic vs. subquadratic computation in SSMs and why Smallest AI favors predictive world-model architectures inspired by Yann LeCun's JEPA.41:22–43:52 · Guest teaching 6/10 Virtual Private Cloud Deployments and Scaling to $2M ARR Siddhartha challenges the operational burden of bespoke on-prem deployments. Sudarshan clarifies that deliveries are containerized in client VPCs, offering 100% software margins and eliminating the GPU autoscaling forecasting headache.43:52–45:54 · Guest teaching 5/10 Horizontal Platform Positioning and System Integrator Partnerships Siddhartha explores enterprise application builds and partner enablement. Sudarshan explains their horizontal platform strategy and why systems integrators like Accenture and Capgemini are better suited for verticalized enterprise implementations.45:54–49:04 · Guest teaching 7/10 Core Research Frontiers: Asynchronous Thinking and Infinite Memory The host asks what the doubled research team will focus on. Sudarshan gives an educational breakdown of asynchronous thinking in voice interactions, continuous learning, and separating compute intelligence from infinite memory layers.49:04–51:24 · Guest teaching 6/10 The Divergence of LLM Context Scaling and Conversational Intelligence Siddhartha presses on why major frontier labs like OpenAI and Anthropic wouldn't capture the entire voice market. Sudarshan uses an airplane-versus-bird analogy to explain that giant context-scaling LLMs represent a different branch of intelligence than agile, low-latency conversational models.51:24–54:11 · Guest teaching 5/10 Mathematical Playbook for Scaling from $1M to $100M ARR The host asks about capital and timeline requirements to hit $100M ARR. Sudarshan breaks down a sales capacity mathematical model (reps doing $1M-$30M ARR combined with cloud and SI channel partnerships).54:11–57:09 · Guest teaching 5/10 Bottom-Up Developer Adoption and Horizontal Voice Use Cases The host questions whether enterprise AI adoption is getting clogged by top-down vendor motions. Sudarshan argues that bottom-up developer self-serve adoption creates organic internal pull, citing Cursor as an example.57:10–1:00:53 · Guest teaching 7/10 Evaluating AI Bubble Risks and Hardware Bottlenecks in Voice Inference Siddhartha presses on AI bubble risks and OpenAI's unverified gross margins. Sudarshan argues that voice modernizes existing utility markets rather than inventing speculative ones, and explains the physical memory bottleneck in standard NVIDIA GPUs for real-time voice inference.1:00:53–1:09:00 · Guest teaching 6/10 The Palantir-Style Sales Strategy and Deep Custom Kernel Engineering The host asks how large accounts were cracked in detail. Sudarshan shares his Palantir-style account targeting playbook, winning over skeptical public-company CTOs through rapid two-day feature turnarounds, and writing custom NVIDIA kernels.0:00–6:38 · Guest disagreement 1/10 Trailer: The Multi-Trillion Dollar Vision of Smallest AI The host opens warmly by expressing regret over missing out on investing in Smallest AI and sets up the narrative of breaking Indian ecosystem ceilings. Sudarshan delivers an extensive monologue recounting his background in physics, computer vision, self-driving cars with Arjun Jain, and the power of model distillation.6:38–9:04 · Guest disagreement 1/10 Domain Inception, Legal Tech Experiments, and Going Full-Time Siddhartha asks clarifying chronological questions about the founding timeline. Sudarshan explains his experiments with legal tech in India, early vector database adoption with Pinecone, context engineering, and the conviction to commit personal capital full-time.9:04–15:07 · Guest disagreement 2/10 Discovering Real-Time Voice AI and the Viral LinkedIn Demo The host tracks the fundraising timeline and seed VC metrics. Sudarshan humorously details the viral LinkedIn demo running off a single GPU and pitching VCs in Bangalore with a torn bag and shirt.15:07–17:23 · Guest disagreement 2/10 Flipping to Delaware: Building a Global Deep-Tech Enterprise The host inquires about banking and money transfers. Sudarshan articulates a decisive stance on flipping immediately to a Delaware C-corp to access global capital, talent, and more ambitious frontier AI problems.17:23–21:09 · Guest disagreement 1/10 Unconventional Hiring and Moving to Silicon Valley Siddhartha asks about round extensions and relocation. Sudarshan outlines his non-traditional hiring philosophy via Twitter/LinkedIn and the density of ambition in Silicon Valley compared to other global tech hubs.21:09–23:11 · Guest disagreement 2/10 Building the Voice AI Infrastructure Portfolio and Pre-Revenue Seed Pitch The host explores traction and initial validation metrics. Sudarshan details the full voice AI infrastructure stack needed to support sub-100ms TTS agents, including synthetic data generation and separating compute from infinite memory.23:11–27:18 · Guest disagreement 1/10 Securing Seed Funding with Sierra Ventures and Enterprise Access Siddhartha shares peer anecdotes about Sierra Ventures' dinner-closing style. Sudarshan describes securing their seed lead Tim Gullery and the immediate high-caliber enterprise introductions that followed.27:18–31:42 · Guest disagreement 2/10 Closing the First Seven-Figure Multi-Year Telephony Contract The host probes the contract timeline and conversion to a seven-figure deal. Sudarshan outlines navigating technical grilling by an enterprise CTO and moving from pilot to a multi-year enterprise contract.31:43–35:22 · Guest disagreement 1/10 Preemptive Series A with Seligman Ventures and Passing on M&A The host asks about acquisition interest and secondary sales. Sudarshan explains passing on a $150M+ acquisition offer to close a preemptive Series A with Seligman Ventures and deferring personal secondaries until sales repeatability is proven.35:23–38:27 · Guest disagreement 3/10 Competitive Landscape: Differentiating Real-Time Voice from ElevenLabs and Cartesia Siddhartha pushes on competitors like ElevenLabs and Bolna. Sudarshan firmly distinguishes applications from model providers and clarifies that ElevenLabs focuses on offline content generation while Smallest AI focuses on real-time conversational low-latency models.38:27–41:22 · Guest disagreement 2/10 Architectural Deep Dive: State Space Models vs. Transformers and JEPA The host asks for a layman explanation of architectural approaches. Sudarshan breaks down quadratic vs. subquadratic computation in SSMs and why Smallest AI favors predictive world-model architectures inspired by Yann LeCun's JEPA.41:22–43:52 · Guest disagreement 2/10 Virtual Private Cloud Deployments and Scaling to $2M ARR Siddhartha challenges the operational burden of bespoke on-prem deployments. Sudarshan clarifies that deliveries are containerized in client VPCs, offering 100% software margins and eliminating the GPU autoscaling forecasting headache.43:52–45:54 · Guest disagreement 1/10 Horizontal Platform Positioning and System Integrator Partnerships Siddhartha explores enterprise application builds and partner enablement. Sudarshan explains their horizontal platform strategy and why systems integrators like Accenture and Capgemini are better suited for verticalized enterprise implementations.45:54–49:04 · Guest disagreement 2/10 Core Research Frontiers: Asynchronous Thinking and Infinite Memory The host asks what the doubled research team will focus on. Sudarshan gives an educational breakdown of asynchronous thinking in voice interactions, continuous learning, and separating compute intelligence from infinite memory layers.49:04–51:24 · Guest disagreement 2/10 The Divergence of LLM Context Scaling and Conversational Intelligence Siddhartha presses on why major frontier labs like OpenAI and Anthropic wouldn't capture the entire voice market. Sudarshan uses an airplane-versus-bird analogy to explain that giant context-scaling LLMs represent a different branch of intelligence than agile, low-latency conversational models.51:24–54:11 · Guest disagreement 1/10 Mathematical Playbook for Scaling from $1M to $100M ARR The host asks about capital and timeline requirements to hit $100M ARR. Sudarshan breaks down a sales capacity mathematical model (reps doing $1M-$30M ARR combined with cloud and SI channel partnerships).54:11–57:09 · Guest disagreement 2/10 Bottom-Up Developer Adoption and Horizontal Voice Use Cases The host questions whether enterprise AI adoption is getting clogged by top-down vendor motions. Sudarshan argues that bottom-up developer self-serve adoption creates organic internal pull, citing Cursor as an example.57:10–1:00:53 · Guest disagreement 3/10 Evaluating AI Bubble Risks and Hardware Bottlenecks in Voice Inference Siddhartha presses on AI bubble risks and OpenAI's unverified gross margins. Sudarshan argues that voice modernizes existing utility markets rather than inventing speculative ones, and explains the physical memory bottleneck in standard NVIDIA GPUs for real-time voice inference.1:00:53–1:09:00 · Guest disagreement 1/10 The Palantir-Style Sales Strategy and Deep Custom Kernel Engineering The host asks how large accounts were cracked in detail. Sudarshan shares his Palantir-style account targeting playbook, winning over skeptical public-company CTOs through rapid two-day feature turnarounds, and writing custom NVIDIA kernels.0:00–6:38 · Siddhartha pushing back 1/10 Trailer: The Multi-Trillion Dollar Vision of Smallest AI The host opens warmly by expressing regret over missing out on investing in Smallest AI and sets up the narrative of breaking Indian ecosystem ceilings. Sudarshan delivers an extensive monologue recounting his background in physics, computer vision, self-driving cars with Arjun Jain, and the power of model distillation.6:38–9:04 · Siddhartha pushing back 1/10 Domain Inception, Legal Tech Experiments, and Going Full-Time Siddhartha asks clarifying chronological questions about the founding timeline. Sudarshan explains his experiments with legal tech in India, early vector database adoption with Pinecone, context engineering, and the conviction to commit personal capital full-time.9:04–15:07 · Siddhartha pushing back 1/10 Discovering Real-Time Voice AI and the Viral LinkedIn Demo The host tracks the fundraising timeline and seed VC metrics. Sudarshan humorously details the viral LinkedIn demo running off a single GPU and pitching VCs in Bangalore with a torn bag and shirt.15:07–17:23 · Siddhartha pushing back 1/10 Flipping to Delaware: Building a Global Deep-Tech Enterprise The host inquires about banking and money transfers. Sudarshan articulates a decisive stance on flipping immediately to a Delaware C-corp to access global capital, talent, and more ambitious frontier AI problems.17:23–21:09 · Siddhartha pushing back 1/10 Unconventional Hiring and Moving to Silicon Valley Siddhartha asks about round extensions and relocation. Sudarshan outlines his non-traditional hiring philosophy via Twitter/LinkedIn and the density of ambition in Silicon Valley compared to other global tech hubs.21:09–23:11 · Siddhartha pushing back 2/10 Building the Voice AI Infrastructure Portfolio and Pre-Revenue Seed Pitch The host explores traction and initial validation metrics. Sudarshan details the full voice AI infrastructure stack needed to support sub-100ms TTS agents, including synthetic data generation and separating compute from infinite memory.23:11–27:18 · Siddhartha pushing back 1/10 Securing Seed Funding with Sierra Ventures and Enterprise Access Siddhartha shares peer anecdotes about Sierra Ventures' dinner-closing style. Sudarshan describes securing their seed lead Tim Gullery and the immediate high-caliber enterprise introductions that followed.27:18–31:42 · Siddhartha pushing back 2/10 Closing the First Seven-Figure Multi-Year Telephony Contract The host probes the contract timeline and conversion to a seven-figure deal. Sudarshan outlines navigating technical grilling by an enterprise CTO and moving from pilot to a multi-year enterprise contract.31:43–35:22 · Siddhartha pushing back 1/10 Preemptive Series A with Seligman Ventures and Passing on M&A The host asks about acquisition interest and secondary sales. Sudarshan explains passing on a $150M+ acquisition offer to close a preemptive Series A with Seligman Ventures and deferring personal secondaries until sales repeatability is proven.35:23–38:27 · Siddhartha pushing back 2/10 Competitive Landscape: Differentiating Real-Time Voice from ElevenLabs and Cartesia Siddhartha pushes on competitors like ElevenLabs and Bolna. Sudarshan firmly distinguishes applications from model providers and clarifies that ElevenLabs focuses on offline content generation while Smallest AI focuses on real-time conversational low-latency models.38:27–41:22 · Siddhartha pushing back 2/10 Architectural Deep Dive: State Space Models vs. Transformers and JEPA The host asks for a layman explanation of architectural approaches. Sudarshan breaks down quadratic vs. subquadratic computation in SSMs and why Smallest AI favors predictive world-model architectures inspired by Yann LeCun's JEPA.41:22–43:52 · Siddhartha pushing back 2/10 Virtual Private Cloud Deployments and Scaling to $2M ARR Siddhartha challenges the operational burden of bespoke on-prem deployments. Sudarshan clarifies that deliveries are containerized in client VPCs, offering 100% software margins and eliminating the GPU autoscaling forecasting headache.43:52–45:54 · Siddhartha pushing back 1/10 Horizontal Platform Positioning and System Integrator Partnerships Siddhartha explores enterprise application builds and partner enablement. Sudarshan explains their horizontal platform strategy and why systems integrators like Accenture and Capgemini are better suited for verticalized enterprise implementations.45:54–49:04 · Siddhartha pushing back 1/10 Core Research Frontiers: Asynchronous Thinking and Infinite Memory The host asks what the doubled research team will focus on. Sudarshan gives an educational breakdown of asynchronous thinking in voice interactions, continuous learning, and separating compute intelligence from infinite memory layers.49:04–51:24 · Siddhartha pushing back 2/10 The Divergence of LLM Context Scaling and Conversational Intelligence Siddhartha presses on why major frontier labs like OpenAI and Anthropic wouldn't capture the entire voice market. Sudarshan uses an airplane-versus-bird analogy to explain that giant context-scaling LLMs represent a different branch of intelligence than agile, low-latency conversational models.51:24–54:11 · Siddhartha pushing back 2/10 Mathematical Playbook for Scaling from $1M to $100M ARR The host asks about capital and timeline requirements to hit $100M ARR. Sudarshan breaks down a sales capacity mathematical model (reps doing $1M-$30M ARR combined with cloud and SI channel partnerships).54:11–57:09 · Siddhartha pushing back 2/10 Bottom-Up Developer Adoption and Horizontal Voice Use Cases The host questions whether enterprise AI adoption is getting clogged by top-down vendor motions. Sudarshan argues that bottom-up developer self-serve adoption creates organic internal pull, citing Cursor as an example.57:10–1:00:53 · Siddhartha pushing back 3/10 Evaluating AI Bubble Risks and Hardware Bottlenecks in Voice Inference Siddhartha presses on AI bubble risks and OpenAI's unverified gross margins. Sudarshan argues that voice modernizes existing utility markets rather than inventing speculative ones, and explains the physical memory bottleneck in standard NVIDIA GPUs for real-time voice inference.1:00:53–1:09:00 · Siddhartha pushing back 1/10 The Palantir-Style Sales Strategy and Deep Custom Kernel Engineering The host asks how large accounts were cracked in detail. Sudarshan shares his Palantir-style account targeting playbook, winning over skeptical public-company CTOs through rapid two-day feature turnarounds, and writing custom NVIDIA kernels.

speaking balance: gold is Siddhartha, purple is the guest (3 minute bins)

0:00 · Siddhartha 0% · guest 100%0:00 · Siddhartha 0% · guest 100%3:00 · Siddhartha 0% · guest 100%3:00 · Siddhartha 0% · guest 100%6:00 · Siddhartha 0% · guest 100%6:00 · Siddhartha 0% · guest 100%9:00 · Siddhartha 0% · guest 100%9:00 · Siddhartha 0% · guest 100%12:00 · Siddhartha 0% · guest 100%12:00 · Siddhartha 0% · guest 100%15:00 · Siddhartha 0% · guest 100%15:00 · Siddhartha 0% · guest 100%18:00 · Siddhartha 0% · guest 100%18:00 · Siddhartha 0% · guest 100%21:00 · Siddhartha 0% · guest 100%21:00 · Siddhartha 0% · guest 100%24:00 · Siddhartha 0% · guest 100%24:00 · Siddhartha 0% · guest 100%27:00 · Siddhartha 0% · guest 100%27:00 · Siddhartha 0% · guest 100%30:00 · Siddhartha 0% · guest 100%30:00 · Siddhartha 0% · guest 100%33:00 · Siddhartha 0% · guest 100%33:00 · Siddhartha 0% · guest 100%36:00 · Siddhartha 0% · guest 100%36:00 · Siddhartha 0% · guest 100%39:00 · Siddhartha 0% · guest 100%39:00 · Siddhartha 0% · guest 100%42:00 · Siddhartha 0% · guest 100%42:00 · Siddhartha 0% · guest 100%45:00 · Siddhartha 0% · guest 100%45:00 · Siddhartha 0% · guest 100%48:00 · Siddhartha 0% · guest 100%48:00 · Siddhartha 0% · guest 100%51:00 · Siddhartha 0% · guest 100%51:00 · Siddhartha 0% · guest 100%54:00 · Siddhartha 0% · guest 100%54:00 · Siddhartha 0% · guest 100%57:00 · Siddhartha 0% · guest 100%57:00 · Siddhartha 0% · guest 100%1:00:00 · Siddhartha 0% · guest 100%1:00:00 · Siddhartha 0% · guest 100%1:03:00 · Siddhartha 0% · guest 100%1:03:00 · Siddhartha 0% · guest 100%1:06:00 · Siddhartha 0% · guest 100%1:06:00 · Siddhartha 0% · guest 100%1:09:00 · Siddhartha 0% · guest 100%1:09:00 · Siddhartha 0% · guest 100%
Sharpest disagreement ▶ 36:20 Rejecting application-level competitor categorization

Sudarshan flatly rejects the host's classification of voice application startups like Bolna as competitors, clarifying that Smallest AI powers them as infrastructure.

Hardest push from Siddhartha ▶ 42:00 Challenging the bespoke support cost of on-premise deployments

Siddhartha directly pushes back against Sudarshan's preference for on-prem deployments, arguing it imposes heavy, non-scalable bespoke support overhead.

Biggest teaching moment ▶ 59:40 Explaining hardware memory-bandwidth bottlenecks in voice inference

Sudarshan delivers a technical breakdown of why NVIDIA GPUs bottleneck real-time voice due to physical distance between compute cores and memory.

Siddhartha holds their own ▶ 58:30 Highlighting OpenAI margin vulnerability and AI bubble risks

Siddhartha challenges AI growth assumptions by highlighting OpenAI's unknown gross margins and the systemic risk if frontier capital drying up bursts the ecosystem.

the scores for every segment, with the reasoning behind each
ChapterTopicSiddhartha as informed peerGuest teachingGuest disagreementSiddhartha pushing backWhy
Trailer: The Multi-Trillion Dollar Vision of Smallest AI 4511 The host opens warmly by expressing regret over missing out on investing in Smallest AI and sets up the narrative of breaking Indian ecosystem ceilings. Sudarshan delivers an extensive monologue recounting his background in physics, computer vision, self-driving cars with Arjun Jain, and the power of model distillation.
Domain Inception, Legal Tech Experiments, and Going Full-Time 3411 Siddhartha asks clarifying chronological questions about the founding timeline. Sudarshan explains his experiments with legal tech in India, early vector database adoption with Pinecone, context engineering, and the conviction to commit personal capital full-time.
Discovering Real-Time Voice AI and the Viral LinkedIn Demo 4421 The host tracks the fundraising timeline and seed VC metrics. Sudarshan humorously details the viral LinkedIn demo running off a single GPU and pitching VCs in Bangalore with a torn bag and shirt.
Flipping to Delaware: Building a Global Deep-Tech Enterprise 3521 The host inquires about banking and money transfers. Sudarshan articulates a decisive stance on flipping immediately to a Delaware C-corp to access global capital, talent, and more ambitious frontier AI problems.
Unconventional Hiring and Moving to Silicon Valley 3411 Siddhartha asks about round extensions and relocation. Sudarshan outlines his non-traditional hiring philosophy via Twitter/LinkedIn and the density of ambition in Silicon Valley compared to other global tech hubs.
Building the Voice AI Infrastructure Portfolio and Pre-Revenue Seed Pitch 4622 The host explores traction and initial validation metrics. Sudarshan details the full voice AI infrastructure stack needed to support sub-100ms TTS agents, including synthetic data generation and separating compute from infinite memory.
Securing Seed Funding with Sierra Ventures and Enterprise Access 4411 Siddhartha shares peer anecdotes about Sierra Ventures' dinner-closing style. Sudarshan describes securing their seed lead Tim Gullery and the immediate high-caliber enterprise introductions that followed.
Closing the First Seven-Figure Multi-Year Telephony Contract 4522 The host probes the contract timeline and conversion to a seven-figure deal. Sudarshan outlines navigating technical grilling by an enterprise CTO and moving from pilot to a multi-year enterprise contract.
Preemptive Series A with Seligman Ventures and Passing on M&A 4511 The host asks about acquisition interest and secondary sales. Sudarshan explains passing on a $150M+ acquisition offer to close a preemptive Series A with Seligman Ventures and deferring personal secondaries until sales repeatability is proven.
Competitive Landscape: Differentiating Real-Time Voice from ElevenLabs and Cartesia 4632 Siddhartha pushes on competitors like ElevenLabs and Bolna. Sudarshan firmly distinguishes applications from model providers and clarifies that ElevenLabs focuses on offline content generation while Smallest AI focuses on real-time conversational low-latency models.
Architectural Deep Dive: State Space Models vs. Transformers and JEPA 5722 The host asks for a layman explanation of architectural approaches. Sudarshan breaks down quadratic vs. subquadratic computation in SSMs and why Smallest AI favors predictive world-model architectures inspired by Yann LeCun's JEPA.
Virtual Private Cloud Deployments and Scaling to $2M ARR 5622 Siddhartha challenges the operational burden of bespoke on-prem deployments. Sudarshan clarifies that deliveries are containerized in client VPCs, offering 100% software margins and eliminating the GPU autoscaling forecasting headache.
Horizontal Platform Positioning and System Integrator Partnerships 4511 Siddhartha explores enterprise application builds and partner enablement. Sudarshan explains their horizontal platform strategy and why systems integrators like Accenture and Capgemini are better suited for verticalized enterprise implementations.
Core Research Frontiers: Asynchronous Thinking and Infinite Memory 4721 The host asks what the doubled research team will focus on. Sudarshan gives an educational breakdown of asynchronous thinking in voice interactions, continuous learning, and separating compute intelligence from infinite memory layers.
The Divergence of LLM Context Scaling and Conversational Intelligence 5622 Siddhartha presses on why major frontier labs like OpenAI and Anthropic wouldn't capture the entire voice market. Sudarshan uses an airplane-versus-bird analogy to explain that giant context-scaling LLMs represent a different branch of intelligence than agile, low-latency conversational models.
Mathematical Playbook for Scaling from $1M to $100M ARR 4512 The host asks about capital and timeline requirements to hit $100M ARR. Sudarshan breaks down a sales capacity mathematical model (reps doing $1M-$30M ARR combined with cloud and SI channel partnerships).
Bottom-Up Developer Adoption and Horizontal Voice Use Cases 4522 The host questions whether enterprise AI adoption is getting clogged by top-down vendor motions. Sudarshan argues that bottom-up developer self-serve adoption creates organic internal pull, citing Cursor as an example.
Evaluating AI Bubble Risks and Hardware Bottlenecks in Voice Inference 5733 Siddhartha presses on AI bubble risks and OpenAI's unverified gross margins. Sudarshan argues that voice modernizes existing utility markets rather than inventing speculative ones, and explains the physical memory bottleneck in standard NVIDIA GPUs for real-time voice inference.
The Palantir-Style Sales Strategy and Deep Custom Kernel Engineering 4611 The host asks how large accounts were cracked in detail. Sudarshan shares his Palantir-style account targeting playbook, winning over skeptical public-company CTOs through rapid two-day feature turnarounds, and writing custom NVIDIA kernels.

Statements from this episode (41)

Assertion Supported
Kamath: Byju's acquired Toppr for $150 million
“We got acquired by juice for like one, fifty million dollars back then.”
Sudarshan Kamath Mar 6, 2026 ▶ 4:03
Prediction Open · timeframe Mar 2029
Kamath: Audio content platform Kuku FM will IPO very soon
“So Cuckoo is gonna do an IPO very soon.”
Sudarshan Kamath Mar 6, 2026 ▶ 4:42
Insight
Kamath: Distilled AI models match large model accuracy at a fraction of size
“When you use that pre-training model as a teacher for this small model, and we could go into what teacher means, but it would basically Act at similar accuracies as the larger model at 100 or one 10th of the size.”
Sudarshan Kamath Mar 6, 2026 ▶ 6:11
Opinion
Kamath: Legal tech startups must target the US market to survive
“If you have to build for the legal market, you have to build from the U S basically, or for the U S at least, right? Every, everywhere else is the market is just too tiny for you to get any money out of that.”
Sudarshan Kamath Mar 6, 2026 ▶ 7:45
Disclosure
Kamath: Founder bootstrapped Smallest AI with $150K of his own money
“I put the first 101 50 K into smallest of my own. And we basically got like some GPUs and we started like training models.”
Sudarshan Kamath Mar 6, 2026 ▶ 8:52
Assertion Partly supported
Kamath: Smallest AI built sub-200ms real-time TTS before ElevenLabs
“11 Labs was completely into offline dubbing and things like that. They did not have a real time model, right? So we were one of the first people to build a real time text-to-speech that ran under 200 milliseconds.”
Sudarshan Kamath Mar 6, 2026 ▶ 9:54
Disclosure
Kamath: Smallest AI raised initial pre-seed checks at a $3M valuation
“Upsparks was like, not much like a hundred K if I am not wrong. Better was another two 50 or so. The total round was three 50 K and then couple of angels of 400 or something like that. This is at like a three million post money sort of around, right?”
Sudarshan Kamath Mar 6, 2026 ▶ 13:41
Disclosure
Kamath: Smallest AI reduced 3one4 Capital's $1M check to prevent dilution
“We asked them to reduce the check size. We said, we can't take one at six. It's like too much dilution for us. So we, I think we ended up at like some 500 or six or something.”
Sudarshan Kamath Mar 6, 2026 ▶ 14:44
Opinion
Kamath: Indian language AI is an operational dataset problem, not scientific
“Indian languages for me was like just an operational problem to solve getting data sets, et cetera. But it was not a scientific enough problem that helps us answer what's like the next level of AI and that how can small models be the future, et cetera.”
Sudarshan Kamath Mar 6, 2026 ▶ 16:34
Assertion Not publicly verifiable
Kamath: Smallest AI raised $900K extension, totaling $1.8M
“We raised another 900 K extension at a higher valuation. And we had like 1.8 of which most of it was in the bank.”
Sudarshan Kamath Mar 6, 2026 ▶ 19:55
Assertion Not checkable as stated
Kamath: Smallest AI cut TTS latency to sub-100ms by seed round
“So we had a text to speech model which basically you could use to give natural realistic voice to the voice agent in under a hundred milliseconds. So initially it was at 200 milliseconds. It came down to a hundred milliseconds by the time our seed round was th…”
Sudarshan Kamath Mar 6, 2026 ▶ 21:21
Disclosure
Kamath: Smallest AI is working with ServiceNow
“For example service now, for example, Right. We are working with them.”
Sudarshan Kamath Mar 6, 2026 ▶ 26:21
Assertion Not checkable as stated
Kamath: Smallest AI signed a seven-figure US telecommunications enterprise deal
“It is converted to a seven figure multi-year multimillion deal. But we will announce it formally sometime like in April or so once we sort of clear a few things, but yeah, it is that like all of those things are in paper right now.”
Sudarshan Kamath Mar 6, 2026 ▶ 29:08
Opinion
Kamath: Only ~20 Indian companies can offer seven-figure enterprise deals
“There are probably 20 companies in India who can give us that kind of a deal size and they would spend an incredible amount of time and effort.”
Sudarshan Kamath Mar 6, 2026 ▶ 31:33
Disclosure
Kamath: Smallest AI closed Series A round with Seligman Ventures
“We closed the series A, we could call it a pre-series A, but like we closed with the Seligman Ventures.”
Sudarshan Kamath Mar 6, 2026 ▶ 31:46
Assertion Not checkable as stated
Kamath: Smallest AI received an acquisition offer over $150M
“We got in between like an acquisition offer for like one 50 plus million dollars.”
Sudarshan Kamath Mar 6, 2026 ▶ 32:13
Assertion Supported
Kamath: AWS is an active reseller of Smallest AI
“AWS is an active reseller of smallest by the way.”
Sudarshan Kamath Mar 6, 2026 ▶ 33:08
Opinion
Kamath: ElevenLabs is focused on content creation and dubbing, not real-time
“Like for example, 11 has a lot of focus on content creation and like dubbing. And if you want to build like an Instagram video or like a AI movie 11 would be a great player. We don't focus on content at all. We are focused on the real time space.”
Sudarshan Kamath Mar 6, 2026 ▶ 38:02
Assertion Not checkable as stated
Kamath: Cartesia is the only other player focused on real-time voice
“Right now, the only other player who's focused on real time is Cartesia.”
Sudarshan Kamath Mar 6, 2026 ▶ 38:19
Assertion Supported
Kamath: Cartesia relies on state space models, Smallest AI does not
“But they are heavily dependent on state space models and we don't train state space models.”
Sudarshan Kamath Mar 6, 2026 ▶ 38:23
Opinion
Kamath: State space models are not a big leap over transformers
“Our strong belief is like state space models are not like that big a leap over transformers.”
Sudarshan Kamath Mar 6, 2026 ▶ 39:42
Opinion
Kamath: JEPA architectures are closer to AGI than state space models
“In our opinion, like those models are actually closer to how we can build AGI than like state space models, which is just a, Play on compute capacity.”
Sudarshan Kamath Mar 6, 2026 ▶ 40:14
Disclosure
Kamath: Smallest AI is building a predictive architecture for voice
“We are trying to build like a predictive view of voice and like conversations that we can have so that you can guess what the answer should be rather than How LLMs typically operate.”
Sudarshan Kamath Mar 6, 2026 ▶ 41:06
Disclosure
Kamath: Smallest AI achieves 100% margins via customer-managed VPC deployments
“We recommend it because for us it's a hundred percent margins, right? If that happens.”
Sudarshan Kamath Mar 6, 2026 ▶ 41:56
Insight
Kamath: AI vendors overprice models to offset GPU auto-scaling costs
“And the biggest issue is today people end up pricing the models extremely expensively because they have to manage the scale up of GP capacity.”
Sudarshan Kamath Mar 6, 2026 ▶ 42:50
Disclosure
Kamath: Smallest AI generates between $1M and $2M in ARR
“It's between one to two million ARR.”
Sudarshan Kamath Mar 6, 2026 ▶ 43:36
Insight
Kamath: AI research hires take six months to show ROI
“Research gives you ROI in like six months. So you can't like hire someone right now and then expect, if you have goals for 2026, like in September, you have to get those people right now.”
Sudarshan Kamath Mar 6, 2026 ▶ 45:43
Prediction Not checkable as stated
Kamath: Future AI will decouple infinite memory from small-model compute
“The way we see the future is you will separate memory from intelligence. So the memory will be captured by infinite layers and the intelligence will be captured by finite layers of, you know, intelligence and that small model that has those finite layers that …”
Sudarshan Kamath Mar 6, 2026 ▶ 47:32
Prediction Not checkable as stated
Kamath: Voice AI will diverge and is definitely not winner-takes-all
“I think the markets will diverge at some point of time. Like so for example, 11 is sort of doubling down on the content space but the real time space, like mimicking how humans do conversations is, it's like a very different problem than mimicking how movies a…”
Sudarshan Kamath Mar 6, 2026 ▶ 48:12
Disclosure
Kamath: Smallest AI raised pre-seed funding from OpenAI personnel
“So some of our investors in like our pre-seed words from OpenAI as I spoke to them, they are building a different kind of intelligence.”
Sudarshan Kamath Mar 6, 2026 ▶ 50:22
Insight
Kamath: Scaling LLMs has diverged from building human-like AI
“Very similarly, right now in AI, we are doing LLMs, which is a different form of intelligence than human intelligence. So they have something and they are playing the sort of let's make LLMs bigger and bigger and smarter. And that is a market as well. But To b…”
Sudarshan Kamath Mar 6, 2026 ▶ 50:56
Prediction Not checkable as stated
Kamath: Smallest AI has enough capital to reach $100M ARR
“Having said that, I still think the capital we have right now is enough to build a hundred million ARR company.”
Sudarshan Kamath Mar 6, 2026 ▶ 52:03
Disclosure
Kamath: Indian financial firm runs 100K+ daily debt calls on Smallest AI
“There is a financial services company based in India, which is doing large scale debt collections using us, like, you know, every day we are like at a 100,000 calls plus using that.”
Sudarshan Kamath Mar 6, 2026 ▶ 56:09
Disclosure
Kamath: US telephony firm records internal meetings with Smallest AI
“There is a telephony company that has all of their like meetings internal meetings, getting like recorded in the US using our voice models.”
Sudarshan Kamath Mar 6, 2026 ▶ 56:25
Prediction Not checkable as stated
Kamath: Most coding agent startups will die or face fire sales
“There are a lot of companies that are just trying to build these like very specific coding agents, which apparently will be like better than that's not going to happen. Right. So I feel like there are like, Few winners in that space. And like the others are ju…”
Sudarshan Kamath Mar 6, 2026 ▶ 57:38
Insight
Kamath: Voice AI expands established markets instead of inventing new ones
“The voice space is a existing market that is getting disrupted. You're not inventing a market. You already have a dictation in mobile apps, which sucks. You already have that desktop note, which was really bad. Now with generative AI, you have much better dict…”
Sudarshan Kamath Mar 6, 2026 ▶ 57:58
Assertion Not checkable as stated
Kamath: Nvidia GPUs suffer severe memory bottlenecks during voice inference
“And media chips are generally very bad for voice inference. Like if you have 40 gigabytes of RAM in a NVIDIA GPU, you can generally, if you want to do real time, like a hundred milliseconds, you can use maybe two or four gigabytes of it for our models. And bey…”
Sudarshan Kamath Mar 6, 2026 ▶ 59:44
Assertion Contradicted
Kamath: Voice AI inference currently costs 1,000x more than text
“Today, the cost of voice is like a thousand times more than the cost of text. We want to like bring it at par basically.”
Sudarshan Kamath Mar 6, 2026 ▶ 1:00:41
Disclosure
Kamath: Smallest AI targets top 50 accounts via Palantir-style sales
“We are literally creating the like list of top 50 dream customers that we have and going aggressively through like a top down motion, a very targeted motion of how could we, you know, get them through a Palantir style approach.”
Sudarshan Kamath Mar 6, 2026 ▶ 1:01:04
Assertion Supported
Kamath: Smallest AI's Pulse speech-to-text operates at 64ms latency
“So we have this amazing voice model that converts voice to text called pulse. It's a speech to text model runs in real time 64 milliseconds latency lot of other features like emotion detection, gender detection, bunch of other things, right?”
Sudarshan Kamath Mar 6, 2026 ▶ 1:02:43
Disclosure
Kamath: Smallest AI writes custom GPU kernels for NVIDIA hardware
“And we go right to the level of The NVIDIA architecture to even write our own kernels. So we are not scared of touching any part of the stack.”
Sudarshan Kamath Mar 6, 2026 ▶ 1:08:39
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.