Jun 9, 2026 · 51m · neon-show
94% CAGR: What the Inference Boom means for your AI costs | Vamshi Ambati
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this episode of The Neon Show, host Siddharth Ahluwalia interviews AI researcher and entrepreneur Vamshi Ambati to analyze the explosive 94% CAGR growth of the inference market, shifting token economics, and strategic playbooks for scaling enterprise AI startups.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →
speaking balance: gold is Siddhartha, purple is the guest (3 minute bins)
The guest directly challenges the host's premise by arguing that compute access remains highly restricted and reasoning token costs will inflate building expenses.
Hardest push from Siddhartha ▶ 13:38 Host challenges guest on 100x compute cost reductionThe host explicitly pushes back against the guest's thesis of rising token costs by citing industry-wide initiatives aimed at making compute 100x cheaper.
Biggest teaching moment ▶ 19:42 Guest provides definitive neural inference breakdownThe guest delivers a foundational explanation of inference versus training, mapping biological neural maturation to LLM parameter forward-pass computations.
Siddhartha holds their own ▶ 8:36 Host uses Nutanix cloud-to-on-prem market share analogyThe host demonstrates deep domain knowledge by invoking historical on-prem survival rates and Nutanix's creation to test the model layer's true enterprise penetration ceiling.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | Siddhartha as informed peer | Guest teaching | Guest disagreement | Siddhartha pushing back | Why |
|---|---|---|---|---|---|---|
| The Three Waves of AI and Unprecedented Innovation Pace | 3 | 6 | 1 | 1 | The host opens with a broad macro question about the AI landscape. The guest provides a comprehensive 20-year retrospective detailing the three distinct waves of AI from symbolic to statistical to neural. | |
| Enterprise AI Adoption, Trust, and SaaS Disruption | 5 | 5 | 2 | 3 | The host utilizes historical context about on-prem versus cloud adoption and mentions Nutanix to query the model layer's market share. The guest acknowledges the host's intuition before explaining why enterprise trust and reliability will determine the pace of diffusion. | |
| The Economics of Token Pricing, Hardware, and Reasoning Costs | 6 | 6 | 3 | 5 | The guest questions the host's assertion that building software has become cheap due to token and compute limits. The host pushes back by pointing out that predictions of rising model costs run counter to industry efforts to make compute 100x cheaper. | |
| The 94% CAGR Inference Market and Defining Inference | 6 | 6 | 1 | 1 | The host cites AWS revenue figures to frame a question on whether inference will outgrow compute spending. The guest delivers market forecast data showing a 94% CAGR and provides a detailed analogy comparing training and inference to human brain development. | |
| Core AI Verticals: Deterministic Code vs. Human Customer Voice | 5 | 5 | 3 | 4 | The guest breaks down the deterministic advantages of coding and expresses skepticism regarding AI automation of human-to-human touchpoints. The host counters this skepticism by citing decacorn valuations achieved by customer support AI platforms like Sierra and Decagon. | |
| Market Trends: Agent Hype vs. Undervalued Voice Modality | 3 | 5 | 2 | 1 | The host asks what is overvalued and undervalued in the current AI market. The guest takes a contrarian stance against near-term agentic hype while highlighting voice modality as an undervalued tool for widespread literacy and adoption. | |
| Origins of Predera: Forward-Deployed Research in Healthcare | 5 | 6 | 1 | 1 | The host explores the transition from a services model to product development. The guest walks through his forward-deployed experience embedded in hospital systems, navigating complex buyer personas and the healthcare 4Ps before expanding across industries. | |
| The Strategic Pivot to LLMOps and Navigating Dual Exits | 4 | 4 | 1 | 2 | The host presses on the financial breakdown and structure of the exits. The guest recounts walking away from a Walmart contract renewal to pivot fully toward LLMOps ahead of two distinct acquisitions. | |
| Services vs. Product Frameworks for AI Startups | 5 | 6 | 1 | 3 | The host questions whether forward-deployed engineering risks turning a startup into a bespoke single-client shop and asks about Palantir's model. The guest shares a strict five-customer rule across distinct verticals required to validate true product repeatability. | |
| Technical Founder to Enterprise Sales: Landing Fortune 500 Deals | 3 | 5 | 1 | 1 | The host asks for parting advice on how a technical founder evolves into an enterprise sales leader. The guest shares lessons learned from driving 50,000 miles across the country, emphasizing customer problem-solving over consulting slide decks. |