Mar 15, 2025 · 1h 36m · a16z

Building the Next Generation of Conversational AI

Ankit Kumar · 1h 9m spoken Anjney Midha · 15m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of the a16z podcast, Sesame CTO Ankit Kumar joins General Partner Anjney Midha to discuss the technical architecture, product philosophy, and long-term vision behind Sesame's human-like conversational AI companion.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The host as informed peer 3.8 Guest teaching 4.7 Guest disagreement 1.5 The host pushing back 2.3
05100:0020:0040:001:00:001:20:000:51–4:08 · The host as informed peer 3/10 Reflectings on the Launch of the Sesame Research Preview Anjney prompts Ankit to reflect on releasing the research preview despite knowing how much better internal iterations are. Ankit explains why creators underestimate their public releases, and Anjney pushes slightly on whether intuition guided the timing.4:08–6:49 · The host as informed peer 4/10 Navigating Intuition and Rigor in ML Product Development Anjney challenges Ankit by asking if the takeaway is just to trust one's gut in ML. Ankit rejects this simplistic framing, arguing for a blend of rigorous component evals and qualitative product taste.6:49–9:58 · The host as informed peer 3/10 Capturing Paralinguistic Cues and Emotional Tone in Audio Ankit explains the technical trade-offs in current voice models, detailing how transcription misses paralinguistic tone and why direct audio-to-audio models are the clear next step.9:58–14:11 · The host as informed peer 4/10 Drawing Inspiration from Pixar and AI-Native Media Anjney brings up their shared historical discussions about Pixar as an aspirational model for AI product taste. Ankit articulates why research labs suffer from a lack of creative direction and humanistic focus.14:11–17:53 · The host as informed peer 4/10 Strategic Resource Allocation and Identifying Unique Problems Anjney points out that much larger labs struggle to match Sesame's output despite vastly superior funding. Ankit notes that focus on naturalness and trade-offs in raw reasoning allow a small team to excel.17:53–20:21 · The host as informed peer 4/10 Contributing to the Open-Source Audio Research Community Anjney asks about the strategic selection of problems in ML development. Ankit explains how a startup must carefully choose what to build in-house versus relying on open-source advancements.20:21–25:10 · The host as informed peer 4/10 Distinguishing Base Speech Weights from the Full Proprietary Demo Ankit corrects public misconceptions about Sesame open-sourcing its proprietary demo characters, clarifying that they are releasing base contextual speech weights rather than full end-to-end agents.25:10–29:03 · The host as informed peer 3/10 Exploring Community Use Cases and Contextual Audio Dynamics Ankit explains the technical distinction between traditional text-to-speech and contextual speech generation, where surrounding conversational audio conditions emotional tone and mirroring.29:03–31:25 · The host as informed peer 3/10 Expanding Context Modalities and the Vision for Smart Glasses Anjney queries whether additional modalities like vision will provide diminishing returns. Ankit argues that visual context is crucial for companion hardware like smart glasses to feel present.31:25–35:51 · The host as informed peer 3/10 Always-On Presence and Eliminating Interface Friction Ankit presents the rationale for smart glasses as the ultimate companion form factor, emphasizing zero-friction, always-available interaction without phone unlocking.35:51–38:01 · The host as informed peer 4/10 Maintaining Singular Focus on the Consumer Companion Product Anjney asks why Sesame doesn't capitalize on massive enterprise demand by releasing a general API. Ankit firmly states that building an API business would be a fatal distraction from their primary companion product.38:01–40:50 · The host as informed peer 4/10 The Complexity of Crafting High-Quality AI Personalities Ankit details why prompt-tuning a generic API cannot achieve high-quality voice personalities, describing conversation as an independent modality requiring specialized research.40:50–43:57 · The host as informed peer 3/10 Mastering Turn-Taking, Interruption, and Conversational Flow Ankit breaks down the subtle mechanics of human conversation—such as backchannels, crosstalk, and constructive interruptions—that models must master to feel natural.43:57–47:01 · The host as informed peer 3/10 Optimizing Backend Infrastructure for Real-Time Interaction Ankit describes the backend systems engineering necessary to achieve sub-500ms voice response times with a lean engineering team of under fifteen people.47:01–51:33 · The host as informed peer 3/10 How Model Scale Improves Context, Homographs, and Name Pronunciation Ankit educates the host on how model parameter scale directly improves homograph disambiguation (e.g., lead vs. lead) and regional name pronunciations in speech generation.51:33–54:38 · The host as informed peer 4/10 Beyond Word Error Rate: Human Preference and Naturalness Evals Ankit details how speech evaluations have shifted beyond word error rate (WER) to human preference rankings and win rates against real human conversational snippets.54:38–58:34 · The host as informed peer 4/10 Resolving the Tension Between Energetic Acting and Casual Conversation Anjney highlights user feedback that Maya can sound overly theatrical or like an actor compared to casual voice chatter. Ankit accepts the criticism and outlines ongoing research to tune naturalness.58:34–1:03:53 · The host as informed peer 3/10 Reddit Fan Reactions and the Upcoming Standalone Companion App Ankit responds to enthusiastic Reddit community feedback, confirming plans for a dedicated companion application while detailing the transition toward integrated audio models.1:03:53–1:09:59 · The host as informed peer 4/10 Moving from Sentence-Based Decisions to 100ms Time Frames Ankit articulates why conversational models must move from making sentence-level decisions to continuous 100ms time-slice decisions to allow real-time fluid interruptions.1:09:59–1:15:57 · The host as informed peer 5/10 Learning the Value of Conversational Imperfections in User Experience Anjney draws comparisons between Sesame's launch moment and ChatGPT's release, asking if personality will be sacrificed to fix errors. Ankit rejects the comparison, maintaining that personality is Sesame's core differentiator.1:15:57–1:21:09 · The host as informed peer 5/10 Bringing the Computer to Life Through Natural Language Interfaces Anjney frames natural language as a paradigm shift in computing interfaces. Ankit expands on how companion interfaces act as the primary orchestration layer between users and downstream compute services.1:21:09–1:25:25 · The host as informed peer 6/10 Drawing Lessons from Steve Jobs, Apple, and Natural Motion in UI Anjney demonstrates deep expertise in historical UI milestones (Steve Jobs, Engelbart, Claude Shannon) to frame UI responsiveness and personality as key UX levers. Ankit agrees and expands on consumer computing history.1:25:25–1:29:23 · The host as informed peer 4/10 Reliability and Multi-Step Agent Integration in Voice AI Ankit explains why multi-step agent actions require 99% reliability to become daily habits, separating the companion interface layer from heavy reasoning execution engines.1:29:23–1:34:06 · The host as informed peer 4/10 Competition and Focus in Building AI Interfaces Ankit outlines why small, hyper-focused teams build superior consumer product experiences compared to large tech conglomerates, concluding with hiring criteria for Sesame.0:51–4:08 · Guest teaching 4/10 Reflectings on the Launch of the Sesame Research Preview Anjney prompts Ankit to reflect on releasing the research preview despite knowing how much better internal iterations are. Ankit explains why creators underestimate their public releases, and Anjney pushes slightly on whether intuition guided the timing.4:08–6:49 · Guest teaching 5/10 Navigating Intuition and Rigor in ML Product Development Anjney challenges Ankit by asking if the takeaway is just to trust one's gut in ML. Ankit rejects this simplistic framing, arguing for a blend of rigorous component evals and qualitative product taste.6:49–9:58 · Guest teaching 5/10 Capturing Paralinguistic Cues and Emotional Tone in Audio Ankit explains the technical trade-offs in current voice models, detailing how transcription misses paralinguistic tone and why direct audio-to-audio models are the clear next step.9:58–14:11 · Guest teaching 4/10 Drawing Inspiration from Pixar and AI-Native Media Anjney brings up their shared historical discussions about Pixar as an aspirational model for AI product taste. Ankit articulates why research labs suffer from a lack of creative direction and humanistic focus.14:11–17:53 · Guest teaching 5/10 Strategic Resource Allocation and Identifying Unique Problems Anjney points out that much larger labs struggle to match Sesame's output despite vastly superior funding. Ankit notes that focus on naturalness and trade-offs in raw reasoning allow a small team to excel.17:53–20:21 · Guest teaching 4/10 Contributing to the Open-Source Audio Research Community Anjney asks about the strategic selection of problems in ML development. Ankit explains how a startup must carefully choose what to build in-house versus relying on open-source advancements.20:21–25:10 · Guest teaching 6/10 Distinguishing Base Speech Weights from the Full Proprietary Demo Ankit corrects public misconceptions about Sesame open-sourcing its proprietary demo characters, clarifying that they are releasing base contextual speech weights rather than full end-to-end agents.25:10–29:03 · Guest teaching 5/10 Exploring Community Use Cases and Contextual Audio Dynamics Ankit explains the technical distinction between traditional text-to-speech and contextual speech generation, where surrounding conversational audio conditions emotional tone and mirroring.29:03–31:25 · Guest teaching 5/10 Expanding Context Modalities and the Vision for Smart Glasses Anjney queries whether additional modalities like vision will provide diminishing returns. Ankit argues that visual context is crucial for companion hardware like smart glasses to feel present.31:25–35:51 · Guest teaching 5/10 Always-On Presence and Eliminating Interface Friction Ankit presents the rationale for smart glasses as the ultimate companion form factor, emphasizing zero-friction, always-available interaction without phone unlocking.35:51–38:01 · Guest teaching 4/10 Maintaining Singular Focus on the Consumer Companion Product Anjney asks why Sesame doesn't capitalize on massive enterprise demand by releasing a general API. Ankit firmly states that building an API business would be a fatal distraction from their primary companion product.38:01–40:50 · Guest teaching 5/10 The Complexity of Crafting High-Quality AI Personalities Ankit details why prompt-tuning a generic API cannot achieve high-quality voice personalities, describing conversation as an independent modality requiring specialized research.40:50–43:57 · Guest teaching 5/10 Mastering Turn-Taking, Interruption, and Conversational Flow Ankit breaks down the subtle mechanics of human conversation—such as backchannels, crosstalk, and constructive interruptions—that models must master to feel natural.43:57–47:01 · Guest teaching 4/10 Optimizing Backend Infrastructure for Real-Time Interaction Ankit describes the backend systems engineering necessary to achieve sub-500ms voice response times with a lean engineering team of under fifteen people.47:01–51:33 · Guest teaching 6/10 How Model Scale Improves Context, Homographs, and Name Pronunciation Ankit educates the host on how model parameter scale directly improves homograph disambiguation (e.g., lead vs. lead) and regional name pronunciations in speech generation.51:33–54:38 · Guest teaching 5/10 Beyond Word Error Rate: Human Preference and Naturalness Evals Ankit details how speech evaluations have shifted beyond word error rate (WER) to human preference rankings and win rates against real human conversational snippets.54:38–58:34 · Guest teaching 5/10 Resolving the Tension Between Energetic Acting and Casual Conversation Anjney highlights user feedback that Maya can sound overly theatrical or like an actor compared to casual voice chatter. Ankit accepts the criticism and outlines ongoing research to tune naturalness.58:34–1:03:53 · Guest teaching 4/10 Reddit Fan Reactions and the Upcoming Standalone Companion App Ankit responds to enthusiastic Reddit community feedback, confirming plans for a dedicated companion application while detailing the transition toward integrated audio models.1:03:53–1:09:59 · Guest teaching 6/10 Moving from Sentence-Based Decisions to 100ms Time Frames Ankit articulates why conversational models must move from making sentence-level decisions to continuous 100ms time-slice decisions to allow real-time fluid interruptions.1:09:59–1:15:57 · Guest teaching 4/10 Learning the Value of Conversational Imperfections in User Experience Anjney draws comparisons between Sesame's launch moment and ChatGPT's release, asking if personality will be sacrificed to fix errors. Ankit rejects the comparison, maintaining that personality is Sesame's core differentiator.1:15:57–1:21:09 · Guest teaching 4/10 Bringing the Computer to Life Through Natural Language Interfaces Anjney frames natural language as a paradigm shift in computing interfaces. Ankit expands on how companion interfaces act as the primary orchestration layer between users and downstream compute services.1:21:09–1:25:25 · Guest teaching 4/10 Drawing Lessons from Steve Jobs, Apple, and Natural Motion in UI Anjney demonstrates deep expertise in historical UI milestones (Steve Jobs, Engelbart, Claude Shannon) to frame UI responsiveness and personality as key UX levers. Ankit agrees and expands on consumer computing history.1:25:25–1:29:23 · Guest teaching 5/10 Reliability and Multi-Step Agent Integration in Voice AI Ankit explains why multi-step agent actions require 99% reliability to become daily habits, separating the companion interface layer from heavy reasoning execution engines.1:29:23–1:34:06 · Guest teaching 4/10 Competition and Focus in Building AI Interfaces Ankit outlines why small, hyper-focused teams build superior consumer product experiences compared to large tech conglomerates, concluding with hiring criteria for Sesame.0:51–4:08 · Guest disagreement 2/10 Reflectings on the Launch of the Sesame Research Preview Anjney prompts Ankit to reflect on releasing the research preview despite knowing how much better internal iterations are. Ankit explains why creators underestimate their public releases, and Anjney pushes slightly on whether intuition guided the timing.4:08–6:49 · Guest disagreement 4/10 Navigating Intuition and Rigor in ML Product Development Anjney challenges Ankit by asking if the takeaway is just to trust one's gut in ML. Ankit rejects this simplistic framing, arguing for a blend of rigorous component evals and qualitative product taste.6:49–9:58 · Guest disagreement 1/10 Capturing Paralinguistic Cues and Emotional Tone in Audio Ankit explains the technical trade-offs in current voice models, detailing how transcription misses paralinguistic tone and why direct audio-to-audio models are the clear next step.9:58–14:11 · Guest disagreement 1/10 Drawing Inspiration from Pixar and AI-Native Media Anjney brings up their shared historical discussions about Pixar as an aspirational model for AI product taste. Ankit articulates why research labs suffer from a lack of creative direction and humanistic focus.14:11–17:53 · Guest disagreement 2/10 Strategic Resource Allocation and Identifying Unique Problems Anjney points out that much larger labs struggle to match Sesame's output despite vastly superior funding. Ankit notes that focus on naturalness and trade-offs in raw reasoning allow a small team to excel.17:53–20:21 · Guest disagreement 1/10 Contributing to the Open-Source Audio Research Community Anjney asks about the strategic selection of problems in ML development. Ankit explains how a startup must carefully choose what to build in-house versus relying on open-source advancements.20:21–25:10 · Guest disagreement 3/10 Distinguishing Base Speech Weights from the Full Proprietary Demo Ankit corrects public misconceptions about Sesame open-sourcing its proprietary demo characters, clarifying that they are releasing base contextual speech weights rather than full end-to-end agents.25:10–29:03 · Guest disagreement 1/10 Exploring Community Use Cases and Contextual Audio Dynamics Ankit explains the technical distinction between traditional text-to-speech and contextual speech generation, where surrounding conversational audio conditions emotional tone and mirroring.29:03–31:25 · Guest disagreement 1/10 Expanding Context Modalities and the Vision for Smart Glasses Anjney queries whether additional modalities like vision will provide diminishing returns. Ankit argues that visual context is crucial for companion hardware like smart glasses to feel present.31:25–35:51 · Guest disagreement 2/10 Always-On Presence and Eliminating Interface Friction Ankit presents the rationale for smart glasses as the ultimate companion form factor, emphasizing zero-friction, always-available interaction without phone unlocking.35:51–38:01 · Guest disagreement 2/10 Maintaining Singular Focus on the Consumer Companion Product Anjney asks why Sesame doesn't capitalize on massive enterprise demand by releasing a general API. Ankit firmly states that building an API business would be a fatal distraction from their primary companion product.38:01–40:50 · Guest disagreement 1/10 The Complexity of Crafting High-Quality AI Personalities Ankit details why prompt-tuning a generic API cannot achieve high-quality voice personalities, describing conversation as an independent modality requiring specialized research.40:50–43:57 · Guest disagreement 1/10 Mastering Turn-Taking, Interruption, and Conversational Flow Ankit breaks down the subtle mechanics of human conversation—such as backchannels, crosstalk, and constructive interruptions—that models must master to feel natural.43:57–47:01 · Guest disagreement 1/10 Optimizing Backend Infrastructure for Real-Time Interaction Ankit describes the backend systems engineering necessary to achieve sub-500ms voice response times with a lean engineering team of under fifteen people.47:01–51:33 · Guest disagreement 1/10 How Model Scale Improves Context, Homographs, and Name Pronunciation Ankit educates the host on how model parameter scale directly improves homograph disambiguation (e.g., lead vs. lead) and regional name pronunciations in speech generation.51:33–54:38 · Guest disagreement 1/10 Beyond Word Error Rate: Human Preference and Naturalness Evals Ankit details how speech evaluations have shifted beyond word error rate (WER) to human preference rankings and win rates against real human conversational snippets.54:38–58:34 · Guest disagreement 2/10 Resolving the Tension Between Energetic Acting and Casual Conversation Anjney highlights user feedback that Maya can sound overly theatrical or like an actor compared to casual voice chatter. Ankit accepts the criticism and outlines ongoing research to tune naturalness.58:34–1:03:53 · Guest disagreement 1/10 Reddit Fan Reactions and the Upcoming Standalone Companion App Ankit responds to enthusiastic Reddit community feedback, confirming plans for a dedicated companion application while detailing the transition toward integrated audio models.1:03:53–1:09:59 · Guest disagreement 2/10 Moving from Sentence-Based Decisions to 100ms Time Frames Ankit articulates why conversational models must move from making sentence-level decisions to continuous 100ms time-slice decisions to allow real-time fluid interruptions.1:09:59–1:15:57 · Guest disagreement 1/10 Learning the Value of Conversational Imperfections in User Experience Anjney draws comparisons between Sesame's launch moment and ChatGPT's release, asking if personality will be sacrificed to fix errors. Ankit rejects the comparison, maintaining that personality is Sesame's core differentiator.1:15:57–1:21:09 · Guest disagreement 1/10 Bringing the Computer to Life Through Natural Language Interfaces Anjney frames natural language as a paradigm shift in computing interfaces. Ankit expands on how companion interfaces act as the primary orchestration layer between users and downstream compute services.1:21:09–1:25:25 · Guest disagreement 2/10 Drawing Lessons from Steve Jobs, Apple, and Natural Motion in UI Anjney demonstrates deep expertise in historical UI milestones (Steve Jobs, Engelbart, Claude Shannon) to frame UI responsiveness and personality as key UX levers. Ankit agrees and expands on consumer computing history.1:25:25–1:29:23 · Guest disagreement 1/10 Reliability and Multi-Step Agent Integration in Voice AI Ankit explains why multi-step agent actions require 99% reliability to become daily habits, separating the companion interface layer from heavy reasoning execution engines.1:29:23–1:34:06 · Guest disagreement 1/10 Competition and Focus in Building AI Interfaces Ankit outlines why small, hyper-focused teams build superior consumer product experiences compared to large tech conglomerates, concluding with hiring criteria for Sesame.0:51–4:08 · The host pushing back 3/10 Reflectings on the Launch of the Sesame Research Preview Anjney prompts Ankit to reflect on releasing the research preview despite knowing how much better internal iterations are. Ankit explains why creators underestimate their public releases, and Anjney pushes slightly on whether intuition guided the timing.4:08–6:49 · The host pushing back 5/10 Navigating Intuition and Rigor in ML Product Development Anjney challenges Ankit by asking if the takeaway is just to trust one's gut in ML. Ankit rejects this simplistic framing, arguing for a blend of rigorous component evals and qualitative product taste.6:49–9:58 · The host pushing back 2/10 Capturing Paralinguistic Cues and Emotional Tone in Audio Ankit explains the technical trade-offs in current voice models, detailing how transcription misses paralinguistic tone and why direct audio-to-audio models are the clear next step.9:58–14:11 · The host pushing back 2/10 Drawing Inspiration from Pixar and AI-Native Media Anjney brings up their shared historical discussions about Pixar as an aspirational model for AI product taste. Ankit articulates why research labs suffer from a lack of creative direction and humanistic focus.14:11–17:53 · The host pushing back 3/10 Strategic Resource Allocation and Identifying Unique Problems Anjney points out that much larger labs struggle to match Sesame's output despite vastly superior funding. Ankit notes that focus on naturalness and trade-offs in raw reasoning allow a small team to excel.17:53–20:21 · The host pushing back 2/10 Contributing to the Open-Source Audio Research Community Anjney asks about the strategic selection of problems in ML development. Ankit explains how a startup must carefully choose what to build in-house versus relying on open-source advancements.20:21–25:10 · The host pushing back 3/10 Distinguishing Base Speech Weights from the Full Proprietary Demo Ankit corrects public misconceptions about Sesame open-sourcing its proprietary demo characters, clarifying that they are releasing base contextual speech weights rather than full end-to-end agents.25:10–29:03 · The host pushing back 2/10 Exploring Community Use Cases and Contextual Audio Dynamics Ankit explains the technical distinction between traditional text-to-speech and contextual speech generation, where surrounding conversational audio conditions emotional tone and mirroring.29:03–31:25 · The host pushing back 2/10 Expanding Context Modalities and the Vision for Smart Glasses Anjney queries whether additional modalities like vision will provide diminishing returns. Ankit argues that visual context is crucial for companion hardware like smart glasses to feel present.31:25–35:51 · The host pushing back 2/10 Always-On Presence and Eliminating Interface Friction Ankit presents the rationale for smart glasses as the ultimate companion form factor, emphasizing zero-friction, always-available interaction without phone unlocking.35:51–38:01 · The host pushing back 3/10 Maintaining Singular Focus on the Consumer Companion Product Anjney asks why Sesame doesn't capitalize on massive enterprise demand by releasing a general API. Ankit firmly states that building an API business would be a fatal distraction from their primary companion product.38:01–40:50 · The host pushing back 2/10 The Complexity of Crafting High-Quality AI Personalities Ankit details why prompt-tuning a generic API cannot achieve high-quality voice personalities, describing conversation as an independent modality requiring specialized research.40:50–43:57 · The host pushing back 1/10 Mastering Turn-Taking, Interruption, and Conversational Flow Ankit breaks down the subtle mechanics of human conversation—such as backchannels, crosstalk, and constructive interruptions—that models must master to feel natural.43:57–47:01 · The host pushing back 2/10 Optimizing Backend Infrastructure for Real-Time Interaction Ankit describes the backend systems engineering necessary to achieve sub-500ms voice response times with a lean engineering team of under fifteen people.47:01–51:33 · The host pushing back 1/10 How Model Scale Improves Context, Homographs, and Name Pronunciation Ankit educates the host on how model parameter scale directly improves homograph disambiguation (e.g., lead vs. lead) and regional name pronunciations in speech generation.51:33–54:38 · The host pushing back 2/10 Beyond Word Error Rate: Human Preference and Naturalness Evals Ankit details how speech evaluations have shifted beyond word error rate (WER) to human preference rankings and win rates against real human conversational snippets.54:38–58:34 · The host pushing back 3/10 Resolving the Tension Between Energetic Acting and Casual Conversation Anjney highlights user feedback that Maya can sound overly theatrical or like an actor compared to casual voice chatter. Ankit accepts the criticism and outlines ongoing research to tune naturalness.58:34–1:03:53 · The host pushing back 1/10 Reddit Fan Reactions and the Upcoming Standalone Companion App Ankit responds to enthusiastic Reddit community feedback, confirming plans for a dedicated companion application while detailing the transition toward integrated audio models.1:03:53–1:09:59 · The host pushing back 2/10 Moving from Sentence-Based Decisions to 100ms Time Frames Ankit articulates why conversational models must move from making sentence-level decisions to continuous 100ms time-slice decisions to allow real-time fluid interruptions.1:09:59–1:15:57 · The host pushing back 3/10 Learning the Value of Conversational Imperfections in User Experience Anjney draws comparisons between Sesame's launch moment and ChatGPT's release, asking if personality will be sacrificed to fix errors. Ankit rejects the comparison, maintaining that personality is Sesame's core differentiator.1:15:57–1:21:09 · The host pushing back 2/10 Bringing the Computer to Life Through Natural Language Interfaces Anjney frames natural language as a paradigm shift in computing interfaces. Ankit expands on how companion interfaces act as the primary orchestration layer between users and downstream compute services.1:21:09–1:25:25 · The host pushing back 3/10 Drawing Lessons from Steve Jobs, Apple, and Natural Motion in UI Anjney demonstrates deep expertise in historical UI milestones (Steve Jobs, Engelbart, Claude Shannon) to frame UI responsiveness and personality as key UX levers. Ankit agrees and expands on consumer computing history.1:25:25–1:29:23 · The host pushing back 2/10 Reliability and Multi-Step Agent Integration in Voice AI Ankit explains why multi-step agent actions require 99% reliability to become daily habits, separating the companion interface layer from heavy reasoning execution engines.1:29:23–1:34:06 · The host pushing back 2/10 Competition and Focus in Building AI Interfaces Ankit outlines why small, hyper-focused teams build superior consumer product experiences compared to large tech conglomerates, concluding with hiring criteria for Sesame.

speaking balance: gold is the host, purple is the guest (3 minute bins)

0:00 · the host 0% · guest 100%0:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%18:00 · the host 0% · guest 100%18:00 · the host 0% · guest 100%21:00 · the host 0% · guest 100%21:00 · the host 0% · guest 100%24:00 · the host 0% · guest 100%24:00 · the host 0% · guest 100%27:00 · the host 0% · guest 100%27:00 · the host 0% · guest 100%30:00 · the host 0% · guest 100%30:00 · the host 0% · guest 100%33:00 · the host 0% · guest 100%33:00 · the host 0% · guest 100%36:00 · the host 0% · guest 100%36:00 · the host 0% · guest 100%39:00 · the host 0% · guest 100%39:00 · the host 0% · guest 100%42:00 · the host 0% · guest 100%42:00 · the host 0% · guest 100%45:00 · the host 0% · guest 100%45:00 · the host 0% · guest 100%48:00 · the host 0% · guest 100%48:00 · the host 0% · guest 100%51:00 · the host 0% · guest 100%51:00 · the host 0% · guest 100%54:00 · the host 0% · guest 100%54:00 · the host 0% · guest 100%57:00 · the host 0% · guest 100%57:00 · the host 0% · guest 100%1:00:00 · the host 0% · guest 100%1:00:00 · the host 0% · guest 100%1:03:00 · the host 0% · guest 100%1:03:00 · the host 0% · guest 100%1:06:00 · the host 0% · guest 100%1:06:00 · the host 0% · guest 100%1:09:00 · the host 0% · guest 100%1:09:00 · the host 0% · guest 100%1:12:00 · the host 0% · guest 100%1:12:00 · the host 0% · guest 100%1:15:00 · the host 0% · guest 100%1:15:00 · the host 0% · guest 100%1:18:00 · the host 0% · guest 100%1:18:00 · the host 0% · guest 100%1:21:00 · the host 0% · guest 100%1:21:00 · the host 0% · guest 100%1:24:00 · the host 0% · guest 100%1:24:00 · the host 0% · guest 100%1:27:00 · the host 0% · guest 100%1:27:00 · the host 0% · guest 100%1:30:00 · the host 0% · guest 100%1:30:00 · the host 0% · guest 100%1:33:00 · the host 0% · guest 100%1:33:00 · the host 0% · guest 100%1:36:00 · the host 0% · guest 100%1:36:00 · the host 0% · guest 100%
Sharpest disagreement ▶ 4:08 Rejection of gut-feeling framing

When Anjney asks if the core lesson in ML product development is simply trusting one's gut, Ankit directly counters that gut feeling alone is an ineffective development mechanism, demanding rigorous evaluations combined with qualitative product taste.

Hardest push from the host ▶ 1:13:00 Pushback on product regression risk

Anjney challenges Ankit by pointing out that previous AI breakthroughs like ChatGPT regressed in personality when forced to fix errors, pressing Ankit on whether Sesame will suffer the same fate.

Biggest teaching moment ▶ 47:05 Detailed breakdown of model scaling benefits

Ankit systematically educates the host on how model scale resolves complex long-tail phonetics, using specific examples like homographs (lead/lead, row/row) and regional accent preservation.

The host holds their own ▶ 1:21:09 Historical contextualization of UI paradigms

Anjney demonstrates expert-level domain knowledge by connecting Claude Shannon, Doug Engelbart, and Steve Jobs's UI breakthroughs to argue that personality in voice AI is the modern equivalent of UI capacitive touch and bounce-back physics.

the scores for every segment, with the reasoning behind each
ChapterTopicThe host as informed peerGuest teachingGuest disagreementThe host pushing backWhy
Reflectings on the Launch of the Sesame Research Preview 3423 Anjney prompts Ankit to reflect on releasing the research preview despite knowing how much better internal iterations are. Ankit explains why creators underestimate their public releases, and Anjney pushes slightly on whether intuition guided the timing.
Navigating Intuition and Rigor in ML Product Development 4545 Anjney challenges Ankit by asking if the takeaway is just to trust one's gut in ML. Ankit rejects this simplistic framing, arguing for a blend of rigorous component evals and qualitative product taste.
Capturing Paralinguistic Cues and Emotional Tone in Audio 3512 Ankit explains the technical trade-offs in current voice models, detailing how transcription misses paralinguistic tone and why direct audio-to-audio models are the clear next step.
Drawing Inspiration from Pixar and AI-Native Media 4412 Anjney brings up their shared historical discussions about Pixar as an aspirational model for AI product taste. Ankit articulates why research labs suffer from a lack of creative direction and humanistic focus.
Strategic Resource Allocation and Identifying Unique Problems 4523 Anjney points out that much larger labs struggle to match Sesame's output despite vastly superior funding. Ankit notes that focus on naturalness and trade-offs in raw reasoning allow a small team to excel.
Contributing to the Open-Source Audio Research Community 4412 Anjney asks about the strategic selection of problems in ML development. Ankit explains how a startup must carefully choose what to build in-house versus relying on open-source advancements.
Distinguishing Base Speech Weights from the Full Proprietary Demo 4633 Ankit corrects public misconceptions about Sesame open-sourcing its proprietary demo characters, clarifying that they are releasing base contextual speech weights rather than full end-to-end agents.
Exploring Community Use Cases and Contextual Audio Dynamics 3512 Ankit explains the technical distinction between traditional text-to-speech and contextual speech generation, where surrounding conversational audio conditions emotional tone and mirroring.
Expanding Context Modalities and the Vision for Smart Glasses 3512 Anjney queries whether additional modalities like vision will provide diminishing returns. Ankit argues that visual context is crucial for companion hardware like smart glasses to feel present.
Always-On Presence and Eliminating Interface Friction 3522 Ankit presents the rationale for smart glasses as the ultimate companion form factor, emphasizing zero-friction, always-available interaction without phone unlocking.
Maintaining Singular Focus on the Consumer Companion Product 4423 Anjney asks why Sesame doesn't capitalize on massive enterprise demand by releasing a general API. Ankit firmly states that building an API business would be a fatal distraction from their primary companion product.
The Complexity of Crafting High-Quality AI Personalities 4512 Ankit details why prompt-tuning a generic API cannot achieve high-quality voice personalities, describing conversation as an independent modality requiring specialized research.
Mastering Turn-Taking, Interruption, and Conversational Flow 3511 Ankit breaks down the subtle mechanics of human conversation—such as backchannels, crosstalk, and constructive interruptions—that models must master to feel natural.
Optimizing Backend Infrastructure for Real-Time Interaction 3412 Ankit describes the backend systems engineering necessary to achieve sub-500ms voice response times with a lean engineering team of under fifteen people.
How Model Scale Improves Context, Homographs, and Name Pronunciation 3611 Ankit educates the host on how model parameter scale directly improves homograph disambiguation (e.g., lead vs. lead) and regional name pronunciations in speech generation.
Beyond Word Error Rate: Human Preference and Naturalness Evals 4512 Ankit details how speech evaluations have shifted beyond word error rate (WER) to human preference rankings and win rates against real human conversational snippets.
Resolving the Tension Between Energetic Acting and Casual Conversation 4523 Anjney highlights user feedback that Maya can sound overly theatrical or like an actor compared to casual voice chatter. Ankit accepts the criticism and outlines ongoing research to tune naturalness.
Reddit Fan Reactions and the Upcoming Standalone Companion App 3411 Ankit responds to enthusiastic Reddit community feedback, confirming plans for a dedicated companion application while detailing the transition toward integrated audio models.
Moving from Sentence-Based Decisions to 100ms Time Frames 4622 Ankit articulates why conversational models must move from making sentence-level decisions to continuous 100ms time-slice decisions to allow real-time fluid interruptions.
Learning the Value of Conversational Imperfections in User Experience 5413 Anjney draws comparisons between Sesame's launch moment and ChatGPT's release, asking if personality will be sacrificed to fix errors. Ankit rejects the comparison, maintaining that personality is Sesame's core differentiator.
Bringing the Computer to Life Through Natural Language Interfaces 5412 Anjney frames natural language as a paradigm shift in computing interfaces. Ankit expands on how companion interfaces act as the primary orchestration layer between users and downstream compute services.
Drawing Lessons from Steve Jobs, Apple, and Natural Motion in UI 6423 Anjney demonstrates deep expertise in historical UI milestones (Steve Jobs, Engelbart, Claude Shannon) to frame UI responsiveness and personality as key UX levers. Ankit agrees and expands on consumer computing history.
Reliability and Multi-Step Agent Integration in Voice AI 4512 Ankit explains why multi-step agent actions require 99% reliability to become daily habits, separating the companion interface layer from heavy reasoning execution engines.
Competition and Focus in Building AI Interfaces 4412 Ankit outlines why small, hyper-focused teams build superior consumer product experiences compared to large tech conglomerates, concluding with hiring criteria for Sesame.

Statements from this episode (58)

Insight
Kumar: Builders underestimate current releases due to internal roadmap gaps
“When you build the thing, right, when you're building the product and using it every day, you know, there are some things that you work on that don't get into the demo because they're going to take longer and you want to ship the demo. You kind of know how big…”
Ankit Kumar Mar 15, 2025 ▶ 1:16
Insight
Kumar: AI product optimization depends on hard-to-quantify qualitative user reactions
“But really, I think with some of these more product experience questions, there's something qualitative about it that is very hard to quantify. That is one of the big challenges internally, actually, is how do you hill climb effectively on what is really an ML…”
Ankit Kumar Mar 15, 2025 ▶ 2:54
Insight
Kumar: Internal ML testing fails when teams exhaust fresh user reactions
“Misleading at times because you tried so much and you don't have, at least when we're trying it internally, you don't have such a diversity of users that you get kind of the first reaction over and over, right? You only get so many first reactions. And then wh…”
Ankit Kumar Mar 15, 2025 ▶ 3:42
Prediction Not checkable as stated
Kumar: Transcription-free conversational AI models are coming soon
“A pretty clear path that a lot of, I think, labs are taking, and we're taking as well, and will be in kind of future versions, is just kind of transcription free. Just go straight into the text component which will kind of obviate transcription entirely. That …”
Ankit Kumar Mar 15, 2025 ▶ 5:37
Disclosure
Kumar: Sesame's current demo cannot detect user emotional tone
“The current demo does not sort of hear the user from the perspective of their paralinguistic kind of emotional tone and so forth.”
Ankit Kumar Mar 15, 2025 ▶ 7:06
Insight
Kumar: Speech-to-text transcription misses essential non-verbal audio cues
“Humans, of course, convey a lot of information through Their speech that is not the words, the content of the speech and transcription misses that entirely.”
Ankit Kumar Mar 15, 2025 ▶ 7:21
Prediction Open · timeframe Mar 2028
Kumar: Next Sesame AI models will feed audio natively into LLMs
“And so the kind of next versions of our models that will take audio and natively into the kind of LLM component will hopefully more and more pick up on those things.”
Ankit Kumar Mar 15, 2025 ▶ 7:32
Disclosure
Kumar: Sesame's entire software and ML team is under 15 people
“The full software team today is still under 15 people, and so we just don't, that's including ML and infrastructure and everything.”
Ankit Kumar Mar 15, 2025 ▶ 8:28
Disclosure
Kumar: Sesame trades complex AI reasoning for natural voice interaction
“So, you know, if you talk to Maya and Miles, you probably will not be able to get the same quality of like reasoning capabilities or intelligence as other As other systems, but in return, you're kind of getting this much more natural fluid interaction.”
Ankit Kumar Mar 15, 2025 ▶ 9:14
Disclosure
Kumar: Sesame is not pre-training frontier LLMs at scale
“You know, we are not A frontier model company. We're not pre-training LLMs at insane scale and so forth.”
Ankit Kumar Mar 15, 2025 ▶ 9:40
Opinion
Kumar: Top AI labs under-invest in creative taste and humanities
“I do think that there is kind of an under-investment or an under-focus in the sort of Strong AI team world on product experience and sort of creative taste and kind of humanities maybe, in a sense, to kind of bring AI to experiences that kind of everyday peopl…”
Ankit Kumar Mar 15, 2025 ▶ 10:32
Prediction Not checkable as stated
Kumar: Storytelling and AI will merge into new AI-native media categories
“And I think that we will see a lot of not just Sesame, but other kinds of media, let's say, that are sort of AI native in a way that Bring some creativity, bring some like storytelling into AI, or maybe bring AI into those categories. And I think they'll make …”
Ankit Kumar Mar 15, 2025 ▶ 10:57
Disclosure
Kumar: Sesame built its open-source voice models from scratch
“Like we had to build the models that we're going to open source from scratch in order to get them to a point where they can achieve this experience.”
Ankit Kumar Mar 15, 2025 ▶ 11:59
Insight
Kumar: Good ML taste means avoiding what APIs will soon commoditize
“I think from my perspective, good taste in ML today, because it's such a fast moving field with so many people working across, you know, open source and APIs and big labs and so forth. Really, you're trying to identify What part of the ecosystem or what part o…”
Ankit Kumar Mar 15, 2025 ▶ 14:52
Prediction Not checkable as stated
Kumar: Open source won't solve AI voice and personality features
“But other parts in particular, kind of some of the personality aspects, some of the voice aspects, the speech generation, we didn't think and we still don't think will just be kind of done by the community. We think we will need to do it because that's kind of…”
Ankit Kumar Mar 15, 2025 ▶ 17:30
Disclosure
Kumar: Sesame is not building an API or developer-facing product
“We are not a developer facing business. We're not making an API.”
Ankit Kumar Mar 15, 2025 ▶ 19:05
Disclosure
Kumar: Sesame is open sourcing its speech model, not the full demo
“We're not open sourcing the demo. We're open sourcing the speech generation model that is powering the voice of the demo.”
Ankit Kumar Mar 15, 2025 ▶ 20:58
Assertion Supported
Kumar: Sesame base model generates any voice with fine-tuning
“We are open sourcing the speech generation base model basically. And so the base model can generate any voice. It's quite conversational, but you do need to fine tune it probably if you want to get a particular personality or a particular kind of voice out of …”
Ankit Kumar Mar 15, 2025 ▶ 22:55
Assertion Supported
Kumar: Sesame achieves voice cloning via in-context learning prompt strings
“The model is this kind of, you know, it has kind of this in context learning style voice cloning. I mean, typically with some other kind of text-to-speech models, the voice cloning is kind of like an explicit feature. So it's sort of the model has dedicated ki…”
Ankit Kumar Mar 15, 2025 ▶ 23:55
Assertion Contradicted
Kumar: No other open-source model generates multi-participant contextual audio
“At least to our knowledge, there's not another model out there that, that is open source that kind of is a sort of contextual thing where you kind of can put two participants in a conversation, even more, three, and generate kind of a conversation between them…”
Ankit Kumar Mar 15, 2025 ▶ 25:47
Insight
Kumar: Traditional text-to-speech sounds flat because non-neutral tones risk sounding inappropriate
“And that's probably why, or it's one of the reasons why historically voice assistants feel so flat is that traditional text of speech, it's kind of like it can only be flat. Or in other words, if it tries to not be flat, it's very likely wrong.”
Ankit Kumar Mar 15, 2025 ▶ 28:42
Prediction Not checkable as stated
Kumar: Speech research community will shift toward contextual AI architectures
“So, so the speech generation research community is very likely, I think, to move to more and more contextual architectures basically.”
Ankit Kumar Mar 15, 2025 ▶ 28:56
Assertion Not checkable as stated
Kumar: Sesame's speech generation is conditioned on full conversation audio
“The speech generation part is conditioned on the, on all the audio of the conversation.”
Ankit Kumar Mar 15, 2025 ▶ 29:18
Prediction Open · timeframe Mar 2028
Kumar: Sesame is actively developing smart glasses for its AI companions
“We mentioned on the website and we mentioned some of our launch content that we are working towards glasses as a form factor for Companions, or kind of this companion interface”
Ankit Kumar Mar 15, 2025 ▶ 30:33
Prediction Not checkable as stated
Kumar: Smartphones and laptops will not be replaced anytime soon
“No one's going to replace phones anytime soon, or laptops for that matter.”
Ankit Kumar Mar 15, 2025 ▶ 32:41
Insight
Kumar: Seemingly easy side products like APIs create massive engineering drag
“Sometimes it feels like an API or something like that is like relatively easy to do. And, you know, it's not like, it's not maybe as hard as some of the other things that we're doing, but everything is a drag on engineering, right?”
Ankit Kumar Mar 15, 2025 ▶ 37:43
Prediction Held up
Kumar: Sesame will not build a one-size-fits-all AI companion
“So we're certainly not going to, we don't see our product as like one companion that's the same for everyone. People have different preferences and that has to be a part of this kind of product category for sure.”
Ankit Kumar Mar 15, 2025 ▶ 38:37
Insight
Kumar: Voice cloning and prompting alone cannot create great AI personalities
“It takes more than just sort of, you know, voice clone plus change the prompt. Now you have a new character that's just as good as it would be if you spent a lot of time on it. It takes, I think today making a great personality voice interface system, we can't…”
Ankit Kumar Mar 15, 2025 ▶ 39:13
Insight
Kumar: AI conversation is a distinct modality requiring core research
“I think that conversation, like human conversation, is kind of its own modality. And it is nowhere near done, right? There's so much more to do in the core research side to make it better.”
Ankit Kumar Mar 15, 2025 ▶ 40:05
Insight
Kumar: True AI naturalness requires modeling turn-taking and backchannels
“I think to get these things to feel very, very natural and real, you do need to model the full conversation, the turn taking, the back channels, everything.”
Ankit Kumar Mar 15, 2025 ▶ 42:38
Assertion Supported
Kumar: Sesame targets sub-500 millisecond response times for voice AI
“We want You know, sub-five hundred millisecond response times, and a lot of things that feel like not a big deal, 50 milliseconds here, 50 milliseconds there, can really add up.”
Ankit Kumar Mar 15, 2025 ▶ 45:15
Insight
Kumar: Early AI startups need flexible systems thinkers over niche specialists
“Especially when you're smaller, you know, you don't really want to harden, like, you know, you have this team that is super, super niche and doing only this thing, because you don't know exactly what the stack is going to look like tomorrow. Things change on t…”
Ankit Kumar Mar 15, 2025 ▶ 46:08
Disclosure
Kumar: Sesame trained 1B, 3B, and 8B parameter models for speech generation
“So we published in the, in our blog post, we trained three variants. We trained 1,000,000,003 1,000,000,008 billion of just speech generation.”
Ankit Kumar Mar 15, 2025 ▶ 47:07
Assertion Supported
Kumar: Larger speech models handle homographs and context-dependent pronunciation better
“And we see that as the models get bigger, they're much better at picking the right pronunciation in examples like this.”
Ankit Kumar Mar 15, 2025 ▶ 48:31
Assertion Not checkable as stated
Kumar: Word error rate metrics for AI speech generation are now saturated
“Earlier on in the speech generation world in the community, very often you'd look at like word error rate where you look at transcription, like you kind of have a sentence and you generate and you transcribe it and you see if it's the same. And those metrics a…”
Ankit Kumar Mar 15, 2025 ▶ 52:51
Disclosure
Kumar: Sesame evaluates speech models against real human conversation continuations
“We also have some data sets that are kind of like Just two people in a conversation or sometimes they're actors, but it's trying to be a real conversation. And so we'll kind of take the conditioning of some snippet of the conversation and then show a human rat…”
Ankit Kumar Mar 15, 2025 ▶ 54:10
Insight
Midha: Most Discord voice activity is casual, low-energy hanging out
“You know, we spend a lot of time at Discord and, you know, a huge amount of Discord usages in voice channels, and when you look, when you actually kind of spend time looking at and trying to observe what the shape of a great conversation and voice is on Discor…”
Anjney Midha Mar 15, 2025 ▶ 55:15
Insight
Kumar: Achieving human realism in voice AI is harder than text
“I think you, I think it's much easier. It would be much easier to make a system that produces text chats with you that feels like you're texting a human because there's such a compression of like what the entity on the other side is into just like text. Wherea…”
Ankit Kumar Mar 15, 2025 ▶ 56:40
Prediction Held up
Kumar: Sesame is developing a companion AI app with persistent memory
“We are making an app. We will make an app. I think for a little bit of time, it's going to still be kind of the demo experience. We want to support people using that for a long time, or, you know, we don't want to, we're not taking it away anytime soon from wh…”
Ankit Kumar Mar 15, 2025 ▶ 59:04
Prediction Didn’t hold up
Kumar: Sesame will build a unified audio-text transformer within months
“The path that we're going to take, I think, over the next few months is making a single transformer that does both audio understanding, content, text content generation, and speech generation.”
Ankit Kumar Mar 15, 2025 ▶ 1:00:20
Insight
Kumar: Adding generative modalities to AI models is harder than understanding
“It's much harder to add a modality to a pre-trained model than it is, add a generative modality, than it is to add an understanding modality.”
Ankit Kumar Mar 15, 2025 ▶ 1:00:30
Prediction Not checkable as stated
Kumar: Conversational AI will eventually rely on single models over heuristic pipelines
“I don't think you want to, in the long term, have those dynamics be like heuristics and so on, which they kind of are now. There are models involved in some heuristics and so forth. I think in the long term, it's just one model that is kind of naturally employ…”
Ankit Kumar Mar 15, 2025 ▶ 1:03:36
Insight
Kumar: Voice AI models must decide every 100 milliseconds for natural interaction
“You need to make decisions at the hundred millisecond, let's say, Time segment so that if you're talking and the other person, you know, starts sort of making some noises that make it seem like they're trying to interrupt you or they want to say something, you…”
Ankit Kumar Mar 15, 2025 ▶ 1:05:01
Prediction Not checkable as stated
Kumar: Near-term AI voice models will still lack dynamic conversational fluidity
“And the models that we have today, like CSM, for example, And probably some of the models that we'll have in the short term that will make the experience better will still not be modeling the conversational dynamics because they're making decisions that kind o…”
Ankit Kumar Mar 15, 2025 ▶ 1:05:25
Disclosure
Kumar: Sesame is developing diffusion-based audio generation models
“We are also working, by the way, on kind of ideas that make the audio generation part diffusion.”
Ankit Kumar Mar 15, 2025 ▶ 1:06:57
Prediction Not checkable as stated
Kumar: Transformers will remain the dominant AI sequence architecture short-term
“And I wouldn't bet against transformers, you know, not in the short term anyways.”
Ankit Kumar Mar 15, 2025 ▶ 1:08:15
Insight
Kumar: Theoretical gains won't unseat transformers without matching years of optimization
“But because of all the engineering work that the community has done around Transformers, it's like, you know, it's very good. And you're not going to just sort of unseat that, you know, just by an idea, right? There's a lot of work to be done.”
Ankit Kumar Mar 15, 2025 ▶ 1:09:38
Disclosure
Kumar: Sesame intentionally includes speech imperfections to make AI sound natural
“My and Miles, they might sort of say the wrong thing or kind of like back up a little bit and say something else or something. And that's on purpose, of course.”
Ankit Kumar Mar 15, 2025 ▶ 1:11:02
Opinion
Midha: Sesame's research preview is the ChatGPT moment for voice AI
“Relative to five, six days ago, I don't think it's too much of an exaggeration anymore to say this research preview has been for voice what ChatGPT was for text.”
Anjney Midha Mar 15, 2025 ▶ 1:12:28
Prediction Not checkable as stated
Kumar: Sesame will preserve AI companion personality as models improve
“They're making assistance. They're making utilities. I love those products. I use them all the time. They're great products. We want to make a companion. And so our prioritization of features and of, let's say, post training kind of personality, et cetera, wil…”
Ankit Kumar Mar 15, 2025 ▶ 1:14:19
Prediction Not checkable as stated
Kumar: Competitors will match Sesame's voice quality; there is no secret sauce
“The other companies, the other sort of chat products and so forth, they will get better voices. They're all, like, it's not gonna, we don't have some magical secret sauce on the technical side that is gonna be, like, impossible to replicate. They're gonna get …”
Ankit Kumar Mar 15, 2025 ▶ 1:14:58
Insight
Kumar: AI voice interface success depends on product experience over model size
“We think that that interface layer, it's really not kind of a core bigger, bigger models, better, better reasoning question. It's really a product experience question. It's really a question of, can you make a system that people actually want to interact with,…”
Ankit Kumar Mar 15, 2025 ▶ 1:19:16
Assertion Supported
Kumar: Sesame's AI voice companions currently cannot execute tasks
“Maya and Miles today, they can't do anything for you”
Ankit Kumar Mar 15, 2025 ▶ 1:26:00
Insight
Kumar: Multi-step AI agents need 99% reliability for daily adoption
“Doing, especially challenging, kind of multi-step things, you know, agents, as people say, I think to make that part of your everyday habits, it has to be like, 99%, you know, and right now, you know, every extra step the thing needs to take, There's some perc…”
Ankit Kumar Mar 15, 2025 ▶ 1:26:00
Opinion
Kumar: Not enough AI teams are focused on the interface layer
“But I don't think there's enough companies and kind of teams working on the interface layer.”
Ankit Kumar Mar 15, 2025 ▶ 1:29:11
Prediction Not checkable as stated
Kumar: Big tech companies will attempt to own conversational AI interface layer
“I think that over time, I think we will see more of these, you know, bigger companies trying to operate this layer. Like I said, I think that there is not enough effort on that right now, making these systems delightful to interact with, you know, and I think …”
Ankit Kumar Mar 15, 2025 ▶ 1:30:09
Disclosure
Kumar: The ChatGPT plugin system I built failed to fully take off
“I did this chat to be plugin system before, and I think it's still there probably. And it kind of didn't fully take off really.”
Ankit Kumar Mar 15, 2025 ▶ 1:32:29
Prediction Not checkable as stated
Kumar: Developer plugins will be essential to future AI interfaces
“I think that the models still need to get better basically to utilize plugins essentially in a way that's kind of reliable enough that someone will go out and look for a plugin for, you know, their kind of downstream service of choice because they just want, y…”
Ankit Kumar Mar 15, 2025 ▶ 1:32:53
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.