Jul 11, 2025 · 1h 4m · latent-space

Personalized AI Language Education — with Andrew Hsu, Speak

Andrew Hsu · 44m spoken Shawn Wang · 8m spoken Alessio Fanelli · 4m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Speak co-founder and CTO Andrew Hsu joins Latent Space to discuss building an AI-native language tutor, covering Speak's custom speech infrastructure, Korean market breakout, pedagogical methodology, and vision for general-purpose AI education.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 21.6% of the talking time here. How this is scored →

The hosts as informed peer 5.2 Guest teaching 4.9 Guest disagreement 1.5 The hosts pushing back 1.9
05100:0015:0030:0045:001:00:000:05–4:34 · The hosts as informed peer 4/10 Thiel Fellowship Beginnings and Early Background Swyx probes into Andrew Hsu's Thiel Fellowship background, playfully questioning his age and asking about famous cohort peers. Andrew politely clarifies timeline details about when he and other fellows like SBF and Vitalik participated.4:35–9:46 · The hosts as informed peer 6/10 Custom Speech Recognition Infrastructure and Latency Alessio and Swyx dive into onboarding UX, state machines, and latency sensitivity in early custom ASR systems. Andrew walks through the trade-offs between open conversational onboarding and guardrailed state machines.9:47–13:45 · The hosts as informed peer 5/10 Three Generations of Language Learning and the Korea Pivot Andrew outlines his taxonomy of three generations of language learning, distinguishing Rosetta Stone (Gen 1) and Duolingo's gamified mobile app (Gen 2) from AI-native functional fluency training (Gen 3). The hosts acknowledge Duolingo's strengths while listening to the pedagogical contrast.13:46–16:31 · The hosts as informed peer 5/10 Company Scale, B2B Growth, and Real-Time Translation vs. Learning Alessio asks about whether real-time translation tech (like Google Beam/Babelfish) will eliminate language learning. Andrew cleanly dismantles the premise using German sentence structure where verbs appear at the end, proving inherent latency constraints, and highlighting the human desire for direct connection.16:32–19:09 · The hosts as informed peer 5/10 Why South Korea Was the Ideal Launch Market Swyx notes the paradox of an American team winning Korea's intense English education market over local teams. Andrew attributes their success to relentless localization and high-density user demand.19:09–23:34 · The hosts as informed peer 4/10 The 'Speak Method' Pedagogy and Consumer Craft Andrew details building their own pedagogical framework and the unconventional decision to build out an engineering hub in Slovenia after finding key talent through referrals.23:34–25:59 · The hosts as informed peer 4/10 The Whisper Breakthrough and Evolution to AI Tutoring Andrew describes the pivotal moment OpenAI released Whisper in late 2022, proving speech recognition could accurately transcribe beginner non-native accents that human native speakers could not decipher.25:59–28:50 · The hosts as informed peer 5/10 Model Saturation and the Broader Future of Learning Alessio asks if foundational model progress makes custom startup work obsolete. Andrew explains the cycle of saturating model capabilities with product layers before scaling beyond language into general AI education.28:51–33:37 · The hosts as informed peer 6/10 Scaling Content Generation with Autonomous AI Agents Swyx and Andrew discuss curriculum generation pipelines, agent scaffolding, and real-world proficiency metrics versus standardized testing benchmarks.33:38–37:46 · The hosts as informed peer 6/10 Evaluation Frameworks and AI Content Leverage Alessio and Swyx explore evaluation frameworks, colloquial phrasing, and dialect variations like Mexican versus Argentine Spanish. Andrew explains prioritizing standard dialects before fine-tuning accents.37:47–42:36 · The hosts as informed peer 6/10 Contextual Learning, Wearables, and Future Hardware Swyx brings up real-world immersive learning tools and browser extensions, as well as future wearable devices from OpenAI and Jony Ive for continuous context capture.42:36–46:15 · The hosts as informed peer 5/10 Multimodal Learning: Video, Audio, and Generative UI Andrew outlines the future of multimodal language tutoring, combining generative UI, real-time image prompting, and dynamic synchronized audio tracks.46:16–52:10 · The hosts as informed peer 6/10 Real-Time Voice Architecture and Learner Voice Activity Detection Swyx suggests router models for multilingual TTS, but Andrew clarifies why subword code-switching makes naive routing fail. Andrew then forcefully reframes the industry obsession with TTFT latency as a vanity metric, educating on the critical role of domain-specific Voice Activity Detection (VAD) for hesitating language learners.52:10–56:20 · The hosts as informed peer 6/10 Internal AI Coding Culture and High-Agency Engineering Andrew discusses cultivating high-agency AI coding tool adoption in engineering. Swyx connects this back to the educational Bloom Two Sigma problem and pedagogical scaffolding.56:20–58:29 · The hosts as informed peer 5/10 Brand Building: Speak.com and Cultural Scale in Korea Alessio and Swyx discuss domain acquisitions and premium branding. Andrew shares the rationale behind buying Speak.com and building mainstream celebrity status in South Korea.58:29–1:00:49 · The hosts as informed peer 4/10 Startup History, Early EdTech, and AI Safety Guardrails Andrew reflects on early edtech startup attempts and shares experiences with user guardrail testing when first launching GPT-4 roleplays.1:00:49–1:03:56 · The hosts as informed peer 7/10 Broad Market Demand and Universal Demographics Andrew presents his contrarian take that despite AI advances, real-world societal inertia means everyday life has barely changed outside Silicon Valley. Swyx pushes back with the safety perspective of 'slow takeoff, short timeline' as an optimal outcome.0:05–4:34 · Guest teaching 3/10 Thiel Fellowship Beginnings and Early Background Swyx probes into Andrew Hsu's Thiel Fellowship background, playfully questioning his age and asking about famous cohort peers. Andrew politely clarifies timeline details about when he and other fellows like SBF and Vitalik participated.4:35–9:46 · Guest teaching 4/10 Custom Speech Recognition Infrastructure and Latency Alessio and Swyx dive into onboarding UX, state machines, and latency sensitivity in early custom ASR systems. Andrew walks through the trade-offs between open conversational onboarding and guardrailed state machines.9:47–13:45 · Guest teaching 6/10 Three Generations of Language Learning and the Korea Pivot Andrew outlines his taxonomy of three generations of language learning, distinguishing Rosetta Stone (Gen 1) and Duolingo's gamified mobile app (Gen 2) from AI-native functional fluency training (Gen 3). The hosts acknowledge Duolingo's strengths while listening to the pedagogical contrast.13:46–16:31 · Guest teaching 7/10 Company Scale, B2B Growth, and Real-Time Translation vs. Learning Alessio asks about whether real-time translation tech (like Google Beam/Babelfish) will eliminate language learning. Andrew cleanly dismantles the premise using German sentence structure where verbs appear at the end, proving inherent latency constraints, and highlighting the human desire for direct connection.16:32–19:09 · Guest teaching 5/10 Why South Korea Was the Ideal Launch Market Swyx notes the paradox of an American team winning Korea's intense English education market over local teams. Andrew attributes their success to relentless localization and high-density user demand.19:09–23:34 · Guest teaching 4/10 The 'Speak Method' Pedagogy and Consumer Craft Andrew details building their own pedagogical framework and the unconventional decision to build out an engineering hub in Slovenia after finding key talent through referrals.23:34–25:59 · Guest teaching 6/10 The Whisper Breakthrough and Evolution to AI Tutoring Andrew describes the pivotal moment OpenAI released Whisper in late 2022, proving speech recognition could accurately transcribe beginner non-native accents that human native speakers could not decipher.25:59–28:50 · Guest teaching 5/10 Model Saturation and the Broader Future of Learning Alessio asks if foundational model progress makes custom startup work obsolete. Andrew explains the cycle of saturating model capabilities with product layers before scaling beyond language into general AI education.28:51–33:37 · Guest teaching 5/10 Scaling Content Generation with Autonomous AI Agents Swyx and Andrew discuss curriculum generation pipelines, agent scaffolding, and real-world proficiency metrics versus standardized testing benchmarks.33:38–37:46 · Guest teaching 5/10 Evaluation Frameworks and AI Content Leverage Alessio and Swyx explore evaluation frameworks, colloquial phrasing, and dialect variations like Mexican versus Argentine Spanish. Andrew explains prioritizing standard dialects before fine-tuning accents.37:47–42:36 · Guest teaching 5/10 Contextual Learning, Wearables, and Future Hardware Swyx brings up real-world immersive learning tools and browser extensions, as well as future wearable devices from OpenAI and Jony Ive for continuous context capture.42:36–46:15 · Guest teaching 5/10 Multimodal Learning: Video, Audio, and Generative UI Andrew outlines the future of multimodal language tutoring, combining generative UI, real-time image prompting, and dynamic synchronized audio tracks.46:16–52:10 · Guest teaching 8/10 Real-Time Voice Architecture and Learner Voice Activity Detection Swyx suggests router models for multilingual TTS, but Andrew clarifies why subword code-switching makes naive routing fail. Andrew then forcefully reframes the industry obsession with TTFT latency as a vanity metric, educating on the critical role of domain-specific Voice Activity Detection (VAD) for hesitating language learners.52:10–56:20 · Guest teaching 4/10 Internal AI Coding Culture and High-Agency Engineering Andrew discusses cultivating high-agency AI coding tool adoption in engineering. Swyx connects this back to the educational Bloom Two Sigma problem and pedagogical scaffolding.56:20–58:29 · Guest teaching 3/10 Brand Building: Speak.com and Cultural Scale in Korea Alessio and Swyx discuss domain acquisitions and premium branding. Andrew shares the rationale behind buying Speak.com and building mainstream celebrity status in South Korea.58:29–1:00:49 · Guest teaching 4/10 Startup History, Early EdTech, and AI Safety Guardrails Andrew reflects on early edtech startup attempts and shares experiences with user guardrail testing when first launching GPT-4 roleplays.1:00:49–1:03:56 · Guest teaching 4/10 Broad Market Demand and Universal Demographics Andrew presents his contrarian take that despite AI advances, real-world societal inertia means everyday life has barely changed outside Silicon Valley. Swyx pushes back with the safety perspective of 'slow takeoff, short timeline' as an optimal outcome.0:05–4:34 · Guest disagreement 1/10 Thiel Fellowship Beginnings and Early Background Swyx probes into Andrew Hsu's Thiel Fellowship background, playfully questioning his age and asking about famous cohort peers. Andrew politely clarifies timeline details about when he and other fellows like SBF and Vitalik participated.4:35–9:46 · Guest disagreement 1/10 Custom Speech Recognition Infrastructure and Latency Alessio and Swyx dive into onboarding UX, state machines, and latency sensitivity in early custom ASR systems. Andrew walks through the trade-offs between open conversational onboarding and guardrailed state machines.9:47–13:45 · Guest disagreement 2/10 Three Generations of Language Learning and the Korea Pivot Andrew outlines his taxonomy of three generations of language learning, distinguishing Rosetta Stone (Gen 1) and Duolingo's gamified mobile app (Gen 2) from AI-native functional fluency training (Gen 3). The hosts acknowledge Duolingo's strengths while listening to the pedagogical contrast.13:46–16:31 · Guest disagreement 3/10 Company Scale, B2B Growth, and Real-Time Translation vs. Learning Alessio asks about whether real-time translation tech (like Google Beam/Babelfish) will eliminate language learning. Andrew cleanly dismantles the premise using German sentence structure where verbs appear at the end, proving inherent latency constraints, and highlighting the human desire for direct connection.16:32–19:09 · Guest disagreement 1/10 Why South Korea Was the Ideal Launch Market Swyx notes the paradox of an American team winning Korea's intense English education market over local teams. Andrew attributes their success to relentless localization and high-density user demand.19:09–23:34 · Guest disagreement 1/10 The 'Speak Method' Pedagogy and Consumer Craft Andrew details building their own pedagogical framework and the unconventional decision to build out an engineering hub in Slovenia after finding key talent through referrals.23:34–25:59 · Guest disagreement 1/10 The Whisper Breakthrough and Evolution to AI Tutoring Andrew describes the pivotal moment OpenAI released Whisper in late 2022, proving speech recognition could accurately transcribe beginner non-native accents that human native speakers could not decipher.25:59–28:50 · Guest disagreement 1/10 Model Saturation and the Broader Future of Learning Alessio asks if foundational model progress makes custom startup work obsolete. Andrew explains the cycle of saturating model capabilities with product layers before scaling beyond language into general AI education.28:51–33:37 · Guest disagreement 2/10 Scaling Content Generation with Autonomous AI Agents Swyx and Andrew discuss curriculum generation pipelines, agent scaffolding, and real-world proficiency metrics versus standardized testing benchmarks.33:38–37:46 · Guest disagreement 1/10 Evaluation Frameworks and AI Content Leverage Alessio and Swyx explore evaluation frameworks, colloquial phrasing, and dialect variations like Mexican versus Argentine Spanish. Andrew explains prioritizing standard dialects before fine-tuning accents.37:47–42:36 · Guest disagreement 1/10 Contextual Learning, Wearables, and Future Hardware Swyx brings up real-world immersive learning tools and browser extensions, as well as future wearable devices from OpenAI and Jony Ive for continuous context capture.42:36–46:15 · Guest disagreement 1/10 Multimodal Learning: Video, Audio, and Generative UI Andrew outlines the future of multimodal language tutoring, combining generative UI, real-time image prompting, and dynamic synchronized audio tracks.46:16–52:10 · Guest disagreement 4/10 Real-Time Voice Architecture and Learner Voice Activity Detection Swyx suggests router models for multilingual TTS, but Andrew clarifies why subword code-switching makes naive routing fail. Andrew then forcefully reframes the industry obsession with TTFT latency as a vanity metric, educating on the critical role of domain-specific Voice Activity Detection (VAD) for hesitating language learners.52:10–56:20 · Guest disagreement 1/10 Internal AI Coding Culture and High-Agency Engineering Andrew discusses cultivating high-agency AI coding tool adoption in engineering. Swyx connects this back to the educational Bloom Two Sigma problem and pedagogical scaffolding.56:20–58:29 · Guest disagreement 1/10 Brand Building: Speak.com and Cultural Scale in Korea Alessio and Swyx discuss domain acquisitions and premium branding. Andrew shares the rationale behind buying Speak.com and building mainstream celebrity status in South Korea.58:29–1:00:49 · Guest disagreement 1/10 Startup History, Early EdTech, and AI Safety Guardrails Andrew reflects on early edtech startup attempts and shares experiences with user guardrail testing when first launching GPT-4 roleplays.1:00:49–1:03:56 · Guest disagreement 2/10 Broad Market Demand and Universal Demographics Andrew presents his contrarian take that despite AI advances, real-world societal inertia means everyday life has barely changed outside Silicon Valley. Swyx pushes back with the safety perspective of 'slow takeoff, short timeline' as an optimal outcome.0:05–4:34 · The hosts pushing back 2/10 Thiel Fellowship Beginnings and Early Background Swyx probes into Andrew Hsu's Thiel Fellowship background, playfully questioning his age and asking about famous cohort peers. Andrew politely clarifies timeline details about when he and other fellows like SBF and Vitalik participated.4:35–9:46 · The hosts pushing back 2/10 Custom Speech Recognition Infrastructure and Latency Alessio and Swyx dive into onboarding UX, state machines, and latency sensitivity in early custom ASR systems. Andrew walks through the trade-offs between open conversational onboarding and guardrailed state machines.9:47–13:45 · The hosts pushing back 1/10 Three Generations of Language Learning and the Korea Pivot Andrew outlines his taxonomy of three generations of language learning, distinguishing Rosetta Stone (Gen 1) and Duolingo's gamified mobile app (Gen 2) from AI-native functional fluency training (Gen 3). The hosts acknowledge Duolingo's strengths while listening to the pedagogical contrast.13:46–16:31 · The hosts pushing back 3/10 Company Scale, B2B Growth, and Real-Time Translation vs. Learning Alessio asks about whether real-time translation tech (like Google Beam/Babelfish) will eliminate language learning. Andrew cleanly dismantles the premise using German sentence structure where verbs appear at the end, proving inherent latency constraints, and highlighting the human desire for direct connection.16:32–19:09 · The hosts pushing back 2/10 Why South Korea Was the Ideal Launch Market Swyx notes the paradox of an American team winning Korea's intense English education market over local teams. Andrew attributes their success to relentless localization and high-density user demand.19:09–23:34 · The hosts pushing back 1/10 The 'Speak Method' Pedagogy and Consumer Craft Andrew details building their own pedagogical framework and the unconventional decision to build out an engineering hub in Slovenia after finding key talent through referrals.23:34–25:59 · The hosts pushing back 1/10 The Whisper Breakthrough and Evolution to AI Tutoring Andrew describes the pivotal moment OpenAI released Whisper in late 2022, proving speech recognition could accurately transcribe beginner non-native accents that human native speakers could not decipher.25:59–28:50 · The hosts pushing back 2/10 Model Saturation and the Broader Future of Learning Alessio asks if foundational model progress makes custom startup work obsolete. Andrew explains the cycle of saturating model capabilities with product layers before scaling beyond language into general AI education.28:51–33:37 · The hosts pushing back 2/10 Scaling Content Generation with Autonomous AI Agents Swyx and Andrew discuss curriculum generation pipelines, agent scaffolding, and real-world proficiency metrics versus standardized testing benchmarks.33:38–37:46 · The hosts pushing back 2/10 Evaluation Frameworks and AI Content Leverage Alessio and Swyx explore evaluation frameworks, colloquial phrasing, and dialect variations like Mexican versus Argentine Spanish. Andrew explains prioritizing standard dialects before fine-tuning accents.37:47–42:36 · The hosts pushing back 2/10 Contextual Learning, Wearables, and Future Hardware Swyx brings up real-world immersive learning tools and browser extensions, as well as future wearable devices from OpenAI and Jony Ive for continuous context capture.42:36–46:15 · The hosts pushing back 1/10 Multimodal Learning: Video, Audio, and Generative UI Andrew outlines the future of multimodal language tutoring, combining generative UI, real-time image prompting, and dynamic synchronized audio tracks.46:16–52:10 · The hosts pushing back 3/10 Real-Time Voice Architecture and Learner Voice Activity Detection Swyx suggests router models for multilingual TTS, but Andrew clarifies why subword code-switching makes naive routing fail. Andrew then forcefully reframes the industry obsession with TTFT latency as a vanity metric, educating on the critical role of domain-specific Voice Activity Detection (VAD) for hesitating language learners.52:10–56:20 · The hosts pushing back 2/10 Internal AI Coding Culture and High-Agency Engineering Andrew discusses cultivating high-agency AI coding tool adoption in engineering. Swyx connects this back to the educational Bloom Two Sigma problem and pedagogical scaffolding.56:20–58:29 · The hosts pushing back 1/10 Brand Building: Speak.com and Cultural Scale in Korea Alessio and Swyx discuss domain acquisitions and premium branding. Andrew shares the rationale behind buying Speak.com and building mainstream celebrity status in South Korea.58:29–1:00:49 · The hosts pushing back 1/10 Startup History, Early EdTech, and AI Safety Guardrails Andrew reflects on early edtech startup attempts and shares experiences with user guardrail testing when first launching GPT-4 roleplays.1:00:49–1:03:56 · The hosts pushing back 4/10 Broad Market Demand and Universal Demographics Andrew presents his contrarian take that despite AI advances, real-world societal inertia means everyday life has barely changed outside Silicon Valley. Swyx pushes back with the safety perspective of 'slow takeoff, short timeline' as an optimal outcome.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 29.1% · guest 70.9%0:00 · the hosts 29.1% · guest 70.9%3:00 · the hosts 13.5% · guest 86.5%3:00 · the hosts 13.5% · guest 86.5%6:00 · the hosts 31.5% · guest 68.5%6:00 · the hosts 31.5% · guest 68.5%9:00 · the hosts 22.3% · guest 77.7%9:00 · the hosts 22.3% · guest 77.7%12:00 · the hosts 6.3% · guest 93.7%12:00 · the hosts 6.3% · guest 93.7%15:00 · the hosts 26.5% · guest 73.5%15:00 · the hosts 26.5% · guest 73.5%18:00 · the hosts 13.9% · guest 86.1%18:00 · the hosts 13.9% · guest 86.1%21:00 · the hosts 21.4% · guest 78.6%21:00 · the hosts 21.4% · guest 78.6%24:00 · the hosts 13.3% · guest 86.7%24:00 · the hosts 13.3% · guest 86.7%27:00 · the hosts 11.3% · guest 88.7%27:00 · the hosts 11.3% · guest 88.7%30:00 · the hosts 8.5% · guest 91.5%30:00 · the hosts 8.5% · guest 91.5%33:00 · the hosts 35.6% · guest 64.4%33:00 · the hosts 35.6% · guest 64.4%36:00 · the hosts 29.4% · guest 70.6%36:00 · the hosts 29.4% · guest 70.6%39:00 · the hosts 23.6% · guest 76.4%39:00 · the hosts 23.6% · guest 76.4%42:00 · the hosts 28.5% · guest 71.5%42:00 · the hosts 28.5% · guest 71.5%45:00 · the hosts 26.4% · guest 73.6%45:00 · the hosts 26.4% · guest 73.6%48:00 · the hosts 11.8% · guest 88.2%48:00 · the hosts 11.8% · guest 88.2%51:00 · the hosts 8.5% · guest 91.5%51:00 · the hosts 8.5% · guest 91.5%54:00 · the hosts 50.1% · guest 49.9%54:00 · the hosts 50.1% · guest 49.9%57:00 · the hosts 16.2% · guest 83.8%57:00 · the hosts 16.2% · guest 83.8%1:00:00 · the hosts 19.1% · guest 80.9%1:00:00 · the hosts 19.1% · guest 80.9%1:03:00 · the hosts 52.5% · guest 47.5%1:03:00 · the hosts 52.5% · guest 47.5%
Sharpest disagreement ▶ 50:46 Calling standard voice latency a vanity metric

Andrew strongly dismisses conventional industry benchmarking around time-to-first-audio, calling it a vanity metric that ignores realistic learner turn-detection and hesitation.

Hardest push from the hosts ▶ 1:03:19 Defending the slow takeoff dynamic

Swyx challenges Andrew's frustration with slow real-world AI adoption, arguing from AI safety principles that a slow takeoff gives humanity necessary preparation time.

Biggest teaching moment ▶ 15:13 Linguistic grammar barriers to real-time translation

Andrew educates the hosts on why universal real-time translation has unavoidable latency barriers due to sentence structures like German clause-final verbs.

The host holds their own ▶ 54:51 Synthesizing the Bloom Two Sigma tutoring framework

Swyx brings deep pedagogical theory to the table, framing Speak's knowledge graph architecture within Bloom's Two Sigma problem of level-adjusting mastery learning.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Thiel Fellowship Beginnings and Early Background 4312 Swyx probes into Andrew Hsu's Thiel Fellowship background, playfully questioning his age and asking about famous cohort peers. Andrew politely clarifies timeline details about when he and other fellows like SBF and Vitalik participated.
Custom Speech Recognition Infrastructure and Latency 6412 Alessio and Swyx dive into onboarding UX, state machines, and latency sensitivity in early custom ASR systems. Andrew walks through the trade-offs between open conversational onboarding and guardrailed state machines.
Three Generations of Language Learning and the Korea Pivot 5621 Andrew outlines his taxonomy of three generations of language learning, distinguishing Rosetta Stone (Gen 1) and Duolingo's gamified mobile app (Gen 2) from AI-native functional fluency training (Gen 3). The hosts acknowledge Duolingo's strengths while listening to the pedagogical contrast.
Company Scale, B2B Growth, and Real-Time Translation vs. Learning 5733 Alessio asks about whether real-time translation tech (like Google Beam/Babelfish) will eliminate language learning. Andrew cleanly dismantles the premise using German sentence structure where verbs appear at the end, proving inherent latency constraints, and highlighting the human desire for direct connection.
Why South Korea Was the Ideal Launch Market 5512 Swyx notes the paradox of an American team winning Korea's intense English education market over local teams. Andrew attributes their success to relentless localization and high-density user demand.
The 'Speak Method' Pedagogy and Consumer Craft 4411 Andrew details building their own pedagogical framework and the unconventional decision to build out an engineering hub in Slovenia after finding key talent through referrals.
The Whisper Breakthrough and Evolution to AI Tutoring 4611 Andrew describes the pivotal moment OpenAI released Whisper in late 2022, proving speech recognition could accurately transcribe beginner non-native accents that human native speakers could not decipher.
Model Saturation and the Broader Future of Learning 5512 Alessio asks if foundational model progress makes custom startup work obsolete. Andrew explains the cycle of saturating model capabilities with product layers before scaling beyond language into general AI education.
Scaling Content Generation with Autonomous AI Agents 6522 Swyx and Andrew discuss curriculum generation pipelines, agent scaffolding, and real-world proficiency metrics versus standardized testing benchmarks.
Evaluation Frameworks and AI Content Leverage 6512 Alessio and Swyx explore evaluation frameworks, colloquial phrasing, and dialect variations like Mexican versus Argentine Spanish. Andrew explains prioritizing standard dialects before fine-tuning accents.
Contextual Learning, Wearables, and Future Hardware 6512 Swyx brings up real-world immersive learning tools and browser extensions, as well as future wearable devices from OpenAI and Jony Ive for continuous context capture.
Multimodal Learning: Video, Audio, and Generative UI 5511 Andrew outlines the future of multimodal language tutoring, combining generative UI, real-time image prompting, and dynamic synchronized audio tracks.
Real-Time Voice Architecture and Learner Voice Activity Detection 6843 Swyx suggests router models for multilingual TTS, but Andrew clarifies why subword code-switching makes naive routing fail. Andrew then forcefully reframes the industry obsession with TTFT latency as a vanity metric, educating on the critical role of domain-specific Voice Activity Detection (VAD) for hesitating language learners.
Internal AI Coding Culture and High-Agency Engineering 6412 Andrew discusses cultivating high-agency AI coding tool adoption in engineering. Swyx connects this back to the educational Bloom Two Sigma problem and pedagogical scaffolding.
Brand Building: Speak.com and Cultural Scale in Korea 5311 Alessio and Swyx discuss domain acquisitions and premium branding. Andrew shares the rationale behind buying Speak.com and building mainstream celebrity status in South Korea.
Startup History, Early EdTech, and AI Safety Guardrails 4411 Andrew reflects on early edtech startup attempts and shares experiences with user guardrail testing when first launching GPT-4 roleplays.
Broad Market Demand and Universal Demographics 7424 Andrew presents his contrarian take that despite AI advances, real-world societal inertia means everyday life has barely changed outside Silicon Valley. Swyx pushes back with the safety perspective of 'slow takeoff, short timeline' as an optimal outcome.

Statements from this episode (40)

Assertion Not checkable as stated
Hsu: Speak Has Never Pivoted Since Inception
“We actually like never pivoted.”
Andrew Hsu Jul 11, 2025 ▶ 3:55
Opinion
Hsu: 80% to 90% of Tech for Superhuman AI Tutors Exists
“It was that as speech models and language models become superhuman, that would let us create an AI language tutor that would help you become fluent faster than any human could. And I think we're like 80 to 90% of the tech is here now.”
Andrew Hsu Jul 11, 2025 ▶ 4:18
Disclosure
Hsu: Speak achieved product-market fit in South Korea around 2019-2020
“Before 20, 22, when Whisper came out, when ChatGPT came out in the years, Before that, you know, like roughly two to three years is when we feel like we found PMF in South Korea and then started growing still only in that market, still only teaching English.”
Andrew Hsu Jul 11, 2025 ▶ 5:06
Disclosure
Hsu: Speak runs custom ASR for core loops alongside Whisper for tutoring
“There's many other sort of product surfaces within the app today that are more LM powered, where it's more open-ended, real tutoring, where we actually give you feedback on what you said in the semantics and so on. So that stuff is more like whisper powered, m…”
Andrew Hsu Jul 11, 2025 ▶ 5:43
Assertion Not checkable as stated
Voice onboarding lowers app signups but boosts trial start rates
“The interesting thing is that in general, because it's speaking based, which is a much higher barrier than just like tapping a multiple choice button, what we see is that install to signup rate is a decent amount lower. But trial start rate is higher.”
Andrew Hsu Jul 11, 2025 ▶ 7:10
Opinion
Hsu: Conversational Onboarding Should Transition from State Machines to Natural LLMs
“I think that things should move in a direction where it's much more of a natural conversation.”
Andrew Hsu Jul 11, 2025 ▶ 8:40
Disclosure
Hsu: Speak Summarizes User Onboarding Audio into Structured Goals via Secondary LLMs
“The way it works is the tutor will ask you some sort of question, like, what are your goals around learning English or the language? And then we will basically use a separate LM prompt to summarize. So it's not the, like, full transcript for what you said that…”
Andrew Hsu Jul 11, 2025 ▶ 9:27
Opinion
Hsu: Duolingo and Gen 2 language apps are comparable to mobile games
“Gen two was basically mobile. So you have these very casual, massively popular mobile apps like Duolingo that I think the comp there is probably closer to a mobile game, something that feels productive, something that's very engaging, very gamified.”
Andrew Hsu Jul 11, 2025 ▶ 10:38
Insight
Hsu: Learning app users suffer decision fatigue and need guided tracks
“We realized people don't want to choose. They're already using some of their motivation on a daily basis just to open the app. They don't want to make another choice after that, right? Just tell me what to do, right? Like, you know, give me a big button and th…”
Andrew Hsu Jul 11, 2025 ▶ 12:31
Insight
Hsu: Dropping a free tier sidesteps user motivation problems
“We also pretty critically, I think, abandoned the free version and just went straight premium. And we kind of sidestepped the motivation question that way, because we knew that there were a ton of users that really wanted to learn English and were already real…”
Andrew Hsu Jul 11, 2025 ▶ 12:48
Assertion Partly supported
Hsu: Speak is the largest English learning app in South Korea
“So we're now the biggest English app in South Korea.”
Andrew Hsu Jul 11, 2025 ▶ 13:56
Assertion Supported
Speak has surpassed $50 million in annual recurring revenue
“In terms of revenue scale, well over fifty million ARR.”
Andrew Hsu Jul 11, 2025 ▶ 14:29
Insight
Hsu: Real-time translation latency is blocked by language syntax rules
“The counterexample that I always have that I think is quite illustrative is in German, the verb is at the end of the sentence. So if you're trying to do real-time translation from German to English, as an example, you can't actually make any progress on the En…”
Andrew Hsu Jul 11, 2025 ▶ 15:24
Insight
Hsu: Asian language learners seek direct human connection, not AI translators
“If you talk to All of our users in Asia. They don't want a translator. The reason that they are trying to learn English is to make themselves a better person, to connect with other people. Like, they want to be able to look you in the eye and speak English, sp…”
Andrew Hsu Jul 11, 2025 ▶ 15:54
Insight
Hsu: Winning South Korea's crowded tutor market proves transferable PMF
“And our logic was basically, if we can really make headway and win this market that is chock full of these human competitor products and all these people that fundamentally care about fluency, then we probably have something pretty real and strong PMF that we …”
Andrew Hsu Jul 11, 2025 ▶ 17:51
Assertion Not checkable as stated
Hsu: Early Korean users were shocked Speak was an American company
“We had a lot of reports from users pretty early on that they were shocked that it was an American company.”
Andrew Hsu Jul 11, 2025 ▶ 18:53
Disclosure
Hsu: Speak Built All Curriculum In-House via Proprietary 'Speak Method'
“Another thing we did that I forgot to mention was we decided we needed to fully own all the content. So the way that we teach all in-house, all sort of thought from first principles, We built this thing called the speak method, which is basically like a pedago…”
Andrew Hsu Jul 11, 2025 ▶ 19:24
Disclosure
Hsu: Speak has 90% of product team in SF and only hires there
“Now we have 90% of our core product development team in San Francisco here. Office in Fideye, we're really only hiring here.”
Andrew Hsu Jul 11, 2025 ▶ 22:06
What-if
Hsu: If building Speak again, he probably wouldn't build remote hub in Slovenia
“So it worked out, but if I had to do it over again, I probably wouldn't do it.”
Andrew Hsu Jul 11, 2025 ▶ 23:26
Assertion Not checkable as stated
Hsu: Whisper outperformed human listeners on accented Korean English clips
“There were four of us in the room, we all closed our eyes, and none of us had any idea, and the model got it right. So, I mean, superhuman.”
Andrew Hsu Jul 11, 2025 ▶ 24:23
Assertion Not checkable as stated
Hsu: Speak reached several million ARR in Korea pre-Whisper
“Still a great product, by the way, you know, still grew to like several million ARR in South Korea.”
Andrew Hsu Jul 11, 2025 ▶ 25:05
Disclosure
Hsu: Speak Is Building a Real-Time Voice Platform for Advanced Lessons
“We're actively building out a real-time voice platform that we can build a lot of more verticalized specific lesson experiences on top of that I'm super, super excited about. I don't think they're going to replace our current lessons. They're going to be more …”
Andrew Hsu Jul 11, 2025 ▶ 27:23
Disclosure
Hsu: Speak runs agentic LLM pipeline to generate curricula and lessons
“We have a tutor agent, we have a curriculum writing agent, we have a giant LM based pipeline that creates curriculum, scaffolds it in the right way, writes the lessons themselves.”
Andrew Hsu Jul 11, 2025 ▶ 30:08
Insight
Hsu: The Frontier of Language Fluency Is Highly Jagged Across Contexts
“You might be really good at that, but be completely unable to, like, talk about your family, right? So the frontier of fluency is very jagged, but”
Andrew Hsu Jul 11, 2025 ▶ 30:58
Disclosure
Speak Building 'Speak Score' Aggregating Fluency Across Knowledge Graphs
“There are aspects of it that are live and it's a very sort of multidimensional system where we think of it as there are many aspects of fluency, right? There's many sub scores and we have a few of them that are currently live and we're actively developing othe…”
Andrew Hsu Jul 11, 2025 ▶ 31:40
Prediction Not checkable as stated
Hsu: Top Curriculum-Writing AI Agents Will Likely Rely on Reinforcement Fine-Tuning
“And I also think like in the future, a really good curriculum or lesson writer agent will probably be like reinforcement fine-tuned on a lot of our internal data as well.”
Andrew Hsu Jul 11, 2025 ▶ 34:32
Insight
Hsu: AI Gives Content and Engineering 100x Leverage While Still Requiring Review
“The way that we see it, really not just for our content team members, but also I think it's perfectly applicable to engineering is that it's leverage. It just allows you to do a hundred X in the same amount of time. We still need human review of the syllabus, …”
Andrew Hsu Jul 11, 2025 ▶ 35:00
Disclosure
Hsu: Speak focuses on casual conversational language over textbook English
“That's one of our fundamental tenets, which is that we don't teach textbook English or textbook language. Like we try very hard to teach Gen Z slang. We don't go quite that far, but slaking. We try to teach Very casual conversational language that is actually …”
Andrew Hsu Jul 11, 2025 ▶ 35:38
Insight
Hsu: Spontaneous communication ability is almost fully orthogonal to pronunciation
“Communication and your ability to speak spontaneously and get us on a concept across, an idea across, is almost fully orthogonal to pronunciation. You can be really bad at pronunciation, but still communicate effectively.”
Andrew Hsu Jul 11, 2025 ▶ 37:50
Insight
Hsu: Language learners face high psychological barrier practicing in front of humans
“It turns out there's like a really key psychological barrier there where people are just not willing to do this in front of a human, even if it's a human that is a teacher that you're paying, right?”
Andrew Hsu Jul 11, 2025 ▶ 38:25
Disclosure
Speak's pronunciation coach fine-tunes Meta's wav2vec on proprietary phonetic transcripts
“We have for English only right now, a pronunciation coach that is basically like a fine-tuned version of WaveDeVec, which is a meta model, but we basically fine-tune it on a bunch of our own phonetic transcripts, like fine-tuned data.”
Andrew Hsu Jul 11, 2025 ▶ 38:59
Assertion Not checkable as stated
Hsu: Only a few TTS models handle multilingual code-switching properly
“It's actually like only a few models are able to speak two languages in the same sentence and then pronounce them properly.”
Andrew Hsu Jul 11, 2025 ▶ 47:33
Disclosure
Speak has no OpenAI Realtime API features in production due to cost
“We are, In the process of building a variety of experiences on top of the real-time API, I want to clarify that actually nothing is in production yet, mostly for price reasons, frankly.”
Andrew Hsu Jul 11, 2025 ▶ 48:05
Opinion
Hsu: Realtime API pricing fits customer support displacement, not consumer apps
“The pricing model of the real-time API makes more sense for something like a customer support agent, where you're very directly replacing somebody that you would pay hourly otherwise, and that's how you're seeing the pricing model for a lot of these initial ag…”
Andrew Hsu Jul 11, 2025 ▶ 48:17
Insight
Hsu: Standard semantic VAD fails completely for language learners
“You can use like the semantic VAD on real time API for regular English conversation. And that will basically classify at every token, how likely it is that you're done speaking as a sort of normal conversational English speaker. Like in this conversation, that…”
Andrew Hsu Jul 11, 2025 ▶ 51:31
Disclosure
Hsu: Speak mandates AI coding tools as default engineering workflow
“We try to set a culture in the engineering team where usage of these tools as much as possible and as a default path is the expectation. And in hiring We are now explicitly asking about this a lot, thinking about what are the types of people that are going to …”
Andrew Hsu Jul 11, 2025 ▶ 52:53
Insight
Hsu: English Language Learning Follows a Linear Path Until Intermediate Levels
“Beginner to intermediate English learners actually like All need to know a bunch of similar concepts. It isn't really until you get intermediate and more advanced where that starts to, like, more sharply diverge and A-zero through B-one, I would say. There's a…”
Andrew Hsu Jul 11, 2025 ▶ 55:19
Assertion Not checkable as stated
Hsu: Early Speak GPT-4 role-play users submitted questionable custom scenarios
“In 2023, when we first launched our AI role plays using GPT-IV, back then people were way more concerned about safety, right? And obviously the models now are much better at like refusals and line sharper between what's appropriate and not. But we did see a lo…”
Andrew Hsu Jul 11, 2025 ▶ 1:00:14
Assertion Not checkable as stated
Hsu: Speak's demographic sweet spot in Korea is 25-45 white-collar professionals
“We have a sweet spot in Korea. It's like, 25 to 45, more professional, more white collar, but it's very, it's like a very long tail on either side.”
Andrew Hsu Jul 11, 2025 ▶ 1:01:35
Opinion
Hsu: Real-World Inertia Leaves AI's Daily Impact Outside Bay Area Near Zero
“And I think if you, like, go to another state outside of the Bay Area, probably even in California, outside of the Bay Area, and then you ask somebody how much their life has materially changed, it's, like, pretty close to zero. Real-world inertia is enormous.…”
Andrew Hsu Jul 11, 2025 ▶ 1:02:23
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.