Everything Andrew Hsu said on any show that made the record, most notable first. Each card names its show and opens the statement there.
Hsu: Real-time translation latency is blocked by language syntax rules
“The counterexample that I always have that I think is quite illustrative is in German, the verb is at the end of the sentence. So if you're trying to do real-time translation from German to English, as an example, you can't actually make any progress on the En…”
Hsu: 80% to 90% of Tech for Superhuman AI Tutors Exists
“It was that as speech models and language models become superhuman, that would let us create an AI language tutor that would help you become fluent faster than any human could. And I think we're like 80 to 90% of the tech is here now.”
Voice onboarding lowers app signups but boosts trial start rates
“The interesting thing is that in general, because it's speaking based, which is a much higher barrier than just like tapping a multiple choice button, what we see is that install to signup rate is a decent amount lower. But trial start rate is higher.”
Hsu: Conversational Onboarding Should Transition from State Machines to Natural LLMs
“I think that things should move in a direction where it's much more of a natural conversation.”
Hsu: Duolingo and Gen 2 language apps are comparable to mobile games
“Gen two was basically mobile. So you have these very casual, massively popular mobile apps like Duolingo that I think the comp there is probably closer to a mobile game, something that feels productive, something that's very engaging, very gamified.”
Hsu: Dropping a free tier sidesteps user motivation problems
“We also pretty critically, I think, abandoned the free version and just went straight premium. And we kind of sidestepped the motivation question that way, because we knew that there were a ton of users that really wanted to learn English and were already real…”
Hsu: Speak is the largest English learning app in South Korea
“So we're now the biggest English app in South Korea.”
Speak has surpassed $50 million in annual recurring revenue
“In terms of revenue scale, well over fifty million ARR.”
Hsu: If building Speak again, he probably wouldn't build remote hub in Slovenia
“So it worked out, but if I had to do it over again, I probably wouldn't do it.”
Hsu: Whisper outperformed human listeners on accented Korean English clips
“There were four of us in the room, we all closed our eyes, and none of us had any idea, and the model got it right. So, I mean, superhuman.”
Hsu: Spontaneous communication ability is almost fully orthogonal to pronunciation
“Communication and your ability to speak spontaneously and get us on a concept across, an idea across, is almost fully orthogonal to pronunciation. You can be really bad at pronunciation, but still communicate effectively.”
Hsu: Standard semantic VAD fails completely for language learners
“You can use like the semantic VAD on real time API for regular English conversation. And that will basically classify at every token, how likely it is that you're done speaking as a sort of normal conversational English speaker. Like in this conversation, that…”
Hsu: Real-World Inertia Leaves AI's Daily Impact Outside Bay Area Near Zero
“And I think if you, like, go to another state outside of the Bay Area, probably even in California, outside of the Bay Area, and then you ask somebody how much their life has materially changed, it's, like, pretty close to zero. Real-world inertia is enormous.…”
Hsu: Learning app users suffer decision fatigue and need guided tracks
“We realized people don't want to choose. They're already using some of their motivation on a daily basis just to open the app. They don't want to make another choice after that, right? Just tell me what to do, right? Like, you know, give me a big button and th…”
Hsu: Asian language learners seek direct human connection, not AI translators
“If you talk to All of our users in Asia. They don't want a translator. The reason that they are trying to learn English is to make themselves a better person, to connect with other people. Like, they want to be able to look you in the eye and speak English, sp…”
Hsu: Winning South Korea's crowded tutor market proves transferable PMF
“And our logic was basically, if we can really make headway and win this market that is chock full of these human competitor products and all these people that fundamentally care about fluency, then we probably have something pretty real and strong PMF that we …”
Hsu: Speak has 90% of product team in SF and only hires there
“Now we have 90% of our core product development team in San Francisco here. Office in Fideye, we're really only hiring here.”
Hsu: Speak reached several million ARR in Korea pre-Whisper
“Still a great product, by the way, you know, still grew to like several million ARR in South Korea.”
Hsu: The Frontier of Language Fluency Is Highly Jagged Across Contexts
“You might be really good at that, but be completely unable to, like, talk about your family, right? So the frontier of fluency is very jagged, but”
Hsu: Top Curriculum-Writing AI Agents Will Likely Rely on Reinforcement Fine-Tuning
“And I also think like in the future, a really good curriculum or lesson writer agent will probably be like reinforcement fine-tuned on a lot of our internal data as well.”
Hsu: AI Gives Content and Engineering 100x Leverage While Still Requiring Review
“The way that we see it, really not just for our content team members, but also I think it's perfectly applicable to engineering is that it's leverage. It just allows you to do a hundred X in the same amount of time. We still need human review of the syllabus, …”
Hsu: Speak focuses on casual conversational language over textbook English
“That's one of our fundamental tenets, which is that we don't teach textbook English or textbook language. Like we try very hard to teach Gen Z slang. We don't go quite that far, but slaking. We try to teach Very casual conversational language that is actually …”
Hsu: Language learners face high psychological barrier practicing in front of humans
“It turns out there's like a really key psychological barrier there where people are just not willing to do this in front of a human, even if it's a human that is a teacher that you're paying, right?”
Hsu: Only a few TTS models handle multilingual code-switching properly
“It's actually like only a few models are able to speak two languages in the same sentence and then pronounce them properly.”