Shrestha Basu Mallick

Product Lead, Google DeepMind · 1 appearance on the record.

computed by AI from the episodes · how this works → · full disclaimer →

operatorengineer@shresbm ↗LinkedIn ↗

Shrestha Basu Mallick leads product initiatives for Google's Gemini Developer APIs, including search grounding, Deep Research, and the real-time Gemini Live API. Before Google DeepMind, she worked at Salesforce Einstein, X, and McKinsey, and earned a Ph.D. in Applied Physics from Stanford University.

12statements → 10claims → 6claims resolved → 67%fully supported → 3.75/5average certainty → 1.42/5average debate potential →

4 supported 1 partly supported 1 contradicted 4 not checkable as stated how the 10 claims stand · each chip opens the sources

3 predictions · 7 assertions · 2 disclosures · every statement was checked. The predictions and assertions are the 10 claims: statements the public record can support or contradict. 6 are resolved, and 4 name no date, number or outcome precise enough to check. Everything else (opinions, insights, what ifs, disclosures) can never be settled by the record, so it carries no assessment.

The record, in short

What the tape says about how Shrestha argues and how the claims held up. Everything they said, and everything said about them, is in the tabs below.

Their most notable supported claim

Assertion Supported
Mallick: Google was first to market with live video API input
“So we were actually the first to market with also video input”
Shrestha Basu Mallick Jun 2, 2025 ▶ 7:53 [AIEWF Preview] Gemini in 2025 and Realtime Voice AI

Their most notable contradicted claim

Assertion Contradicted
Mallick: Gemini recognizes distinct voices as an unsupported emergent behavior
“This is not officially supported yet. The model just does it.”
Shrestha Basu Mallick Jun 2, 2025 ▶ 21:31 [AIEWF Preview] Gemini in 2025 and Realtime Voice AI

How they sound: not measured why? →

We measure speaking style by listening to the audio itself, and a fair number needs at least 2,000 words from one person on tape we have measured. There is too little of Shrestha Basu Mallick on measured tape to publish a rate. This says nothing about how they speak.

Everything Shrestha Basu Mallick said on Latent Space that made the record, most notable first. Filter by type, assessment or year in the ledger →

Prediction Not checkable as stated
Basu Mallick: Most voice use cases will transition to audio-to-audio models
“I do think perhaps eventually For most use cases, as these audio to audio architecture models get better, a lot of use cases will probably transition to that.”
Shrestha Basu Mallick Jun 2, 2025 ▶ 11:29 [AIEWF Preview] Gemini in 2025 and Realtime Voice AI
Assertion Contradicted
Mallick: Gemini recognizes distinct voices as an unsupported emergent behavior
“This is not officially supported yet. The model just does it.”
Shrestha Basu Mallick Jun 2, 2025 ▶ 21:31 [AIEWF Preview] Gemini in 2025 and Realtime Voice AI
Assertion Supported
Mallick: Google was first to market with live video API input
“So we were actually the first to market with also video input”
Shrestha Basu Mallick Jun 2, 2025 ▶ 7:53 [AIEWF Preview] Gemini in 2025 and Realtime Voice AI
Assertion Supported
Basu Mallick: Google was first to introduce tool chaining for live APIs
“Again, we were very proud because we introduced tool chaining first, so you could change search and code execution to all kinds of analysis.”
Shrestha Basu Mallick Jun 2, 2025 ▶ 8:33 [AIEWF Preview] Gemini in 2025 and Realtime Voice AI
Prediction Not checkable as stated
Mallick: Google's URL Context tool will enable developer research agents
“And I think that this will unlock new use cases. Like if people want to build their own version of a research agent, which is something developers ask us for a lot.”
Shrestha Basu Mallick Jun 2, 2025 ▶ 3:43 [AIEWF Preview] Gemini in 2025 and Realtime Voice AI
Assertion Not checkable as stated
Basu Mallick: Transcription was one of Gemini Live API's biggest early use cases
“Transcription actually, even before we released native audio, now of course you get text and audio interleaved in the output, but transcription used to be one of the biggest use cases we had on the live API.”
Shrestha Basu Mallick Jun 2, 2025 ▶ 7:25 [AIEWF Preview] Gemini in 2025 and Realtime Voice AI
Assertion Partly supported
Basu Mallick: Early Gemini Live sessions were capped at 20m audio, 5m video
“Like when we started, you could do like 15 to 20 minutes of audio, I'm sorry, and about five minutes of video.”
Shrestha Basu Mallick Jun 2, 2025 ▶ 8:05 [AIEWF Preview] Gemini in 2025 and Realtime Voice AI
Assertion Supported
Mallick: Gemini Live API allows custom VAD tuning and third-party integration
“Now developers can actually tune the sensitivity on our voice activity detection model as well as, you know, how much of the prefix, like how much of a time duration at the beginning, at the start or stop of saying things. And we also have a mode where now you…”
Shrestha Basu Mallick Jun 2, 2025 ▶ 17:30 [AIEWF Preview] Gemini in 2025 and Realtime Voice AI
Disclosure
Mallick: Google released experimental proactive audio for Gemini native audio
“One of the features that we've pushed out A little more experimental, but would love for people to test it is what we're calling proactive audio, and it's available only in the native audio, in the audio to audio architecture right now. And what this feature d…”
Shrestha Basu Mallick Jun 2, 2025 ▶ 20:19 [AIEWF Preview] Gemini in 2025 and Realtime Voice AI
Disclosure
Mallick: Google launched async function calling for cascaded voice models
“One thing that we launched on the cascaded architecture that we hope to eventually bring to the native audio as well is asynchronous function calling.”
Shrestha Basu Mallick Jun 2, 2025 ▶ 22:07 [AIEWF Preview] Gemini in 2025 and Realtime Voice AI
Assertion Supported
Mallick: Gemini officially supports 24 languages but responds in Klingon
“We officially support 24 languages, but you can try talking to the model and cling on and it'll respond to you.”
Shrestha Basu Mallick Jun 2, 2025 ▶ 23:28 [AIEWF Preview] Gemini in 2025 and Realtime Voice AI
Prediction Not checkable as stated
Mallick: Gemini will expand global language support well before next I/O
“So I think we'll get there way before next IO, but I just think more and more capabilities into the main model.”
Shrestha Basu Mallick Jun 2, 2025 ▶ 23:37 [AIEWF Preview] Gemini in 2025 and Realtime Voice AI

Appearances (1)

EpisodeDateSpeaking time
[AIEWF Preview] Gemini in 2025 and Realtime Voice AI Jun 2, 2025 8m
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.