“So it went from seven and a half cents per million tokens to 10 cents. The sort of way that this was offset was we used to distinguish. I don't know if folks are familiar with this, but we used to distinguish based on input token volume. So it was like over a 128 K tokens would be twice as much. So it was like 15 cents. Now it's just like 10 cents straight up, whether it's one token or a million tokens”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Logan Kilpatrick
PredictionNot checkable as stated
Kilpatrick: Generative UI will be the killer use case for diffusion LLMs
“But I do think that's going to be the killer use case will be like this generative UI experience that doesn't exist today because the models just take too long to generate tokens.”
Logan KilpatrickJun 2, 2025▶ 6:51[AIEWF Preview] Gemini in 2025 and Realtime Voice AI
Disclosure
Kilpatrick: Google's strategy is building Gemini as one single unified model
“Like we're here to make one model and like that model is Gemini.
And like, I think you do need to just trust this point, like to make the capabilities work in some cases, like you do need to have these forks that like go off and make that capability and harde…”
Logan KilpatrickJun 2, 2025▶ 12:25[AIEWF Preview] Gemini in 2025 and Realtime Voice AI
AssertionNot checkable as stated
Kilpatrick: Gemini's SOTA video performance resulted from reasoning, not video engineering
“With reasoning is a great example of this where like multimodal with video understanding ended up like having this huge, like it's having this beautiful moment.
The model is like soda out of the box because of all the reasoning capabilities that were baked in…”
Logan KilpatrickJun 2, 2025▶ 13:19[AIEWF Preview] Gemini in 2025 and Realtime Voice AI
PredictionNot checkable as stated
Next billions of AI users will onboard via SMS and phone
“I don't think those people are going to come in through, you know, some front end website somewhere. Like those people are going to come in through audio from a telephone to texting to email. Like that's just the most obvious outcome.”
Logan KilpatrickFeb 28, 2025▶ 20:15Gemini 2.0 Flash and Flash Thinking: the new SOTA models for the agentic era
Opinion
Embeddings will not scale for cross-session memory in real-time AI systems
“I don't think that like embeddings are going to be able to scale to, I think they work well for some of this as like kind of the MVP version of the experience, but I think you're going to need a different experience and it's Probably something like really smar…”
Logan KilpatrickFeb 28, 2025▶ 23:41Gemini 2.0 Flash and Flash Thinking: the new SOTA models for the agentic era
Insight
Kilpatrick: Live AI APIs create severe vendor lock-in due to bespoke infra
“I think if you look at a lot of the live API infrastructure right now, like you really do need to commit that you're like gonna, you know, there's, it's not easily interoperable between different model providers. Like everyone's infrastructure is all bespoke a…”
Logan KilpatrickJun 2, 2025▶ 9:16[AIEWF Preview] Gemini in 2025 and Realtime Voice AI
Made with StarZero
Turn any episode into a week of clips.
This entire site, over 200 episodes transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.