“So I would say, and maybe I'm biased because one of my best friends is, is on the original attention paper, but that was the real breakthrough. It's just like figuring out that you have this attention mechanism that actually allows you to yeah, to do a much better job at, Sort of representation learning, and as a result, kind of generating correct sequences autoregressively.”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Douwe Kiela
Opinion
Kiela: Core language model development is almost solved and plateauing
“It's not even really about language models anymore. That has almost been solved, right? That's kind of why you see things plateauing off a little bit as well.”
Douwe KielaMar 6, 2025▶ 11:20Top AI Researcher on GPT 4.5, DeepSeek and Agentic RAG | Douwe Kiela, CEO, Contextual AI
Opinion
Kiela: Long-context LLMs are inherently incredibly wasteful
“Long context models are inherently incredibly wasteful. You're paying for all this compute, and that's maybe why some of the companies that are trying to really sell long context model, long context window models, they will make more money from that, right?”
Douwe KielaMar 6, 2025▶ 29:23Top AI Researcher on GPT 4.5, DeepSeek and Agentic RAG | Douwe Kiela, CEO, Contextual AI
Insight
Kiela: DeepSeek proved frontier AI models can rely on synthetic data
“We have kind of an existence proof now that it's actually not that hard to do this and so you don't need to invest all that much in, in data, and you can use synthetic data and get a pretty good model out of that”
Douwe KielaMar 6, 2025▶ 3:12Top AI Researcher on GPT 4.5, DeepSeek and Agentic RAG | Douwe Kiela, CEO, Contextual AI
AssertionNot checkable as stated
Kiela: DeepSeek's total development cost was at least 100x its $6M training
“So I would guess that they spent at least a hundred X The amount of that, that single training run, right?”
Douwe KielaMar 6, 2025▶ 9:58Top AI Researcher on GPT 4.5, DeepSeek and Agentic RAG | Douwe Kiela, CEO, Contextual AI
Insight
Kiela: Fine-tuning cannot inject new knowledge into AI models
“One common misconception about fine tuning is a lot of people think that you can inject new knowledge into a model using fine tuning. And that is not true.”
Douwe KielaMar 6, 2025▶ 28:19Top AI Researcher on GPT 4.5, DeepSeek and Agentic RAG | Douwe Kiela, CEO, Contextual AI
Insight
Kiela: Advanced RAG systems break down when scaling to a million PDFs
“You can build a very awesome demo on a couple of PDFs and things will probably work. But then you have to scale it up to a million PDFs, and then everything breaks down. And the reason for that is that a lot of these kind of advanced RAG systems still actually…”
Douwe KielaMar 6, 2025▶ 35:14Top AI Researcher on GPT 4.5, DeepSeek and Agentic RAG | Douwe Kiela, CEO, Contextual AI
Made with StarZero
Turn any episode into a week of clips.
This entire site, over 400 conversations transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.