context window

also referred to as: context windows

14 statements across 12 episodes · 5 bullish · 4 bearish · 12 people on the record · first statement Nov 3, 2023 by Michael Royzen · across every show →

Everything said about context window, oldest first

Nov 3, 2023 positive
Insight
Royzen: Large context windows outperform RAG chunking for code
“Like, I think it's generally been shown that if you have the space to just put The raw files inside of a big context window. That is still better than chunking and retrieval. It just is.”
Michael Royzen Nov 3, 2023 ▶ 39:24 Beating GPT-4 with Open Source Models - with Michael Royzen of Phind
Apr 11, 2024 negative
Insight
Stuhlmüller: Pure long-context LLMs are significantly harder to debug than RAG
“In one sense, I think you're right that the throw everything into the context window thing is easier to maintain because you just can swap out a model. In another sense, it's, if things go wrong, it's harder to debug, where, like, if you know, here's the proce…”
Andreas Stuhlmüller Apr 11, 2024 ▶ 51:48 Supervise the Process of AI Research — with Jungwon Byun and Andreas Stuhlmüller of Elicit
May 31, 2024 negative
Insight
Huang: Placing answers at the end of long contexts breaks attention
“You could create, like, a long context dataset where, like, every single time the last 200 tokens can answer the entire question, and that's never gonna make the model attend to anything.”
Mark Huang May 31, 2024 ▶ 30:34 How to train a Million Context LLM — with Mark Huang of Gradient.ai
May 31, 2024 positive
Insight
Huang: LLM session state management will require huge context windows
“Making the model track state and have state management over time is really, really hard. And it's an incredibly hard evaluation that will probably only really work when you have a huge context.”
Mark Huang May 31, 2024 ▶ 51:37 How to train a Million Context LLM — with Mark Huang of Gradient.ai
Aug 22, 2024
Insight
Autonomous coding agents require grounding all generated code in context windows
“Fundamentally to build a product like this, you need to get as much information in front of the model as possible and make sure that everything ever writes in output can be Traced back to something in the context window, so it's not hallucinating it.”
Alistair Pullen Aug 22, 2024 ▶ 11:32 Is finetuning GPT4o worth it?
Oct 4, 2024 bullish
Prediction Held up
Altman: 10-million-token fast context windows are coming within months
“Even getting to the, like, Ten million tokens of very fast and accurate context, which I expect to measure in, like, months, something like that.”
Sam Altman Oct 4, 2024 ▶ 2:05:30 Building AGI in Real Time (OpenAI Dev Day 2024)
Feb 18, 2025 neutral
Disclosure
Sridhar: Gemini Deep Research falls back to RAG beyond context limits
“We also have we have retrieval mechanisms, if required. So we natively try to use the context as much as it's available beyond which you know, we have, like, a rag setup to figure out”
Mukund Sridhar Feb 18, 2025 ▶ 19:03 Why is everyone cloning Deep Research?
Feb 18, 2025 positive
Insight
Sehgal: Keep recent research tasks in context, relegating older ones to RAG
“Just to add to that, I think like, just like a simple rule of thumb that we use is like, if it's the most recent set of research tasks where the user is likely to ask lots of follow-up questions, that should be in context. But like, as stuff gets 10 tasks ago,…”
Arush Sehgal Feb 18, 2025 ▶ 21:11 Why is everyone cloning Deep Research?
May 29, 2025 neutral
Insight
Massive context windows will not eliminate the need for RAG retrieval
“Like if you do have billion token context window model, you throw it all in there. It's still gonna be more expensive. The reason why retrieval is so important for us is because even if there is a model that's going to have these larger context windows, and ce…”
Matan Grinberg May 29, 2025 ▶ 35:43 The AI Coding Factory
Aug 5, 2025
Assertion Not checkable as stated
Dax Reed notes developers avoid context limits via frequent session restarts.
“To be honest, most people haven't complained about this. Like it comes up occasionally. But people tend to start new sessions pretty frequently. They tend to not have super long running things that often. So it is a problem that needs to be solved, but it's on…”
Dax Reed Aug 5, 2025 ▶ 16:44 ⚡️OpenCode: Claude Code but Open Source, with Any Model, and frontier TUI - with Dax Reed (@thdxr)
Aug 19, 2025 negative
Insight
Huber: LLM Performance and Reasoning Degrade as Token Counts Increase
“The performance of LLMs is not invariant to how many tokens you use. As you use more and more tokens, the model can pay attention to less, and then also can reason sort of less effectively.”
Jeff Huber Aug 19, 2025 ▶ 14:05 Long Live Context Engineering - with Jeff Huber of Chroma
Mar 5, 2026 bullish
Prediction Not checkable as stated
Huber: Model self-pruning of context windows will become standard
“And so, like, I think pruning is also going to be, like, really, it's already becoming a thing, right? But, like, letting models, like, self-prune their context windows.”
Jeff Huber Mar 5, 2026 ▶ 30:19 Why Every Agent Needs a Box — Aaron Levie, Box
Mar 6, 2026 neutral
Insight
Horthy: Beginners should compact LLM context at 40% of window capacity
“If you don't know what you're doing, and you don't really know what the AI model is capable of, and you don't have a lot of experience, like, you know, training wheels is like, when you get to 40%, start thinking about wrapping it up, or like doing a, like, yo…”
Dex Horthy Mar 6, 2026 ▶ 16:16 Why Your AI Agents Don’t Work with Dex Horthy of HumanLayer | In-Context Cooking
Jul 13, 2026 bearish
Prediction Not checkable as stated
Biderman: Model accuracy will still degrade at 10M context window scale
“But two is like, for the agentic tasks of 18 months from now, inside those major repositories of knowledge, and asking the models more and more things in underspecified ways, I suspect that the accuracy of the models would go down. The phenomenon of context fr…”
Dan Biderman Jul 13, 2026 ▶ 17:31 The AI Memory Problem: Why Long Context Isn’t Enough — Dan Biderman, Engram Co-founder & CEO
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.