Aug 19, 2025 · 57m · latent-space

Long Live Context Engineering - with Jeff Huber of Chroma

Jeff Huber · 36m spoken Shawn Wang · 11m spoken Alessio Fanelli · 3m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of Latent Space, Chroma co-founder and CEO Jeff Huber explores the evolution of AI search infrastructure, introducing 'context engineering' to resolve context rot while advocating for high-craftsmanship engineering and empirical benchmarking.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 28.2% of the talking time here. How this is scored →

The hosts as informed peer 4.6 Guest teaching 4.2 Guest disagreement 2.2 The hosts pushing back 1.8
05100:0015:0030:0045:002:56–8:08 · The hosts as informed peer 4/10 Defining Modern Search Infrastructure for AI Swyx asks Jeff to clarify terminology between information retrieval and search, prompting Jeff to systematically delineate modern distributed search infrastructure and the four distinct ways AI changes search requirements.8:08–12:14 · The hosts as informed peer 4/10 Chroma Cloud Architecture and Developer Experience Swyx brings up Chroma's download and star metrics and explores Chroma Cloud's serverless architecture. When Swyx speculates that SQLite wrappers are pip installable, Jeff gently corrects the technical history.12:16–18:52 · The hosts as informed peer 5/10 Defining Context Engineering and the Threat of Context Rot Jeff rejects buzzwords like RAG and ambiguous agent definitions while laying out context engineering and context rot findings. Swyx engages with technical theories about reasoning models and context utilization.18:52–23:36 · The hosts as informed peer 6/10 Research Incentives and Frontier Model Dynamics Swyx offers direct pushback against Jeff's claim that frontier labs solely optimize for consumers, pointing to OpenAI's ChatGPT memory features. Alessio presses Jeff on whether this problem falls under the Bitter Lesson.23:37–27:02 · The hosts as informed peer 4/10 Emerging Paradigms in Context Engineering and LLM Re-ranking Jeff outlines the transition from first-stage retrieval to LLM-based re-ranking, correcting Swyx's assumption that practitioners only use dedicated lightweight re-rankers rather than prompting general LLMs.27:02–30:00 · The hosts as informed peer 5/10 Code Retrieval, Indexing Tradeoffs, and Index Forking Swyx brings up how coding tools like Claude Code handle retrieval without traditional indexing. Jeff deconstructs indexing as fundamentally trading write-time cost for query-time speed and introduces Chroma index forking.30:02–34:04 · The hosts as informed peer 4/10 Evaluating Search Strategies and Data Ingestion Pipelines Alessio asks about developer vs agent experiences in code representation. Jeff educates on chunk rewriting at ingestion and uses a Google Drive spreadsheet analogy to explain when lexical vs embedding search works.34:05–37:47 · The hosts as informed peer 6/10 The Architectural Future of Retrieval and Latent Space Swyx synthesizes an architectural overview of how the industry decoupled transformer encoders and decoders across vector databases. Jeff reacts with vision for continual retrieval and staying inside latent space while jokingly shutting down the phrase 'agentic RAG'.37:48–42:18 · The hosts as informed peer 5/10 Demystifying AI Memory and Compaction Alessio and Swyx explore memory taxonomy and sleep/garbage collection cycles. Jeff dismisses overcomplicated memory taxonomy charts, re-grounding AI memory in classical database compaction and continuous re-indexing.42:18–45:27 · The hosts as informed peer 5/10 Generative Benchmarking and the Power of Small Labeled Data Jeff explains Chroma's generative benchmarking paper to solve the missing query problem in golden datasets. Swyx agrees strongly, refining Jeff's slogan from 'look at your data' to 'label your data'.45:27–49:58 · The hosts as informed peer 4/10 Conviction, Craft, and Countering Tech Nihilism Swyx asks Jeff about his background with Standard Cyborg and how religious conviction informs his view of startup impact against Valley nihilism. Jeff critiques AGI hype as a modern secular religion.49:59–52:35 · The hosts as informed peer 3/10 Taste, Design Philosophy, and Brand Intentionality Alessio asks about Chroma's strong aesthetic and design culture. Jeff explains the founder's duty to act as a curator of taste to maintain company coherence across every touchpoint.2:56–8:08 · Guest teaching 4/10 Defining Modern Search Infrastructure for AI Swyx asks Jeff to clarify terminology between information retrieval and search, prompting Jeff to systematically delineate modern distributed search infrastructure and the four distinct ways AI changes search requirements.8:08–12:14 · Guest teaching 4/10 Chroma Cloud Architecture and Developer Experience Swyx brings up Chroma's download and star metrics and explores Chroma Cloud's serverless architecture. When Swyx speculates that SQLite wrappers are pip installable, Jeff gently corrects the technical history.12:16–18:52 · Guest teaching 5/10 Defining Context Engineering and the Threat of Context Rot Jeff rejects buzzwords like RAG and ambiguous agent definitions while laying out context engineering and context rot findings. Swyx engages with technical theories about reasoning models and context utilization.18:52–23:36 · Guest teaching 4/10 Research Incentives and Frontier Model Dynamics Swyx offers direct pushback against Jeff's claim that frontier labs solely optimize for consumers, pointing to OpenAI's ChatGPT memory features. Alessio presses Jeff on whether this problem falls under the Bitter Lesson.23:37–27:02 · Guest teaching 5/10 Emerging Paradigms in Context Engineering and LLM Re-ranking Jeff outlines the transition from first-stage retrieval to LLM-based re-ranking, correcting Swyx's assumption that practitioners only use dedicated lightweight re-rankers rather than prompting general LLMs.27:02–30:00 · Guest teaching 5/10 Code Retrieval, Indexing Tradeoffs, and Index Forking Swyx brings up how coding tools like Claude Code handle retrieval without traditional indexing. Jeff deconstructs indexing as fundamentally trading write-time cost for query-time speed and introduces Chroma index forking.30:02–34:04 · Guest teaching 5/10 Evaluating Search Strategies and Data Ingestion Pipelines Alessio asks about developer vs agent experiences in code representation. Jeff educates on chunk rewriting at ingestion and uses a Google Drive spreadsheet analogy to explain when lexical vs embedding search works.34:05–37:47 · Guest teaching 4/10 The Architectural Future of Retrieval and Latent Space Swyx synthesizes an architectural overview of how the industry decoupled transformer encoders and decoders across vector databases. Jeff reacts with vision for continual retrieval and staying inside latent space while jokingly shutting down the phrase 'agentic RAG'.37:48–42:18 · Guest teaching 5/10 Demystifying AI Memory and Compaction Alessio and Swyx explore memory taxonomy and sleep/garbage collection cycles. Jeff dismisses overcomplicated memory taxonomy charts, re-grounding AI memory in classical database compaction and continuous re-indexing.42:18–45:27 · Guest teaching 4/10 Generative Benchmarking and the Power of Small Labeled Data Jeff explains Chroma's generative benchmarking paper to solve the missing query problem in golden datasets. Swyx agrees strongly, refining Jeff's slogan from 'look at your data' to 'label your data'.45:27–49:58 · Guest teaching 2/10 Conviction, Craft, and Countering Tech Nihilism Swyx asks Jeff about his background with Standard Cyborg and how religious conviction informs his view of startup impact against Valley nihilism. Jeff critiques AGI hype as a modern secular religion.49:59–52:35 · Guest teaching 3/10 Taste, Design Philosophy, and Brand Intentionality Alessio asks about Chroma's strong aesthetic and design culture. Jeff explains the founder's duty to act as a curator of taste to maintain company coherence across every touchpoint.2:56–8:08 · Guest disagreement 2/10 Defining Modern Search Infrastructure for AI Swyx asks Jeff to clarify terminology between information retrieval and search, prompting Jeff to systematically delineate modern distributed search infrastructure and the four distinct ways AI changes search requirements.8:08–12:14 · Guest disagreement 2/10 Chroma Cloud Architecture and Developer Experience Swyx brings up Chroma's download and star metrics and explores Chroma Cloud's serverless architecture. When Swyx speculates that SQLite wrappers are pip installable, Jeff gently corrects the technical history.12:16–18:52 · Guest disagreement 3/10 Defining Context Engineering and the Threat of Context Rot Jeff rejects buzzwords like RAG and ambiguous agent definitions while laying out context engineering and context rot findings. Swyx engages with technical theories about reasoning models and context utilization.18:52–23:36 · Guest disagreement 3/10 Research Incentives and Frontier Model Dynamics Swyx offers direct pushback against Jeff's claim that frontier labs solely optimize for consumers, pointing to OpenAI's ChatGPT memory features. Alessio presses Jeff on whether this problem falls under the Bitter Lesson.23:37–27:02 · Guest disagreement 2/10 Emerging Paradigms in Context Engineering and LLM Re-ranking Jeff outlines the transition from first-stage retrieval to LLM-based re-ranking, correcting Swyx's assumption that practitioners only use dedicated lightweight re-rankers rather than prompting general LLMs.27:02–30:00 · Guest disagreement 2/10 Code Retrieval, Indexing Tradeoffs, and Index Forking Swyx brings up how coding tools like Claude Code handle retrieval without traditional indexing. Jeff deconstructs indexing as fundamentally trading write-time cost for query-time speed and introduces Chroma index forking.30:02–34:04 · Guest disagreement 2/10 Evaluating Search Strategies and Data Ingestion Pipelines Alessio asks about developer vs agent experiences in code representation. Jeff educates on chunk rewriting at ingestion and uses a Google Drive spreadsheet analogy to explain when lexical vs embedding search works.34:05–37:47 · Guest disagreement 3/10 The Architectural Future of Retrieval and Latent Space Swyx synthesizes an architectural overview of how the industry decoupled transformer encoders and decoders across vector databases. Jeff reacts with vision for continual retrieval and staying inside latent space while jokingly shutting down the phrase 'agentic RAG'.37:48–42:18 · Guest disagreement 3/10 Demystifying AI Memory and Compaction Alessio and Swyx explore memory taxonomy and sleep/garbage collection cycles. Jeff dismisses overcomplicated memory taxonomy charts, re-grounding AI memory in classical database compaction and continuous re-indexing.42:18–45:27 · Guest disagreement 1/10 Generative Benchmarking and the Power of Small Labeled Data Jeff explains Chroma's generative benchmarking paper to solve the missing query problem in golden datasets. Swyx agrees strongly, refining Jeff's slogan from 'look at your data' to 'label your data'.45:27–49:58 · Guest disagreement 2/10 Conviction, Craft, and Countering Tech Nihilism Swyx asks Jeff about his background with Standard Cyborg and how religious conviction informs his view of startup impact against Valley nihilism. Jeff critiques AGI hype as a modern secular religion.49:59–52:35 · Guest disagreement 1/10 Taste, Design Philosophy, and Brand Intentionality Alessio asks about Chroma's strong aesthetic and design culture. Jeff explains the founder's duty to act as a curator of taste to maintain company coherence across every touchpoint.2:56–8:08 · The hosts pushing back 1/10 Defining Modern Search Infrastructure for AI Swyx asks Jeff to clarify terminology between information retrieval and search, prompting Jeff to systematically delineate modern distributed search infrastructure and the four distinct ways AI changes search requirements.8:08–12:14 · The hosts pushing back 2/10 Chroma Cloud Architecture and Developer Experience Swyx brings up Chroma's download and star metrics and explores Chroma Cloud's serverless architecture. When Swyx speculates that SQLite wrappers are pip installable, Jeff gently corrects the technical history.12:16–18:52 · The hosts pushing back 2/10 Defining Context Engineering and the Threat of Context Rot Jeff rejects buzzwords like RAG and ambiguous agent definitions while laying out context engineering and context rot findings. Swyx engages with technical theories about reasoning models and context utilization.18:52–23:36 · The hosts pushing back 6/10 Research Incentives and Frontier Model Dynamics Swyx offers direct pushback against Jeff's claim that frontier labs solely optimize for consumers, pointing to OpenAI's ChatGPT memory features. Alessio presses Jeff on whether this problem falls under the Bitter Lesson.23:37–27:02 · The hosts pushing back 2/10 Emerging Paradigms in Context Engineering and LLM Re-ranking Jeff outlines the transition from first-stage retrieval to LLM-based re-ranking, correcting Swyx's assumption that practitioners only use dedicated lightweight re-rankers rather than prompting general LLMs.27:02–30:00 · The hosts pushing back 1/10 Code Retrieval, Indexing Tradeoffs, and Index Forking Swyx brings up how coding tools like Claude Code handle retrieval without traditional indexing. Jeff deconstructs indexing as fundamentally trading write-time cost for query-time speed and introduces Chroma index forking.30:02–34:04 · The hosts pushing back 1/10 Evaluating Search Strategies and Data Ingestion Pipelines Alessio asks about developer vs agent experiences in code representation. Jeff educates on chunk rewriting at ingestion and uses a Google Drive spreadsheet analogy to explain when lexical vs embedding search works.34:05–37:47 · The hosts pushing back 3/10 The Architectural Future of Retrieval and Latent Space Swyx synthesizes an architectural overview of how the industry decoupled transformer encoders and decoders across vector databases. Jeff reacts with vision for continual retrieval and staying inside latent space while jokingly shutting down the phrase 'agentic RAG'.37:48–42:18 · The hosts pushing back 1/10 Demystifying AI Memory and Compaction Alessio and Swyx explore memory taxonomy and sleep/garbage collection cycles. Jeff dismisses overcomplicated memory taxonomy charts, re-grounding AI memory in classical database compaction and continuous re-indexing.42:18–45:27 · The hosts pushing back 1/10 Generative Benchmarking and the Power of Small Labeled Data Jeff explains Chroma's generative benchmarking paper to solve the missing query problem in golden datasets. Swyx agrees strongly, refining Jeff's slogan from 'look at your data' to 'label your data'.45:27–49:58 · The hosts pushing back 1/10 Conviction, Craft, and Countering Tech Nihilism Swyx asks Jeff about his background with Standard Cyborg and how religious conviction informs his view of startup impact against Valley nihilism. Jeff critiques AGI hype as a modern secular religion.49:59–52:35 · The hosts pushing back 0/10 Taste, Design Philosophy, and Brand Intentionality Alessio asks about Chroma's strong aesthetic and design culture. Jeff explains the founder's duty to act as a curator of taste to maintain company coherence across every touchpoint.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 26.2% · guest 73.8%0:00 · the hosts 26.2% · guest 73.8%3:00 · the hosts 21.1% · guest 78.9%3:00 · the hosts 21.1% · guest 78.9%6:00 · the hosts 35.5% · guest 64.5%6:00 · the hosts 35.5% · guest 64.5%9:00 · the hosts 14.5% · guest 85.5%9:00 · the hosts 14.5% · guest 85.5%12:00 · the hosts 24.6% · guest 75.4%12:00 · the hosts 24.6% · guest 75.4%15:00 · the hosts 27% · guest 73%15:00 · the hosts 27% · guest 73%18:00 · the hosts 31.9% · guest 68.1%18:00 · the hosts 31.9% · guest 68.1%21:00 · the hosts 43.2% · guest 56.8%21:00 · the hosts 43.2% · guest 56.8%24:00 · the hosts 1.9% · guest 98.1%24:00 · the hosts 1.9% · guest 98.1%27:00 · the hosts 26.7% · guest 73.3%27:00 · the hosts 26.7% · guest 73.3%30:00 · the hosts 9.5% · guest 90.5%30:00 · the hosts 9.5% · guest 90.5%33:00 · the hosts 48.6% · guest 51.4%33:00 · the hosts 48.6% · guest 51.4%36:00 · the hosts 25.4% · guest 74.6%36:00 · the hosts 25.4% · guest 74.6%39:00 · the hosts 31.6% · guest 68.4%39:00 · the hosts 31.6% · guest 68.4%42:00 · the hosts 29.3% · guest 70.7%42:00 · the hosts 29.3% · guest 70.7%45:00 · the hosts 33.2% · guest 66.8%45:00 · the hosts 33.2% · guest 66.8%48:00 · the hosts 33.7% · guest 66.3%48:00 · the hosts 33.7% · guest 66.3%51:00 · the hosts 29% · guest 71%51:00 · the hosts 29% · guest 71%54:00 · the hosts 44.8% · guest 55.2%54:00 · the hosts 44.8% · guest 55.2%57:00 · the hosts 0% · guest 0%57:00 · the hosts 0% · guest 0%
Sharpest disagreement ▶ 13:08 Dismissal of RAG terminology

Jeff bluntly rejects the popular industry buzzword RAG as dumb, confusing, and reductive, arguing that it obscures genuine context engineering.

Hardest push from the hosts ▶ 21:02 Host pushes back on lab consumer focus

Swyx directly challenges Jeff's claim that frontier labs ignore developers, arguing that OpenAI's heavy investment in consumer ChatGPT memory proves context engineering matters universally.

Biggest teaching moment ▶ 30:15 Masterclass on lexical vs semantic search

Jeff provides an intuitive explanation using the CapTable file search example to clarify exactly when full-text lexical search outperforms embedding search.

The host holds their own ▶ 34:20 Synthesizing the decoupled transformer architecture

Swyx presents a high-level theoretical framework showing how the AI ecosystem decoupled original encoder-decoder transformers into encoder vector databases and decoder generation LLMs.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Defining Modern Search Infrastructure for AI 4421 Swyx asks Jeff to clarify terminology between information retrieval and search, prompting Jeff to systematically delineate modern distributed search infrastructure and the four distinct ways AI changes search requirements.
Chroma Cloud Architecture and Developer Experience 4422 Swyx brings up Chroma's download and star metrics and explores Chroma Cloud's serverless architecture. When Swyx speculates that SQLite wrappers are pip installable, Jeff gently corrects the technical history.
Defining Context Engineering and the Threat of Context Rot 5532 Jeff rejects buzzwords like RAG and ambiguous agent definitions while laying out context engineering and context rot findings. Swyx engages with technical theories about reasoning models and context utilization.
Research Incentives and Frontier Model Dynamics 6436 Swyx offers direct pushback against Jeff's claim that frontier labs solely optimize for consumers, pointing to OpenAI's ChatGPT memory features. Alessio presses Jeff on whether this problem falls under the Bitter Lesson.
Emerging Paradigms in Context Engineering and LLM Re-ranking 4522 Jeff outlines the transition from first-stage retrieval to LLM-based re-ranking, correcting Swyx's assumption that practitioners only use dedicated lightweight re-rankers rather than prompting general LLMs.
Code Retrieval, Indexing Tradeoffs, and Index Forking 5521 Swyx brings up how coding tools like Claude Code handle retrieval without traditional indexing. Jeff deconstructs indexing as fundamentally trading write-time cost for query-time speed and introduces Chroma index forking.
Evaluating Search Strategies and Data Ingestion Pipelines 4521 Alessio asks about developer vs agent experiences in code representation. Jeff educates on chunk rewriting at ingestion and uses a Google Drive spreadsheet analogy to explain when lexical vs embedding search works.
The Architectural Future of Retrieval and Latent Space 6433 Swyx synthesizes an architectural overview of how the industry decoupled transformer encoders and decoders across vector databases. Jeff reacts with vision for continual retrieval and staying inside latent space while jokingly shutting down the phrase 'agentic RAG'.
Demystifying AI Memory and Compaction 5531 Alessio and Swyx explore memory taxonomy and sleep/garbage collection cycles. Jeff dismisses overcomplicated memory taxonomy charts, re-grounding AI memory in classical database compaction and continuous re-indexing.
Generative Benchmarking and the Power of Small Labeled Data 5411 Jeff explains Chroma's generative benchmarking paper to solve the missing query problem in golden datasets. Swyx agrees strongly, refining Jeff's slogan from 'look at your data' to 'label your data'.
Conviction, Craft, and Countering Tech Nihilism 4221 Swyx asks Jeff about his background with Standard Cyborg and how religious conviction informs his view of startup impact against Valley nihilism. Jeff critiques AGI hype as a modern secular religion.
Taste, Design Philosophy, and Brand Intentionality 3310 Alessio asks about Chroma's strong aesthetic and design culture. Jeff explains the founder's duty to act as a curator of taste to maintain company coherence across every touchpoint.

Statements from this episode (21)

Disclosure
Huber: Chroma is written in Rust and uses object storage
“Chrome is written in Rust. It's fully multi-tenant. We have, we use object storage as a key Assistance tier and, like, data layer for Chroma distributed in Chroma Cloud as well.”
Jeff Huber Aug 19, 2025 ▶ 3:28
Insight
Huber: Gradient-descent user feedback creates lowest-common-denominator products
“My critique of that would be that if you follow that methodology, you will probably end up building a dating app for middle schoolers, because that just seems to be like the lowest base take of what humans want to some degree.”
Jeff Huber Aug 19, 2025 ▶ 5:16
Assertion Not checkable as stated
Huber: Chroma is currently serving hundreds of thousands of developers
“And obviously I'm incredibly proud that it exists today and that it's like serving hundreds of thousands of developers and they love it, but it was hard to get there.”
Jeff Huber Aug 19, 2025 ▶ 6:32
Assertion Supported
Huber: Chroma is the most used project across LangChain and LlamaIndex
“For many years running, Chrome has been the number one used project broadly, but also within communities like LinkChain and Llama Index.”
Jeff Huber Aug 19, 2025 ▶ 8:43
Disclosure
Huber: Chroma Distributed control and data planes are open source Apache 2.0
“Chroma Distributed is also a part of the same monorepo. That's open source Apache two. And then the control and data plane are both fully open source Apache two.”
Jeff Huber Aug 19, 2025 ▶ 11:27
Insight
Huber: LLM Performance and Reasoning Degrade as Token Counts Increase
“The performance of LLMs is not invariant to how many tokens you use. As you use more and more tokens, the model can pay attention to less, and then also can reason sort of less effectively.”
Jeff Huber Aug 19, 2025 ▶ 14:05
Opinion
Huber: Successful AI Startups Fundamentally Excel at Context Engineering
“This is what, frankly, most AI startups, any AI stuff that you know of, that you think of today that's doing very well, like what are they fundamentally good at? What is the one thing that they're good at? It is context engineering.”
Jeff Huber Aug 19, 2025 ▶ 14:36
Opinion
Swix: Reasoning Models Have Better Context Utilization Than Standard LLMs
“I have a theory also that reasoning models have better context utilization because they can loop back. Normal auto-aggressive models, they just kind of go left to right, but reasoning models, in theory, they can loop back and look for things that they needed c…”
Shawn Wang Aug 19, 2025 ▶ 18:28
Opinion
Huber: Consumer focus leaves LLM labs unmotivated to help developers
“Increasingly is the market to be a good LLM provider, the main market seems to be consumer. You're just not that motivated to, like, help developers.”
Jeff Huber Aug 19, 2025 ▶ 20:20
Opinion
Huber: Flawless 60k-token reasoning is more valuable than 5M-token context
“I would rather have a model that has a 60,000 context, token context window, that is able to perfectly pay attention to, and perfectly reason over those 60,000 tokens, than a model that's like five million tokens. Like, just as a developer, the former is like …”
Jeff Huber Aug 19, 2025 ▶ 22:17
Assertion Not checkable as stated
Huber: Context caching improves cost and speed but ignores context rot
“And yeah, they're using context caching and that certainly helps, but like their cost and speed, but like isn't helping the context raw problem at all.”
Jeff Huber Aug 19, 2025 ▶ 23:48
Prediction Not checkable as stated
Huber: LLMs will largely replace purpose-built re-rankers
“I think that, like, this is going to be the dominant paradigm. I actually think that, like, probably purpose-built re-rankers will go away, and the same way that, like, purpose-built, they'll still exist, right? Like, if you're at extreme scale, extreme cost, …”
Jeff Huber Aug 19, 2025 ▶ 25:50
Opinion
Huber: Regex handles 90% of code queries; embeddings add marginal improvement
“My guess is that, like, for code today, it's something like, 90% of queries or 85% of queries can be satisfactorily run with regex. Regex is obviously, like, the dominant pattern used by Google code search, GitHub code search, but you maybe can get, like, 15% …”
Jeff Huber Aug 19, 2025 ▶ 31:09
Insight
Huber: Pre-Baking Metadata and Chunk Rewriting at Ingestion Simplifies Retrieval
“As much structured information as you can put into your write or your ingestion pipeline, you should. So all of the metadata you can extract, do it at ingestion. All of the chunk rewriting you can do, do it at ingestion. If you really invest in, like, trying t…”
Jeff Huber Aug 19, 2025 ▶ 32:48
Prediction Not checkable as stated
Huber: Future retrieval systems will operate entirely within latent space
“I think, like, there's a few things that I think might be true about retrieval systems in the future. So, like, number one, they just stay in latent space, they don't go back to natural language.”
Jeff Huber Aug 19, 2025 ▶ 35:39
Insight
Huber: Frozen architectures and static corpora hurt early retrieval-model adoption
“A lot of those have the problem where, like, either the retriever or the language model has to be frozen, and then, like, the corpus can't change, which most developers don't want to, like, deal with the developer experience around.”
Jeff Huber Aug 19, 2025 ▶ 36:35
Prediction Not checkable as stated
Huber: Offline compute driving continuous AI self-improvement is a sure bet
“Like, the idea that there's going to be, like, a lot of offline compute and inference under the hood that helps make AI systems continuously self-improve is a sure bet.”
Jeff Huber Aug 19, 2025 ▶ 42:09
Insight
Huber: Synthetic QA pair generation is underrated for retrieval benchmarking
“So I think generating QA pairs is really important for benchmarking your retrieval system, golden dataset. Frankly, it's also the same dataset that you would use to fine tune in many cases. And so, yeah, there's definitely something like very underrated there.”
Jeff Huber Aug 19, 2025 ▶ 43:45
Insight
Huber: A couple hundred high-quality labeled examples offer massive ML returns
“Having worked in applied machine learning developer tools now for 10 years, like the returns to a very high quality small label data set are so high. Everybody thinks you have to have like a million examples or whatever. No, actually just like a couple hundred…”
Jeff Huber Aug 19, 2025 ▶ 44:37
Opinion
Huber: Silicon Valley treats AGI as a secular religion
“I think AGI is also a religion. It has a problem of evil. We don't have enough intelligence. It has a solution, a deus ex machina. It has the second coming of Christ that AGI, the singularity is going to come. It's going to save humanity because we will now ha…”
Jeff Huber Aug 19, 2025 ▶ 48:45
Opinion
Huber: No AI coding tools are particularly good at Rust
“So far we've still not found that really any AI coding tools are particularly good at rust though.”
Jeff Huber Aug 19, 2025 ▶ 55:54
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.