The Ledger, every show
Every statement that passed quotation and attribution checks, across all 44 shows. Pick shows below, then mix any filter with any other.
shows 




every show 44 of 44
Huber: Silicon Valley treats AGI as a secular religion
“I think AGI is also a religion. It has a problem of evil. We don't have enough intelligence. It has a solution, a deus ex machina. It has the second coming of Christ that AGI, the singularity is going to come. It's going to save humanity because we will now ha…”
Huber: Frontier models repeat mistakes if failed actions remain in context
“A few of the insights is, like, everyone, frontier model is not good at search. Humans have this natural explore-exploit trade-off, where we kind of understand, like, when to stop doing something. Also, humans are pretty good at, like, forgetting, actually, li…”
Huber: LLM Performance and Reasoning Degrade as Token Counts Increase
“The performance of LLMs is not invariant to how many tokens you use. As you use more and more tokens, the model can pay attention to less, and then also can reason sort of less effectively.”
Huber: LLMs will largely replace purpose-built re-rankers
“I think that, like, this is going to be the dominant paradigm. I actually think that, like, probably purpose-built re-rankers will go away, and the same way that, like, purpose-built, they'll still exist, right? Like, if you're at extreme scale, extreme cost, …”
Huber: Regex handles 90% of code queries; embeddings add marginal improvement
“My guess is that, like, for code today, it's something like, 90% of queries or 85% of queries can be satisfactorily run with regex. Regex is obviously, like, the dominant pattern used by Google code search, GitHub code search, but you maybe can get, like, 15% …”
Huber: Future retrieval systems will operate entirely within latent space
“I think, like, there's a few things that I think might be true about retrieval systems in the future. So, like, number one, they just stay in latent space, they don't go back to natural language.”
Huber: LLMs are like CPUs, not operating systems
“I don't think of an LLM as an operating system. I think an LLM is much more like a CPU, right? It's an information processing unit.”
Huber: Open-source models win B2B through developer focus, not beating GPT-5
“Focus on the developers. I think that's the beachhead. That's how you win the B to B market. If you win the B to B market with your open source models, Like, you get all of the sort of downstream effects that you want. You know, you don't need to beat you know…”
Huber: 10x compute increases are not producing 10x better AI models
“Diminishing, they're clearly diminishing marginal returns, right? We're sort of spending 10 X on compute. We're not getting 10 X or better models, at least evidently not yet.”
Huber: AI will probably drive GDP growth exceeding the Industrial Revolution
“It's a, you know, technology is probably as important as the invention of electricity. It will probably, you know, bring about a increase in GDP that is on the order of the industrial revolution or greater.”
Huber: In 10 years, the poorest could have better healthcare than today's billionaires
“Like it is very possible the poorest people on earth today, or, you know, in 10 years, we'll have access to better healthcare better legal representation you know, better financial services than, like, billionaires have today.”
Huber: Current SOTA LLMs Lack Reliability for Multi-Agent Workflows
“Now, of course, for those of you that have actually played with technology, I think it's questionable whether the current state of the art Language models, embedding models, et cetera, will give you the reliability you want from, ah, you know, agents working t…”
Huber: RBAC is dead and AI vendors are ignoring authorization
“Well, I'm just saying, I think, like, RBAC does seem to be dead, right? Like, if you want to say something is dead, probably RBAC is dead. And like the auth story to me seems like incredibly unsolved and unaddressed by like the existing state of like AI vendor…”
Huber: Frontier AI models are not actually good at agentic search
“We've like sort of stress tested like frontier models and their ability to search. And they are not actually that good at searching.”
Huber: Graph structures emerge dynamically in AI agents rather than schemas
“I think that the actual graph structure is emergent in the mind of the agent, ah, in the same way it is in the mind of the human. And that's a more powerful graph, because it actually evolved over time.”
Huber: Gradient-descent user feedback creates lowest-common-denominator products
“My critique of that would be that if you follow that methodology, you will probably end up building a dating app for middle schoolers, because that just seems to be like the lowest base take of what humans want to some degree.”
Huber: Successful AI Startups Fundamentally Excel at Context Engineering
“This is what, frankly, most AI startups, any AI stuff that you know of, that you think of today that's doing very well, like what are they fundamentally good at? What is the one thing that they're good at? It is context engineering.”
Huber: Consumer focus leaves LLM labs unmotivated to help developers
“Increasingly is the market to be a good LLM provider, the main market seems to be consumer. You're just not that motivated to, like, help developers.”
Huber: Flawless 60k-token reasoning is more valuable than 5M-token context
“I would rather have a model that has a 60,000 context, token context window, that is able to perfectly pay attention to, and perfectly reason over those 60,000 tokens, than a model that's like five million tokens. Like, just as a developer, the former is like …”
Huber: Offline compute driving continuous AI self-improvement is a sure bet
“Like, the idea that there's going to be, like, a lot of offline compute and inference under the hood that helps make AI systems continuously self-improve is a sure bet.”
Huber: No AI coding tools are particularly good at Rust
“So far we've still not found that really any AI coding tools are particularly good at rust though.”
Huber: Language models require a multi-tiered memory hierarchy like traditional computers
“In the same way that we have a memory hierarchy in classic computers, right, we have the CPU, RAM, disk, and network we are also going to have a similar memory hierarchy in language models. And again, it already exists today. We have the actual sort of transfo…”
Huber: Needle-in-a-haystack tests do not prove real-world long context reliability
“Even these, like, needle-in-a-haystack tests, like, are not actually that representative of, like, real-world utility and reliability of long context windows.”
Huber: Never bet against Zuckerberg and Meta's distribution power
“Distribution is incredibly important as long as, you know, sort of the incumbents can wake up and can catch up. You know, I would not bet against Zuck And a hundred billion dollars of profit per year.”
Huber: Most businesses prefer open-source models over closed-source AI
“Most businesses don't love using closed source models. They want to use open source models. For all kinds of reasons, you know, privacy, security, continuity, cost”
Huber: Fine-tuning model weights fails enterprise AI due to lack of deterministic control
“Updating the weights of the model is not a very good idea because you cannot really deterministically control that. You can fine tune, but what you're going to get the other end, you know, again, you don't really control.”
Huber: Current AI models have an immense capability overhang
“We think the capability overhang we have in the models that we already have today, and we will have absolutely in six months is immense.”
Huber: Chroma is committed to remaining fully open source
“Chroma will always, we are committed to building the ubiquitous open source standard.”
Huber: Vector databases must support both transactional and analytical workloads
“And we think that both certainly transactional has to be the case because it is a online database. It's gonna sit in the loop of applications. Again, you've already seen demos of this happening tonight. But also to make this technology useful for developers, y…”
Huber: Asking whether vector databases replace classic databases is dumb
“People, there's this, like, you know, big question about, oh, are vector databases gonna replace classic databases? Are these competitive in some way? And I think it's just kind of a dumb question.”
Huber: Model self-pruning of context windows will become standard
“And so, like, I think pruning is also going to be, like, really, it's already becoming a thing, right? But, like, letting models, like, self-prune their context windows.”
Huber: Most companies operate as apprenticeships with unwritten tacit knowledge
“Most companies are practically apprenticeships. Like every new employee who joins the team, like you spend one to three months, like wrapping them up. All that tested knowledge is not written down.”
Huber: Chroma is the most used project across LangChain and LlamaIndex
“For many years running, Chrome has been the number one used project broadly, but also within communities like LinkChain and Llama Index.”
Huber: Context caching improves cost and speed but ignores context rot
“And yeah, they're using context caching and that certainly helps, but like their cost and speed, but like isn't helping the context raw problem at all.”
Huber: Pre-Baking Metadata and Chunk Rewriting at Ingestion Simplifies Retrieval
“As much structured information as you can put into your write or your ingestion pipeline, you should. So all of the metadata you can extract, do it at ingestion. All of the chunk rewriting you can do, do it at ingestion. If you really invest in, like, trying t…”
Huber: Frozen architectures and static corpora hurt early retrieval-model adoption
“A lot of those have the problem where, like, either the retriever or the language model has to be frozen, and then, like, the corpus can't change, which most developers don't want to, like, deal with the developer experience around.”
Huber: Synthetic QA pair generation is underrated for retrieval benchmarking
“So I think generating QA pairs is really important for benchmarking your retrieval system, golden dataset. Frankly, it's also the same dataset that you would use to fine tune in many cases. And so, yeah, there's definitely something like very underrated there.”
Huber: A couple hundred high-quality labeled examples offer massive ML returns
“Having worked in applied machine learning developer tools now for 10 years, like the returns to a very high quality small label data set are so high. Everybody thinks you have to have like a million examples or whatever. No, actually just like a couple hundred…”
Huber: Silicon Valley tends to be extremely intellectually shallow
“Silicon Valley has a tendency to be sort of extremely intellectually shallow. This is both a strength and a weakness of the Valley, to be clear.”
Huber: Model releases often top leaderboards but lack practical developer hooks
“You know, you see a lot of like model drops that come out, but they don't actually provide the real hooks and they do very well in the benchmarks, right? They do very well on like kind of the public leaderboards. But they don't actually provide the hooks that …”
Huber: Over 90% of enterprise AI use cases are retrieval-augmented generation
“I think like today, 90 plus percent of it in enterprises is retrieval event generation, or it's, you know, using retrieval, it's sort of a chat on top of unstructured data.”
Embedding search and analytics enable developers to improve model reliability
“By looking at embeddings doing embedding search, doing analytics over embedding space you could give developers you know, at the minimum of a divining rod, if not a compass, to be able to improve their models and get to the level of reliability they want to ha…”
Huber: Programmable memory enables reliable LLMs across all use cases
“Chroma's belief is that programmable memory, so developers being able to set terministically Hey, language model, this is the knowledge you should know about, this is the knowledge you should use, these are the tools you should know about, these are the tools …”
Huber: 'Chat Your Data' AI Use Case Will See Mass Adoption
“So I think that use case, even just the Chat Your Data use case, truly will go to the ends of the earth.”
Huber: Steering LLMs at embedding layer is not exposed in closed-source models
“So, there's lots of stuff around steering language models at the embedding layer itself, and not using text but this is not yet exposed To at least closed source models, so.”
Huber: AI-native databases will be much thicker than traditional databases
“We think the database will be, like, much thicker than it's been before.”
Huber: Multimodal models will run directly inside application code and databases
“There'll be language models running inside the application code, obviously, language models running inside the database as well or large models more broadly, multimodal models will run, you know, everywhere as well.”
Huber: Most enterprises will deploy language models within three years
“I certainly think probably most enterprises, organizations, companies on earth will have brought language models Into the company, probably in pretty meaningful ways. At minimum, the customer service department, the sales department, ops, back end, legal and h…”