The Ledger, every show
Every statement that passed quotation and attribution checks, across all 44 shows. Pick shows below, then mix any filter with any other.
shows 




every show 44 of 44
Misra: LLMs are just silicon doing matrix multiplication, not conscious beings
“You can rule out they're conscious. I mean, come on. And I said, you know, Anthropic makes great products. Cloud code is fantastic. Co-work is fantastic, but they are grains of silicon doing matrix multiplication. They don't have consciousness. They don't have…”
Misra: Scale will not solve everything; AGI requires a different architecture
“One of the misconceptions that exists today is that scale will solve everything. Scale will not solve everything. You need a different kind of architecture, and this continual learning is a difficult problem.”
Misra: Current LLM architectures cannot recursively self-improve into new paradigms
“That kind of self-improvement is not possible with these architectures. They can refine these. They can fill out these rows where the answer already exists.”
Transformers compute precise Bayesian posteriors down to 10^-3 bits accuracy
“We trained these models and we found that the transformer got the precise Bayesian posterior down to 10 to the power minus three bits accuracy. It was matching the distribution perfectly. So it is actually doing Bayesian in the mathematical sense, given a task…”
Transformers perform all Bayesian tasks, Mamba does most, and MLPs fail
“Transformer does everything. Mamba does most of it. LSTMs do only partially, and MLPs fail completely.”
Misra: Deep learning performs correlation rather than causation
“All of deep learning is, ah, doing correlations. It's not doing causation. Causal models are the ones that are able to do simulations and intervention.”
Deep learning remains in the Shannon entropy world, lacking Kolmogorov complexity
“I think deep learning is still in the Shannon entropy world. It has not crossed over to the Kolmogorov complexity and the causal world.”
Misra: AGI requires real-time plasticity and causal world models
“To get to what is called AGI, I think there are two things that need to happen. One is this plasticity. Which has to be implemented through container learning. Secondly, we have to move from correlation to causation.”
An LLM trained on pre-1916 physics will never derive general relativity
“Take an LLM and train it on pre-nineteen-sixteen or nineteen-eleven physics and see if it can come up with the theory of relativity. If it does, then we have AGI. I mean, it's a high bar, but, you know, we should have high bars. It won't.”
Misra: AGI requires generating new science beyond training data
“AGI will be when we are able to create new science, new results, new math. When an AGI comes up with a theory of relativity, it has to go beyond what it has been trained on. To come up with new paradigms, new science. That's my definition of AGI.”
Misra: LLM output is bounded by the inductive closure of training data
“So you know another phrase that we have been using recently is, you know, the output of the LLM is the inductive closure of what it has been trained on.”
Misra: Adding data to LLMs cannot create new mathematical manifolds
“There has to be an architectural lead that is able to create these manifolds and just throwing new data will not do it. It'll just smoothen out the already existing manifolds.”
Misra: Prompt engineering is prompt twiddling, not true engineering
“One term I really dislike is prompt engineering. You know, engineering used to mean sending a man to the moon or providing five nines reliability. Prompt engineering is prompt twiddling.”
Misra claims he built the first known implementation of RAG using GPT-3
“And I got GPD three to do in context learning, few short learning, and you know, it was kind of the First, at least to me, it was the first known implementation of RAG, Retrieval Augmented Generation, which I used to solve this problem of querying, getting GPT…”
Misra: LLMs function as compressed representations of prompt-probability matrices
“So in kind of an abstract way, what all these LLMs are doing is coming, coming up with a compressed representation of this matrix. And when you give a prompt, They try to approximate what the true distribution should have been and try to generate it.”
Misra: Geometric Bayesian signatures persist in large open-weight LLMs
“We took these frontier production LLMs, which have open weights so that we could look inside them. And we did our testing and we saw that the geometries that we saw in the small models persisted in models, which are, you know, hundreds of millions of parameter…”
Misra: LLM deception is driven by training data, not architecture
“That's not a function of the architecture. That's a function of the training data. It has been fed, you know, articles on Reddit or Asimo or whatever.”
Misra: LLMs trained on pre-1915 physics could not discover relativity
“Any LLM that was trained on pre-nineteen-fifteen physics would never have come up with a theory of relativity.”
Misra: Pure language processing is insufficient for human-level intelligence
“Language is great, but language is not the answer. You know, when I'm looking at catching a ball that is coming to me, I'm mentally doing that simulation in my head. I'm not translating it to language to figure out where it'll land.”
ESPN deployed a GPT-3 database query system in production in September 2021
“We deployed this In production at ESPN in September, 21.”
Misra: LLMs hallucinate when straying from learned Bayesian manifolds
“And as long as the LLM is going in sort of traversing through these manifolds, it is confident. And it can produce something which is, which makes sense. The moment it sort of wears away from the manifold, then it starts hallucinating and start spotting nonsen…”
Misra: Chain-of-thought prompting works by reducing LLM prediction entropy
“That's why chain of heart works. What happens with chain of thought is you ask the LLM to do something chain of thought. It starts breaking the problem into small steps. These steps it has seen in the past. It has been trained on maybe with some different numb…”
LLM prompt combinations exceed the number of electrons in the universe
“If you look at all possible combinations of 8000 tokens and 50,000, ah, vocabulary, the number of rows In this matrix is more than the number of electrons across all galaxies, right? So, so there's no way that these LLMs can represent it exactly.”
Misra: Adding context to prompts reduces LLM prediction entropy
“The moment you add more context. You make the prompt information rich. The prediction entropy reduces.”
Misra: GPT-3 prompt matrix rows exceed atoms in all known galaxies
“If you just take just the old first generation GPT-III model, which had a context window of 2000 tokens and a vocabulary of 50,000 next tokens or 50,000 tokens, then the size of it, the number of rows in this matrix is more than the number of atoms across all …”