why aren't all 41 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 1 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Not checkable as stated
Evans: AI foundation models lack winner-takes-all network effects
“And the problem is that as far as we can see, there is no winner takes all effect or network effect in a foundation model.”
Prediction Not checkable as stated
Evans: VR headsets will not replace smartphones as primary computing platforms
“You're not going to wear a headset, no matter how light and cheap it is, all day outdoors. I would not have worn it walking here, even if it weighed a hundred grams and had completely perfect pass-through. Therefore, it can't replace your phone. Therefore, we …”
Assertion Partly supported
Evans: OpenAI has 900M weekly active users, but only 5% pay
“You've got nine hundred million weekly active users, but most of them are not using it every day and can't think of anything to do with it. And only five percent of them are paying for it.”
Assertion Not checkable as stated
Evans: AI frontier models are commodities spread across half a dozen organizations
“The thing that's become very clear, if it wasn't clear a year ago, is that the models themselves are sort of commodities in that, you know, there's half a dozen people who have a state-of-the-art model.”
Assertion Partly supported
Evans: OpenAI's Deep Research inverted numbers in its own marketing example
“Then the interns typed the number in wrong. Like, it was literally the wrong percentage. It was like, 65, 35, instead of 35, 65.”
Prediction Not checkable as stated
Evans: AI is not currently on a path to 100% factual accuracy
“Are you telling me this is going up to the point that I'm going to be able to use deep research and the numbers will all be right, and I'll know that they're all right? Because I don't think we're on a path to that. Or at least I don't think we know that we're…”
Assertion Not checkable as stated
Evans: No standalone breakout consumer apps exist wrapping ChatGPT API
“We don't have a breakout, there's no standalone breakout consumer app. There's no one, there's no, there's all these, there's, there are all these enterprise SaaS stuff. There is not really a consumer equivalent. There aren't hundreds of consumer apps using th…”
Assertion Not checkable as stated
Evans: Public discourse around AI doomerism has effectively vanished
“Yes, all of the dumerism's gone away.”
Assertion Not checkable as stated
Evans: Fully autonomous end-to-end AI workflows require AGI
“But you kind of know at the same time that in practice, what I've just described would kind of require AGI. In practice, there's like 10 ways that that's going to break And never mind the hallucination problem, which is a whole separate conversation. There's j…”
Assertion Not checkable as stated
Benedict Evans: Science lacks an underlying theory for how LLMs work
“We don't have any equivalent set of theories for intelligence or artificial intelligence. We have a lot of theories of how some bits of it might work. But we do not have a theory of what we have and what, and in what sense is what we have is different and the …”
Assertion Not checkable as stated
Benedict Evans: The tech industry lacks enough human data to 100x LLMs
“And so that means you kind of can't do like a prediction. There's no Moore's law here where you can say, well, it'll get to that power of compute level at this, but at this point set aside the fact that we actually don't have enough data to give it 10 or a hun…”
Prediction Open · timeframe Mar 2031
Evans: Microsoft Bing will never catch up with Google in search
“Doesn't matter how much money and how hard Microsoft works, Bing will never catch up with Google.”
Assertion Supported
Evans: Submitting 1,000 prompts placed ChatGPT users in top 20%
“Turns out that if you did a thousand posts, if you did a thousand prompts last year, you're in the top 20%.”
Prediction Not checkable as stated
Evans: AI coding will lead to far more total software volume
“There will be way more software, and that will pick up many more of those use cases, either that weren't automated before, either because they were too small, or because you could actually couldn't automate that thing with software before, and now with AI you …”
Prediction Not checkable as stated
Evans: No one will vibe code their own ERP software
“No, no one will vibe code their own ERP or their own frame.io, but they may ask Anthropic or Gemini or ChatGPT, can you do this thing for me?”
Assertion Not checkable as stated
Evans: Initial enterprise deployments of Microsoft Copilot were largely unsuccessful
“Everyone deployed co-pilot and went, oh, okay, that wasn't very successful.”
Assertion Not checkable as stated
Evans: Current AI models still cannot replace existing software like Excel
“But it still can't actually replace any of the software you use. It can't replace Excel, it can't.”
Assertion Not checkable as stated
Evans: Generative AI is deployed in thousands of companies, unlike crypto
“There are hundreds and hundreds of companies who've already got this in production, doing stuff that's really useful where it works, where you understand what it is. So this is just objectively wrong to say that it's useless. It's already not, not in the way t…”
Assertion Supported
Evans: Google and Meta delayed LLMs in 2022 due to high error rates
“This is why Google and Meta didn't launch their own LLMs in twenty-twenty-two when they had them as well, because they looked at them and said, well, they're wrong too much.”
Assertion Not checkable as stated
Evans: Foundational LLMs are similar, but OpenAI holds consumer mindshare
“The models are all sort of the same, but OpenAI is the only one that anyone uses, that, that, that has consumer mindshare.”
Assertion Partly supported
DeepMind AI discovered unknown biological differences between male and female retinas
“DeepMind did a project with Moorfields, which is iHospital in the UK. And they were looking at retinas. And their system discovered a difference between male and female retinas, and apparently medical science didn't actually know there was a difference between…”
Assertion Not checkable as stated
Evans: Meta Quest 3 lacks traction with roughly 10M active users
“Meta has the Quest III, which is a perfectly good, credible consumer device. It does not have traction. Does not have a, it's probably got, I don't know, maybe ten million active users. Huge abandonment rate in the past. It's not really good for anything other…”
Assertion Not checkable as stated
Evans: Three to six organizations can create frontier AI models
“So that means that we've got, like, pick a number between three and six or maybe more organizations that can make a frontier model, and they keep leapfrogging each other every couple of weeks or every couple of months.”
Assertion Not checkable as stated
Evans: 10% of population uses LLMs daily, 50% monthly
“If you look at the usage data, something like 10% of the population is using these things every day, but another 50% are using it every week or every month.”
Assertion Supported
Evans: Typical big US company uses 400 to 500 vertical SaaS apps
“And so this is why the typical big US company today has, depending on your numbers, like four to 500 vertical SaaS apps.”
Assertion Supported
Evans: Spreadsheets increased finance employment rather than causing job losses
“Spreadsheets did not result in a collapse in the number of people working in finance. They're quite the opposite. You have way more people in finance, because now it's possible to do all this more, all this new stuff that you couldn't have done before.”
Assertion Contradicted
Evans: Corporate audit costs stayed flat since early 2000s despite software progress
“Average audit data, all sorts of data for audit costs since the Barnes Oakley, which has basically been flat since about since the early 2000, despite everything that's happened in software”
Assertion Supported
Evans: Nvidia sells integrated computer systems, not just GPU chips
“People still think of NVIDIA as making GPUs in the sense that they make chips and sell chips. That's not what they do. They sell computers. They sell custom computers, kind of like Sun Microsystems did with a whole networking stack and a software stack on top …”
Assertion Supported
Evans: Big Tech data center spending will exceed $300B this year
“Google, Meta AWS, not Amazon overall, AWS only, and Microsoft spent about two hundred twenty billion dollars building data centers last year, and will spend about 300, maybe over 300 this year, depending on where their numbers come out.”
Assertion Not checkable as stated
Evans: Apple has user interest graph data but refuses to use it
“Apple also in principle has an interest graph. It just refuses to use it.”
Assertion Partly supported
Evans: OpenAI lacks visibility into user search, purchases, and social media activity
“Back to open AI, they've got a partial view on you, but like, they don't know what you've bought. They don't know what you've searched for. They don't know where you go. They don't know what Instagram you look at and what TikTok you look at and what YouTube yo…”
Assertion Supported
Evans: Accenture recorded $1.4B in generative AI bookings in one quarter
“Accenture built 1.4 billion, booked 1.4 billion of new generative AI bookings last quarter.”
Assertion Partly supported
Evans: 20% to 30% of large companies have deployed AI projects
“Every big company is now, like, 20 to 30% of big companies have got stuff in deployment, but every big company's got pilots.”
Assertion Not checkable as stated
Evans: Siri and Alexa were decision trees despite natural language capabilities
“There was like a trap with Siri and Alexa, which was that natural language processing worked, so you thought it was AI, and it wasn't. It was actually still just an IVR, it was still just a tree.”
Assertion Not checkable as stated
Evans: 2023's immediate generative AI use cases were coding and brainstorming
“But we're still, and we had, I think in 20, 23, that wave of the initial, oh my God, you can use it for that right now things, which is basically coding and brainstorming.”
Assertion Supported
Evans: The typical enterprise today runs 400 to 500 SaaS apps
“Why is it that the typical enterprise today has four to 500 SaaS apps?”
Assertion Supported
Evans: Predictions that only tech giants had enough AI data were wrong
“And this is clearly what happened with the last wave of machine learning. There was a brief moment where people said it needs all this data. Only Google's got all the data. There's going to be, like, three people who've got enough data to do AI, and that turne…”
Assertion Not checkable as stated
Benedict Evans: We cannot predict what happens when doubling LLM training data
“We can't do that with LLMs either. We don't know what will happen if you put double the data in or why, or we don't know why it works with this much data or not.”
Prediction Not checkable as stated
Benedict Evans: Generative AI will inevitably experience a market bubble
“Clearly, if we're not in a bubble now, we're going to have a bubble. It's because that's just like the nature of the light, the cycle of life. There will be a bubble around each new technology.”
Assertion Supported
Evans: Nvidia hit over $70B in trailing 12-month free cash flow
“I mean, I haven't looked at, I haven't updated my number here, but like, I think Q three last year, I think NVIDIA had something over seventy billion dollars of trading 12 months free cash flow.”
Assertion Supported
Evans: ChatGPT usage on Google Trends drops sharply in summer and Christmas
“I mean, it's funny if you look at Google Trends. There's a big sag in the summer and then a big sag in the Christmas week.”