why aren't all 853 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Supported
Baidu's Deep Speech engine achieves superhuman accuracy on short queries
“The speech engine that we've built at Baidu called Deep Speech it's actually superhuman for these short queries.”
Prediction Held up
Greene: Enterprise tech will consolidate into four clouds immune to startup disruption
“Everything's going to be in just a few, maybe four clouds or so, you know, there's not going to be that many of them, because like, we spend about ten billion a year in infrastructure, and so it's not going to Get disrupted by a startup, unless they invent qua…”
Prediction Held up
Benet: Decentralized storage nodes will beat AWS on unit economics
“Our bet there is that yes, that there's a whole bunch of places and cases where Certain individuals or groups in the world have access to either really cheap storage or storage that's positioned well in the network that is kind of like, you know, somewhere bet…”
Assertion Supported
Altman: Chesky asked if candidates would work with one year to live
“He used to ask people before he hired them at Airbnb if they would take the job if they got a medical diagnosis that they had one year left to live.”
Assertion Supported
Stuart: Multi-GPU compilers often produce kernels slower than baseline
“Two is to use compiler-based approaches, but we find these compilers, compilers to be quite suboptimal. They produce kernels that are sometimes slower than non-overlapped baselines”
Assertion Supported
Stuart: Bypassing NCCL intermediate buffers speeds up all-reduce by 80%
“For example, Nickel's default mode forces intermediate buffers which adds extra data movement between the sender and the receiver. And for fine-grained communication, this overhead really accumulates, and by stripping it out, you can speed up an operation as s…”
Assertion Supported
Stuart: ParallelKittens matches hand-optimized kernels in 50 to 100 lines
“And what we find is that with roughly 50 to 100 lines of device code, PK is able to surpass or match hand-optimized kernels that are often hundreds to thousands lines of code.”
Assertion Contradicted
Scholl: Douglas Aircraft in 1921 was the last startup to build airliners
“The last company entrepreneur founded that went on to build a commercial airliner was Douglas Aircraft in 1921.”
Assertion Contradicted
Scholl: The FAA approved Boom's XB-1 flight permit in just 90 minutes
“The approvals process for XB-One, that even for an R&D airplane can stretch on for months, actually happened in 90 minutes. From when we said we think we're ready to fly, and FAA said, here's the paperwork, we agree, go. We got the first ever permit for flying…”
Assertion Partly supported
Scholl: US legalized supersonic flight 115 days after XB-1 broke sound barrier
“And so it was about 24 hours from breaking the sound barrier to be invited to the West Wing, and a 115 days from breaking the sound barrier to an executive order making supersonic flight legal again in the U.S.”
Assertion Supported
Scholl: XB-1 Broke Sound Barrier Six Times With No Ground Boom
“XB one broke the sound barrier on two flights, six times through the sound barrier. Each time, no audible sonic boom in the ground.”
Assertion Supported
Anthropic Deleted 80% of Claude Code System Prompt for Opus 5
“So yeah, we deleted 80% of the system prompt.”
Assertion Supported
Cherny: Claude Code runs in production on Bun's AI-rewritten Rust codebase
“This is in production now. This is what quad code uses now when you're running it.”
Assertion Supported
Huang: Radiology jobs grew 20% over recent years despite AI scan automation
“The task of reading radiology scans has been automated, but the number of radiology jobs has increased some 20% in the last several years, even though AI has taken over the whole field.”
Assertion Partly supported
Pomel: Major Datadog Investor Dumped All Shares When COVID Lockdowns Began
“Our lockup expired the day of the COVID lockdowns. So I don't know if you remember, there was a big market crash that day. It was also a lockup expiring. One of our biggest investors got scared and he dumped everything that day. So I think, I don't know, we're…”
Assertion Supported
Pincus: Early backers sold Facebook stock because none foresaw its ultimate scale
“I don't think Reed or I Ever imagined how, I know none of us ever imagined how big this, Peter Thiel, all of us, we all sold our stock in Facebook, so we voted with our shares, you know, none of us could have ever described how big this all got.”
Assertion Contradicted
Bryant Chou: Webflow was the first no-code application
“We created this visual interface over HTML, CSS, and JavaScript, and it was sort of, you know, the first no-code application that was out there.”
Assertion Contradicted
Tan: Delaware C Corps Require Relentless Pursuit of Profit
“Basically, you know, if you're a Delaware C Corp, you have to relentlessly pursue profit. Otherwise there's grounds to remove you.”
Assertion Contradicted
Ries: FTX bankruptcy's Anthropic stake exceeds the entire fraud value
“Apparently the stake that the bankruptcy has of those shares is worth more than the whole, than all of the entire fraud by a lot.”
Assertion Partly supported
Ries: Costco routinely receives the worst corporate governance ratings
“Costco came under attack in the early 2000 for having these non-standard governance practices. In fact, Costco routinely gets the worst possible governance rating from governance rating people.”
Assertion Supported
Ries: Twilio fired Jeff Lawson despite revenue growth since IPO
“So at the time he was fired, the stock was down like 80% from the peak. And it's like, oh, well case closed. But if you measure from the IPO or even from the pandemic peak, revenue was up. It's like, did the business go down? Was revenue down? Was there some k…”
Assertion Partly supported
Graham: Silicon Valley VCs get better returns than European VCs
“Silicon Valley investors get better returns than European investors despite having to decide so quickly.”
Assertion Supported
Tan: Professional engineers average only 30 to 50 lines of code daily
“If you look at the literature about software engineering, going back to like 2000. 1990. I mean, it's pretty clear that the average number of lines of code that a professional software engineer that's like tested and production ready, it's not like a hundred l…”
Assertion Contradicted
Chaubard: 7M parameter TRM scored 87% on ARC Prize 1
“And so it's a twenty-eight million parameter model for HRM. Now she brings it down to a seven million parameter model. It actually gets from 70% to 87% on on ArcPrize one. And does actually quite well on ArcPrize two as well.”
Prediction Partly held up
Hassabis: Frontier model capabilities reach tiny edge models within 6-12 months
“But I think for now there's the assumption we make is that, you know, a year later after one of our leading, you know, pro models or frontier models goes out half a year later, you'll have them in the,
The really tiny, almost edge models.”
Assertion Supported
Tan: GStack has more GitHub stars than Ruby on Rails in three weeks
“I built GStack to encode this three weeks ago, and now it has more GitHub stars than Ruby on Rails.”
Assertion Supported
Vuong: OpenX generalist robot model outperformed embodiment-specific specialists by 50%
“You can compare it to the specialist that has been optimized to work well on a particular embodiment. How does it compare? And the interesting result from OpenX is it was 50% better.”
Assertion Supported
Chollet: Base LLMs Score Under 10% on ARC-AGI-1
“So basal alarms were scoring extremely low on V-one, like sub-ten percent, basically. And, I mean, it was true of the original, like, GPT-III actually scoring zero, but that's even true of the latest basal alarms today, you know, as of March.”
Assertion Supported
Hu: Confluence Labs Saturated the ARC-2 Benchmark at 97% Accuracy
“I actually worked with a company in the winter 26 batch not too long ago called Confluence Lab, which actually ended up saturating the V-two results with 97%, and I think their task cost was a lot more efficient too.”
Assertion Contradicted
Hodak: Science trial achieved the first coherent form vision restoration
“As far as I know, our clinical trial was the first time ever that form vision had been, like, had created, like, a coherent image in the mind's eye of a person.”
Assertion Contradicted
Hodak: A major organ perfusion firm makes more from private jet logistics
“Like one of the big companies in the space, it turns out that they're like private jet logistics business is bigger than their medical device business.”
Assertion Partly supported
Hodak: Science retinal implant allowed a blind patient to read eye charts
“And with our retinal prosthesis that we, what we saw in the trial was we can take a patient who's been unable to see faces for a decade and allow them to read every letter on an eye chart.”
Assertion Supported
Humans can learn to control individual cortical neurons within minutes via feedback
“If I put an electrode almost anywhere in your brain and then wake you up in during surgery, And I show you a flashing light that is, that flashes proportionally to how much that neuron is firing, at least almost anywhere in cortex. Within a couple of minutes, …”
Assertion Supported
Hodak: Prima implant produces normal black-and-white vision in blind patients
“The qualia of Prima is, is normal sight. It's black and white. It's only a, it's a small field of view, but it's vision.”
Assertion Supported
Hodak: Directly stimulating retinal bipolar cells produces images in the mind
“It was an empirical discovery of our study that if you excite the bipolar cells with an image, you get an image in the mind's eye, because that is clearly the critical processing step in the retina that you wanted to preserve.”
Assertion Supported
Hodak: Science found optogenetic proteins sensitive to ambient indoor office lighting
“What we were able to do were find optogenetic proteins that are so sensitive that they're sensitive to like indoor office lighting”
Assertion Supported
Hodak: Science's neural grafts formed biological brain connections in animal models
“But it comes with the potential of growing throughout the brain, forming biological connections all over the place. And I mean, that's what we've seen in, in the animal models. That's not in humans yet”
Assertion Supported
Poetiq outperformed Gemini 3 Deep Think on ARC-AGI-2 at half the cost
“Yeah, so the interesting thing is that we were half the cost of Gemini Three Deep Think because we were building on top of Gemini Three Pro, which is a much cheaper model. But we still got in the end, a nine percentage point improvement on the official verific…”
Assertion Supported
Poetiq scored 55% on Humanity's Last Exam, outperforming Claude Opus 4.6
“AI hasn't passed it yet, but we got to 55%, which is almost two percentage points higher than the previous state of the art. Which came out just last week from Anthropic with Claude Opus 4.6. They got 53.1%, and we got 55% on it.”
Assertion Contradicted
Mercury data shows 70 percent of startups choose Claude
“There was some stuff from Mercury that like, 70% of startups are, you know, choosing quad as their model of choice.”
Assertion Supported
SemiAnalysis reports Claude Code generates four percent of all public commits
“There were some other stuff from like semi-analysis that four percent of all public commits are made by quad code.”
Assertion Supported
NASA JPL used Claude Code to navigate the Perseverance Mars rover
“It plotted the course for perseverance, like for like the Mars Rover.”
Assertion Supported
French-Owen: Claude Code and Codex Use Grep Rather Than Semantic Embeddings
“Like I think cursor takes an approach where they actually do semantic search, where they embed everything and figure out like, Hey, what query is closest to this? If you look at a codex or a cloud code they actually just use like grep.”
Assertion Supported
French-Owen: Codex Uses Periodic Compaction to Support Long-Running Tasks
“Where it will run compaction, like, periodically after each turn, and so Codex can continue to run for a very long time, and if you look at the percentage in the CLI, you'll see it, like, move up and down as compaction runs.”
Assertion Supported
Tan: Gamma reached $100M ARR with only 50 employees
“Gamma was interesting to see, like one of the biggest things that they said in their launch that I think is a very good trend is they said they got to a hundred million dollars in ARR with only 50 employees.”
Assertion Supported
Kaplan: Scaling laws apply to reinforcement learning in AI training
“You can see scaling laws in the reinforcement learning phase of AI training.”
Assertion Supported
Kaplan: METR found AI task duration doubles roughly every seven months
“And an organization meter studied this very systematically and found yet another scaling trend. They found that If you look at the length of tasks that AI models can do, it's doubling roughly every seven months.”
Prediction Held up
Kaplan: AI training and inference will see 3x-10x yearly efficiency gains
“I think that over time, as AI becomes more and more widespread, I think that we're going to really drive down the cost of inference and training dramatically from where we are right now. Sort of, three X to 10 X gains algorithmically, and in sort of scaling up…”