why aren't all 24 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Prediction Not checkable as stated
Wang: The path to AGI resembles curing cancer, not a vaccine
“My biggest belief here is that the path to AGI is is one that looks a lot more like curing cancer than developing a vaccine. And what I mean by that is I think that the path to build AGI is going to be in, in, you know, you're going to have to solve a bunch of…”
Assertion Contradicted
Wang: AI gets no positive transfer across modalities like video to text
“My understanding there's no positive transfer from learning in one modality to other modalities. So like training off of a bunch of video doesn't really help you that much with your text problems and vice versa.”
Assertion Not checkable as stated
Wang: No strong evidence that video creates useful AI world models
“I don't think there's strong scientific evidence of that yet. Maybe there will be eventually.”
Prediction Not checkable as stated
Alexandr Wang: Achieving AGI will take multiple decades of solving individual problems.
“My biggest belief here is that the path to AGI is is one that looks a lot more like curing cancer than developing a vaccine. And what I mean by that is I think that the path to build AGI is going to be in, in, you know, you're going to have to solve a bunch of…”
Assertion Partly supported
Alexandr Wang: Cross-modality training produces no positive transfer between video and text.
“I think the main thing, fundamentally, is I think there's very limited generality that we get from these models and even for multimodality, for example my understanding there's no positive transfer from learning in one modality to other modalities. So like tra…”
Assertion Supported
Wang: Standard academic benchmarks are contaminated by training data overfitting
“Most of the benchmarks that we as a community look at... Academic benchmarks that are what the industry used to measure the performance of these algorithms are fraught with issues. Many of the models are overfit on these benchmarks. They're sort of in the trai…”
Insight
Alexandr Wang: Data abundance is the fundamental bottleneck for post-GPT-4 models.
“The key to the scaling of these large language models and the, you know, these language models in general is the ability to scale data. And I think that one of the fundamental bottlenecks to, you know, what's, what's in the way of us getting from GPT-IV to GPT…”
Assertion Not checkable as stated
Alexandr Wang: The AI industry has exhausted all easy internet training data.
“And we've sort of, as a community, we have, we've had easy data, which is all the data on the internet and we've kind of exhausted all the easy data, and now it's about, you know, forward data production that has high supervisory signal that is basically very …”
Prediction Not checkable as stated
Alexandr Wang: Human-AI teams will outperform standalone models for a long time.
“The question is, is a human plus a model together going to be able to produce better output than a model alone? And I think that'll be the case for A very, very, very long time. That, that humans are still, you know, human intelligence is complementary to mach…”
Assertion Supported
Alexandr Wang: Academic AI benchmarks are contaminated and models are overfit.
“The academic benchmarks that are what the industry used to measure the performance of these algorithms are fraught with issues. Many of the models are overfit on these benchmarks. They're sort of in the training data sets of these models.”
Assertion Supported
Alexandr Wang: Held-out benchmarks reveal several AI models underperform their reported scores.
“So we, one of the things we did is we published DSM-I-K, which was a held out eval. So we basically produced a new evaluation of the math capabilities of models. That there's no way it would ever exist in the training data set to really see how much of the, ho…”
Opinion
Alexandr Wang: Multimodality is a lateral move; the industry needs smarter models.
“So, you know, we got multi-modality capability.
That's exciting.
It's more of a lateral expansion of the models, and the industry needs smarter models.
We need GPT-V, or we need Gemini-II, or whatever that, those models are going to be.
and so to me it was, y…”
Prediction Not checkable as stated
Wang: Society will adapt smoothly due to slow AI progress
“I think it's actually a pretty bullish case for society adapting the technology because I think it's going to be, you know, consistent, slow progress. For quite some time, and society will have time to fully sort of acclimate to the technology that develops.”
Assertion Supported
Alexandr Wang: Scale built data infrastructure for DoD's first AI program.
“So we built the very first data engines to support government data. This would support mostly geospatial and satellite and over, other overhead imagery. This ended up fueling the first AI program of record for the US DoD.”
Assertion Not checkable as stated
Alexandr Wang: Scale AI fuels basically every major LLM developer.
“Today, you know, fast forward to today our data foundry fuels basically every major large language model in the industry. Work with OpenAI, Meta, Microsoft, many of the other players.”
Assertion Not checkable as stated
Alexandr Wang: Frontier AI models can no longer learn much from Reddit.
“It's not Any more the case that these models can learn that much more from, you know, various comments on Reddit or whatnot. They need, ah, they need truly frontier data.”
Disclosure
Alexandr Wang: Scale AI's core thesis is hybrid human-AI synthetic data.
“And our perspective is that the critical thing is, is what we call hybrid human AI synthetic data. So how can you build hybrid human AI systems such that AI are doing a lot of the heavy lifting, but human experts and people, you know, the basically best, brigh…”
Opinion
Alexandr Wang: GPT-4 was too early a model to sustain application hype.
“GPT-IV, I think, as a model, was a little early of a technology for us to have this entire hype wave around, and I think we, you know, the community very quickly discovered all the limitations of GPT-IV... It was probably a few generations too early of a model…”
Insight
Alexandr Wang: High-quality frontier data is 10,000x more valuable than enterprise data.
“One of the things that every, you know, all the model developers understand well, but the enterprises understand super well is that you know, not all data is created equal and high quality data or frontier data is, is, can be, you know, 10,000 times more valua…”
Assertion Supported
Alexandr Wang: Scale partnered with OpenAI on the first GPT-2 RLHF experiments.
“So we partnered with OpenAI at that time to do the very first experiments on RLHF on top of GPT-II.”
Assertion Supported
Alexandr Wang: JPMorgan has 150 petabytes of data; GPT-4 used under one.
“JP Morgan's proprietary data set is a 150 petabytes of data. GPT-IV is trained on less than one petabyte. Of data.”
Disclosure
Alexandr Wang: Scale AI will launch recurring held-out LLM benchmark leaderboards.
“So one is that we're going to launch these private held out evaluations and have leaderboards associated with these evals for the leading LLMs in the ecosystem. And we're going to rerun this contest periodically. So every few months we're going to do a new set…”
Insight
Alexandr Wang: Multimodality faces a scarcity of quality data for personal agents.
“So multimodality as an entire space is one where for the same reasons that we've like exhaust a lot of the internet data, there's a lot of scarcity for good multimodal data that can empower these personal agents and these personal
Use cases.”
Assertion Partly supported
Alexandr Wang: Scale AI built the first sensor-fusion data engine for AVs.
“And so we built, The very first data engine that supported sensor fused data. So support a combination of two D data plus three D data. So lidars plus cameras that were built on onto the vehicles. And then that very quickly became an industry standard across a…”