Assertion Supported AI assessment confidence: 95% certainty 4/5 debate potential 3/5

Alexandr Wang: Held-out benchmarks reveal several AI models underperform their reported scores.

Alexandr Wang · No Priors Ep. 65 | With Scale AI CEO Alexandr Wang · May 22, 2024 · at 24:20

Scale AI CEO Alexandr Wang discusses benchmark contamination and Scale AI's GSM1K evaluation paper with No Priors host Sarah Guo.

0:00 / 0:27exact quote · 27.6s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“So we, one of the things we did is we published DSM-I-K, which was a held out eval. So we basically produced a new evaluation of the math capabilities of models. That there's no way it would ever exist in the training data set to really see how much of the, how were the performance of the models were the reported performance of the model capability versus the actual capability. And what you notice is some of the models perform really well, but some of them perform much worse than the reported performance.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Alexandr Wang

Prediction Not checkable as stated
Wang: The path to AGI resembles curing cancer, not a vaccine
“My biggest belief here is that the path to AGI is is one that looks a lot more like curing cancer than developing a vaccine. And what I mean by that is I think that the path to build AGI is going to be in, in, you know, you're going to have to solve a bunch of…”
Alexandr Wang Dec 26, 2024 ▶ 23:49 No Priors Ep. 95 | Best of 2024
Assertion Contradicted
Wang: AI gets no positive transfer across modalities like video to text
“My understanding there's no positive transfer from learning in one modality to other modalities. So like training off of a bunch of video doesn't really help you that much with your text problems and vice versa.”
Alexandr Wang Dec 26, 2024 ▶ 25:42 No Priors Ep. 95 | Best of 2024
Assertion Not checkable as stated
Wang: No strong evidence that video creates useful AI world models
“I don't think there's strong scientific evidence of that yet. Maybe there will be eventually.”
Alexandr Wang Dec 26, 2024 ▶ 26:11 No Priors Ep. 95 | Best of 2024
Prediction Not checkable as stated
Alexandr Wang: Achieving AGI will take multiple decades of solving individual problems.
“My biggest belief here is that the path to AGI is is one that looks a lot more like curing cancer than developing a vaccine. And what I mean by that is I think that the path to build AGI is going to be in, in, you know, you're going to have to solve a bunch of…”
Alexandr Wang May 22, 2024 ▶ 34:53 No Priors Ep. 65 | With Scale AI CEO Alexandr Wang
Assertion Partly supported
Alexandr Wang: Cross-modality training produces no positive transfer between video and text.
“I think the main thing, fundamentally, is I think there's very limited generality that we get from these models and even for multimodality, for example my understanding there's no positive transfer from learning in one modality to other modalities. So like tra…”
Alexandr Wang May 22, 2024 ▶ 36:37 No Priors Ep. 65 | With Scale AI CEO Alexandr Wang
Assertion Supported
Wang: Standard academic benchmarks are contaminated by training data overfitting
“Most of the benchmarks that we as a community look at... Academic benchmarks that are what the industry used to measure the performance of these algorithms are fraught with issues. Many of the models are overfit on these benchmarks. They're sort of in the trai…”
Alexandr Wang Jul 11, 2024 ▶ 23:21 No Priors Ep. 71: The Best of 2024 (so far) with Sarah Guo and Elad Gil
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.