Scale AI, every mention
29 scenes, the whole family · ← back to Scale AI
tap a year for its mentions
every year anyone Nathan Lambert 9Shawn Wang 8Varun Mohan 5Mike Conover 4Alessio Fanelli 4Yi Tay 2Pranav Reddy 1Martin Casado 1Kyle Corbitt 1Joon Sung Park 1
Verbatim, from the transcripts: the passages where Scale AI comes up
Simulating Humanity: from Generative Agents to 8 Billion Digital Twins — Joon Sung Park, Simile AI
- ▶ 30:33 Joon Sung Park So you go get their data from places like my core scale.
The $15B Physical AI Company: Simulation, Autonomy OS, Neural Sim, & 1K Engineers—Applied Intuition
- ▶ 4:28 Alessio Fanelli Domitian, I remember initially, I think people put you and Scale.ai very similarly for some things about being kind of like on the data infrastructure side of things. 2 times in the scene
The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals
- ▶ 10:48 unnamed speaker And then, uh, CBench Pro will, will be some of the next one, which is an effort from scale.
Inside AI’s $10B+ Capital Flywheel — Martin Casado & Sarah Wang of a16z
- ▶ 21:22 Martin Casado I would say scale AI was actually a horizontal one for, you know, for robotics early on, so that sort of thing we're very, very interested, but the actual, like, robot interacting with the world is probably better for a different team.
Terminal-Bench 2.0: the most impt coding agent benchmark of 2025 gets a v2! Launch + Q&A w/ founders
- ▶ 28:11 Andy Konwinski Aaron Shaw, Ph.D.: companies like surge and scale and the rest and being able to do to find an incentive structure and the open ecosystem that gets extremely high quality training training and or, in this case, evaluation data generated is…
Why RL Won — Kyle Corbitt, OpenPipe (acq. CoreWeave)
- ▶ 32:09 Kyle Corbitt Um, but you know, like, look, you can say the same about scale AI and all of their competitors that are like, you know, many billion dollar companies that have basically the exact same customer set.
A Technical History of Generative Media
- ▶ 52:58 Gorkem Yurtseven Or like scale AI for Imogen video models, like data, data collection, like more, more prepared data sets for, for video models, like effects, different camera angles.
Better Data is All You Need — Ari Morcos, Datology
- ▶ 10:20 Alessio Fanelli Because I think from the outside, you have companies like scale that obviously have become super successful. 2 times in the scene
- ▶ 1:16:32 Ari Morcos I mean, I think one thing that's pretty cool about it, obviously, is it showcases the importance of data, um, that Meta is willing to spend, uh, quite this much on, uh, you know, scale, uh, kind of, not acquisition, not acquisition, uh,…
The AI Agenda: GPT5 leaks and the business of AI News — Steph Palazzolo, The Information
- ▶ 45:29 Shawn Wang I mean, I think the, the thing that is uncertain for me is like, and there's been like five of these, a character scale, all these.
The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
- ▶ 15:18 Nathan Lambert I mean, it's almost like how I see scale. 3 times in the scene
⚡️Using RFT to Build Clinical Superintelligence
- ▶ 17:42 Shawn Wang Yeah, I think like there's the sort of test environments, uh, where you sort of, let's say you hire domain experts and clinicians to be labellers for you, like a scale AI or whatever, but like also that's pretty artificial.
🕰️ The Oral History of Windsurf (ft. Varun Mohan, Scott Wu, Jeff Wang, Kevin Hou, Anshul R)
- ▶ 2:37:38 Varun Mohan He was a year younger than us at MIT, and he's figured out how to constantly change the direction of the company. 5 times in the scene
2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents
- ▶ 13:48 Shawn Wang She's pointing to the, uh, the scale seal leaderboard.
- ▶ 14:02 Shawn Wang And scale is one of maybe three or four people this year that has like really made an effort in, in doing a credible private test set leaderboard.
- ▶ 33:04 Shawn Wang I think there was a bit of a fight between Scale AI and the synthetic data community, because Scale published a paper saying that synthetic data doesn't work. 2 times in the scene
The State of AI Startups in 2024 [LS Live @ NeurIPS]
- ▶ 5:32 Pranav Reddy So this is from the scale leaderboards, which is a set of independent evals that are not contaminated.
[Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu
- ▶ 21:42 unnamed speaker Because it's not like anyone from Mechanical Turk or anyone from Scale can actually just do this.
[Paper Club] Who Validates the Validators? Aligning LLM-Judges with Humans (w/ Eugene Yan)
- ▶ 19:30 Eugene Yan Sometimes if you're not labeling the data yourself, if you're using something like scale AI,
The Ultimate Guide to Prompting - with Sander Schulhoff from LearnPrompting.org
- ▶ 20:18 Shawn Wang I'll call out two recent papers, which people might want to look into, which is a Salesforce yesterday released a paper called diversity empowered intelligence, which is, uh, I think a shot at the bow of, of, for scale AI.
[LLM Paper Club] Llama 3.1 Paper: The Llama Family of Models
Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
- ▶ 35:32 Shawn Wang So just infinite amounts more while, you know, scale AI and all the, the annotation providers are very happy to hear that.
The 10,000x Yolo Researcher Metagame — with Yi Tay of Reka
How AI is Eating Finance - with Mike Conover of Brightwave
- ▶ 29:41 Mike Conover But you also saw scale. 4 times in the scene
The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
- ▶ 42:05 Nathan Lambert I think places like scale are also really big in this, where a lot of the labeler companies will help control, like,
- ▶ 52:05 unnamed speaker Um, just in a domain of, uh, human preference data suppliers, uh, Scali, I very happily will tell you that they, they supplied, uh, all that data for Lama too.
- ▶ 1:30:44 Nathan Lambert On the research side, there's a lot of interesting things as, like, does getting your data from scale or a Discord army change the quality of the data based on, like, professional contexts? 3 times in the scene
- ▶ 1:34:33 Nathan Lambert I mean, Scale's probably trying to do it. 2 times in the scene