Everything Brandon McKinzie said on any show that made the record, most notable first. Each card names its show and opens the statement there.
McKinzie: Tools prevent reasoning models from degrading during test-time compute
“We've in the past for our reasoning models talked a lot about test time scaling, and I think for a lot of problems you know, without tools, test time scaling might occasionally work and, but at some point the model is just kind of ranting in its internal chain…”
McKinzie: General reasoning models could unify with robotics foundation models
“And I personally don't see any reason why we couldn't have this, these be this, the same model.”
McKinzie: Tool use noticeably changes test-time scaling for visual reasoning
“We've seen exactly that, like the test time scaling slopes for, without tool use and with tool use for visual reasoning specifically are very noticeably different.”
McKinzie: Reinforcement learning is the key differentiator behind o3 reasoning
“I guess the short answer is reinforcement learning is, is the biggest one. So yeah, rather than just having to predict the next token and some large pre-training corpus from, you know you know, everywhere essentially now we have a more focused goal of the mode…”
McKinzie: Tool use improves test-time scaling slopes for visual reasoning
“And we've seen exactly that, like the test time scaling slopes for without tool use and with tool use for visual reasoning specifically are very noticeably different.”
McKinzie: OpenAI has run out of reliable evaluation benchmarks for recent models
“Especially with some of our recent models where we've kind of run out of Reliable evals to track because they kind of just solved a few of those.”
McKinzie: OpenAI is developing models with precise uncertainty understanding
“And I hope we can get to a place where our models have a more precise understanding of their own level of uncertainty. Because you know, if they already know the answer, they should just kind of tell you it. And if it takes them a day to actually figure it out…”
McKinzie: OpenAI models hit an inflection point navigating internal codebases
“I think our models are getting a lot better very quickly at being actually useful. And it seems like they were kind of reaching some kind of inflection point where They are useful enough to want to reach out to and use like multiple times a day for me at least…”
McKinzie: AI research consists of modular tasks ripe for automated optimization
“And there's so many like different components of research too. There's, it's not just you know, sitting off in the ivory tower thinking about things, but there's like hardware there's you know, various components of training and evaluation and stuff like this.…”
McKinzie: OpenAI models use external tools with weirdly human-like intuition
“It's also surprising to me how intuitively our models do use the tools we give them access to. It's like weirdly human-like, but I guess that's not too surprising given the data they've seen before,”
McKinzie: Multi-agent RL is a good baseline for human collaboration
“There's no reason you can't scale all this up so that models are trained to be really good at cooperating with each other. I mean, there's a lot of already existing literature on multi-agent RL and yeah, if you want the model to be good at something like colla…”
McKinzie: Over 90% of internet clock images show 10:10, biasing vision models
“It's like over 90% or something like that of all clocks on the internet are 10 10.”
McKinzie: Humans are an extremely expensive tool call for AI
“Yeah, we are a super expensive tool call. You know, if you're a model, you can either ask me, you know, meat bag over here to you know, help with something and I'll try to think really slowly. In the meantime, it could have like used browser and read like a hu…”