why aren't all 13 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Insight
Lambert: OpenAI o1 uses large-scale RL on verifiable outcomes, not MCTS
“You should take open AI at their face value, which they are doing very large scale RL on the verifiable outcomes is what I've added, especially in context of the RL API that they've released, which I'll talk about more, but most of the reasons to believe in mo…”
Insight
Karina Nguyen: OpenAI o1 excels when given explicit hard constraints
“If you give a one like hard, like constraints of like what you're looking for, basically the model would be, we'll have a much easier time to like, kind of like select the candidates and match like the candidate that is most like, fulfill the criteria that you…”
Insight
Beauchamp: AI intelligence is generative LLMs combined with tree search
“I think if you want to talk about what would intelligence look like, it looks much more like tree search. Combining the generative nature of these LLMs with a really good tree search. And that's what opening I've done with O-one and O-three.”
Opinion
OpenAI likely reaches frontier capabilities via search, then distills into mini models
“The only way you reach the frontier with the full size models of O-one and O-three is with that stuff. And then you can distill to the minis, the O-one mini, O-three mini. So in my writeup, I said like, maybe this is the formula for O-one mini, O-three mini. T…”
Insight
Hylak: Users willing to wait five minutes for AI will wait an hour
“I think that it's like the number of tasks that you're willing to wait, you know, like 3:05 minutes for it. It's like probably pretty similar to the number of tasks you're willing to wait like an hour for, which is interesting.”
Assertion Not checkable as stated
McAteer: o1 is the first model to achieve one-shot codebase implementation
“Using O-one was, it was the first time where I would connect it to my IDE. I would provide the full context of my code base. I'll just create a file that concatenates all my files into one just simple text file, give it to O-one, and then say, hey, based on th…”
Disclosure
Fireworks AI will release a reasoning model inspired by OpenAI's o1
“So another announcement is we will also announce a, our next
Declarative system is going to be appear as a model that has extremely high quality, and this model is inspired by O-one announcement from OpenAI.
You should see that by the time we announce this o…”
Insight
Hylak: OpenAI o1 struggles to match personal tone and writing styles
“I think that I've had a very hard time getting it to actually write stuff. I know that I've heard of people using it for writing where it's like processing diffs, more like providing critiques or feedback, but at least for myself, I haven't found a good way to…”
Opinion
McAteer: OpenAI o1 is the first model that grows more impressive over time
“O-one is actually the first model where I'm getting more impressed by it the more I use it. So, like, when ChatGPT first came out, right, I think it was GPT-III. And at first it seemed like, oh wow, this is amazing, it can actually create text that sounds like…”
Insight
McAteer: Asking o1 for full code files instead of diffs is ineffective
“At first I was asking it to generate full files of code, like as I'm making changes. That can get a little bit confusing, and then if you're trying to use like a coding agent to bring it over into your projects. So I feel like that's not the best format for co…”
Insight
Hylak: Hidden reasoning tokens create an information asymmetry between OpenAI and developers
“I think that what makes a one even trickier than other models is that there is actually an asymmetric miss to how well open AI understands the model and how well we, for example, the fact that like reasoning tokens are hidden, right? So there's all this stuff …”
Opinion
Hylak: OpenAI o1 is OpenAI's most capable yet hardest model to use
“We're finding that like, oh, one is the most capable model. I think that opening eye has made. And it's also, I think the hardest to use.”
Assertion Not checkable as stated
Lambert: Early DeepSeek and Qwen reasoning models are substantially narrower than o1
“And I think that these models are really substantially narrower than these full O-one models from OpenAI. So OpenAI is, if you use O-one, you can do it for a lot more tasks. If you use, like I was using the DeepSeq model, and it's supposed to be for math or co…”