Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q I guess for V two, what, what caused the, so one was clearly reasoning, two, a benchmark doesn't care how you solve it. I guess embedded in what you said, like, were people using code gen to then solve?
A That's right. So not, not necessarily code gen, uh, per se, but, uh, the Frontier Labs has been targeting ARC V two and, uh, the progress you saw on ARC V two is actually a result. Uh, this very, very large scale targeting. So what you can do to solve RG-II is you ask your reasoning model to make more tasks like those in the benchmark. Uh, and then you try to solve them using, let's say, let's say program induction, for instance, uh, uh, still using your reasoning model. Then you verify the solution. Again, it's very viable. So you can, you can trust, uh, the answer. Um, and then you fine tune the model on the successful reasoning chains. And then you keep repeating, like, you generate new tasks, you solve them, you verify the solution, you fine tune the model on the reasoning chains, and, um, you can keep doing this millions of times, right? Like, you just need to spend more money.
AI assessment note: “not necessarily code gen, per se, but, the Frontier Labs has been targeting ARC V two”
Answered raw tape
D 4 · C 5 · P 4 · Cm 4 4.30
Q out. I guess it becomes more nebulous when you go a couple degrees off where there are fields that are not naturally formally verified and you need to come up with a, again, with some sort of a function. To come up with that reward that makes it verifiable with very fuzzy things like, let's say English language and composing the perfect essay. How do you make that formally verifiable?
A Yeah, yeah, absolutely. I mean, writing SS is, you know, the typical example of a domain that's not verifiable. And so what you're going to see is that progress of reasoning models and base LLMs on this type of, of, of domain is, is, you know, it's going to be very slow because the stack we're using, like the LLM stack is very, very reliant on its trained data. It's basically just operationalizing the trained data. And for writing SS, the trained data is coming from Uh, human experts, like annotating, uh, answers, and that's costly. So you're going to see this very, very slow progress. Maybe, maybe it's even going to stall. But for any, any very favorable domain, like take code for instance, which was the big unlock is, uh, when, uh, when people started creating this code-based like training environment, uh, for, for post-training. Uh, where the, the, the reward signal, the verification signal is provided by things like, uh, unit tests and so on. And so that means that, uh, the model was not just working from human provider annotations. It was actually trying some things, uh, verifying the answer and, uh, and generating a lot, lot more string data in the process, a much denser coverage of the problem space. And not just coverage in terms of like, is, is the answer right or wrong? But also starting to build models of the execution traces, right, so that the models could start in…
AI assessment note: “writing SS is, you know, the typical example of a domain that's not verifiable.”