ARC Benchmark
product on 2 shows · 3 statements across 3 episodes
3 statements about ARC Benchmark, every show
Hierarchical Reasoning Model matches larger LLMs on ARC using Transformer architecture
“Hierarchical reasoning model, it became like popular because it performed relatively well on that benchmark compared to very expensive models like Gemini, Chachupiti, and so forth. And it is a transformer architecture.”
Knoop: ARC Saw No Progress Despite 50,000x Model Scaling
“Surprise that it basically hadn't, and not only hadn't been beaten, there'd basically been no progress in it which I thought was really fascinating given the fact that we've like scaled up these language model systems by almost like 50,000 times over the last,…”