evaluations for language models
1 statements across 1 episodes · 1 bullish · 0 bearish · 1 people on the record · first statement Jul 9, 2026 by Atin Sanyal · across every show →
Everything said about evaluations for language models, oldest first
Jul 9, 2026 positive
Sanyal: LLM evaluation does not require general-purpose reasoning models
“We feel like if you really focus on the fundamental problem that, hey, evaluations for language models is a very task-specific, constrained problem, and it does not require you to use you know, general purpose reasoning for that, and then there's fine-tuning y…”