Agentic Tasks
topic on 2 shows · 2 statements across 2 episodes
2 statements about Agentic Tasks, every show
General LLM benchmark rankings do not translate to agent performance
“What we have understood from our previous experiments is that it's not necessary that what you see as top models, In let's say LM arena or other specific evaluation or general evaluations, they might not be also the same ranking for other tasks like agentic ta…”