Visual Question Answering
topic on 1 show · 2 statements across 1 episodes
2 statements about Visual Question Answering, every show
Bordes: Visual QA models saturate around 60% accuracy versus 95% for humans
“Basically many methods actually saturated at like, I don't know, 55 or 60% accuracy. Ah, the simpler or the most complicated were actually in the same ballpark. Whereas human can actually go up to 95”
Visual question answering is harder than captioning because it requires reasoning
“So, what people try to do now is that to move to caption what's called caption generation, which was actually super promising, but actually people realized that actually the machine wasn't that good, to what's called now visual question answering, which is mor…”