Optical Character Recognition
topic on 3 shows · 4 statements across 4 episodes
Latent Space
How I Built This
the a16z Podcast
4 statements about Optical Character Recognition, every show
Zhang: Meta intentionally avoided OCR-heavy images during SAM 3 training data sampling
“In fact, during our data engine, we intentionally do not sample OCR-heavy images.”
von Ahn: Early book scanning software failed on 30% of words
“Computers cannot, or at the time, could not recognize many of the words, about 30% of the words computers could not recognize.”