Token Generation
topic on 2 shows · 2 statements across 2 episodes
2 statements about Token Generation, every show
Ross: Splitting LLM pre-fill and generation across different chips is a mistake
“What we realized was, and this is what most people get wrong when they're trying to do this themselves. They'll, they'll take the reading of tokens or what's called pre-fill, and they'll do that on one piece of hardware, and then they'll put the generation of …”