Reasoning Tokens
topic on 2 shows · 2 statements across 2 episodes
2 statements about Reasoning Tokens, every show
Huang: Inference and reasoning token generation are growing exponentially
“In fact, that token generation rate for inference, especially reasoning tokens are growing so fast. Several exponentials at the same time, it seems.”
Sohmers: Reasoning models shift inference workloads to 100 output tokens per input
“If you go back a year from today in July of last year, the ratios of like input to output for LLMs were very, very heavily on, on inputs where you could be doing, you know, 10, 10, 15 to one ratio of input to output. But that has completely flipped and it's ob…”