On Chip Cache
topic on 1 show · 1 statements across 1 episodes
1 statements about On Chip Cache, every show
Sohmers: Matrix-vector multiplication in transformer inference is fundamentally uncacheable
“So the second level of this is that matrix vector multiplication is basically uncacheable. When you're doing transformer inference, matrix A is the weights of your model. And so if you're talking about model weights that are tens of gigabytes, hundreds of giga…”