Adam Optimizer
topic on 1 show · 2 statements across 2 episodes
2 statements about Adam Optimizer, every show
Bakouch: DeepSeek-V3 uses the same Adam optimizer parameters as Llama 2
“And for example, a good a good way to view that is that DeepSeq rig three is still using the same Adam parameter than Lama two.”
First-time Mojo developers built a GPU training system in one day
“The winning team for the hackathon took that four person team. And one day they had not used Mojo before they hadn't programmed GPUs before. And they had built a training system. They wrote an atom optimizer, a bunch of training kernels. They built a simple ba…”