Benedikt Jenik describes using partial differential equations directly as dense self-improvement signals in model training.
Assertion Not checkable as stated
Jenik: Accelerated Understanding achieved 5-trillion token inference context length
“Like we're able to train up to a trillion context input. We're able to train With like, even inputs, outputs, both trillion context length, we are able to do inference at five trillion contexts.”
Assertion Not checkable as stated
Jenik: Accelerated Understanding has trained physics models up to 1-trillion parameters
“We've trained, done hundreds of training runs. We've trained up to a trillion parameter models. So we've really shown the stuff to take off.”
Insight
Jenik: Broadening PDE training domains solves the physical sim-to-real gap
“One interesting thing with PDEs is we are actually fairly confident, like the math is known, that when you solve a PDE correctly, you're doing the physics correctly. Like, obviously, you still need to make sure that you're representing the task that you're try…”
Assertion Not checkable as stated
Jenik: A single 5-trillion-token model run generated 22 terabytes of output
“Like, for example, our five trillion run that we did, the outputs were 22 terabytes, and you want that kind of stuff in accelerator memory.”
Assertion Not checkable as stated
Jenik: Traditional chip design loses performance by separating digital logic and physics
“Especially when you look at the chip design itself, it was much more a, let's start in the digital, let's freeze the digital in, let's send it through some physics for a one time check, like the PDK dictates, I have to have The following feature, otherwise TSN…”
Assertion Not checkable as stated
Jenik: Accelerated Understanding's medium-sized physics models run inference on Mac Studios
“We have done inference on, like, we can do small models with smaller context, or even medium big models with smaller context fit on a MacBook or Mac Studio.”