Insight certainty 4/5 debate potential 3/5

Editing isolated facts in MLP weights fails to update multi-hop downstream reasoning

Ali Taha · Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten · Aug 3, 2026 · at 1:40:11

Baseten engineer Ali Taha explains why modifying model weights directly is inadequate for continual learning in language models.

0:00 / 0:22exact quote · 22.1s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“So if you know that the best university in the world is Waterloo, then the answer should be Waterloo. But if I wasn't just one-shotting the question and I was to ask it to like, use its knowledge to think and then give me a second answer, or like, should I hire an intern from Waterloo or MIT? It'd be like, oh yeah, both are good. I literally, I just edited in your knowledge base that Waterloo is the best. Why didn't you use that to do reasoning? So that's the fundamental problem with trying to change a fact in an MLP within the way.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Ali Taha

Opinion
AI ASIC startups are doomed as NVIDIA GPUs become domain-specialized
“How can you look at this trend and then still be bullish on companies that are coming up with Asics for AI.”
Ali Taha Aug 3, 2026 ▶ 1:04:48 Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
Disclosure
GLM-5.2 autonomously wrote and guided production GPU kernels for Baseten's inference engine
“Some of the GPU kernels that were on GLM-Five-two within our inference engine is written by GLM-Five-two. And the trace and the kernels were guided by GLM-Five-two as the driver.”
Ali Taha Aug 3, 2026 ▶ 1:34:09 Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
Insight
Quantizing more layers can actually improve model fidelity via error cancellation
“It is possible that the model in which I quantized more information is going to perform better because the quantization errors have canceled out. And so what Joshua showed in his mathematical proof where he had like a verifier in is that you can predict which …”
Ali Taha Aug 3, 2026 ▶ 31:54 Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
Insight
Optimal inference parallelism cannot be mathematically calculated; it must be auto-tuned
“And with training, it's more of like a math, like you can run the math and see the flops and maximize it. With inference, it's more of like an auto-tuning, like GPU kernel auto-tuning... You shadow the same traffic, like real traffic, and you just see which co…”
Ali Taha Aug 3, 2026 ▶ 56:06 Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
Assertion Not checkable as stated
Companies developing fused mega kernels rarely run them in production
“Even the companies that have worked at or people that have spoken to who work at companies that do fused mega kernels, they, Very, very often don't end up running those in production because the TRTL and modular kernels that will launch are faster because you …”
Ali Taha Aug 3, 2026 ▶ 58:38 Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
Prediction Not checkable as stated
NVIDIA's Vera Rubin GPU architecture effectively renders mega kernel research obsolete
“The GPU is, is It's designed in such a way that it basically kills megakernels. You don't need to use megakernels that much anymore. So it seems like that entire research field goes into, like, won't be continued, but yeah.”
Ali Taha Aug 3, 2026 ▶ 59:20 Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.