May 31, 2024 · 1h 12m · latent-space

How to train a Million Context LLM — with Mark Huang of Gradient.ai

Mark Huang · 52m spoken Shawn Wang · 8m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Mark Huang, co-founder of Gradient.ai, joins the Latent Space podcast to break down how his team extended Llama-3 to a one-million token context window, discussing RoPE theta scaling, Ring Attention, data curation, and the future of enterprise AI agents.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 13.1% of the talking time here. How this is scored →

The hosts as informed peer 4.9 Guest teaching 4.4 Guest disagreement 1.2 The hosts pushing back 1.4
05100:0015:0030:0045:001:00:003:04–5:51 · The hosts as informed peer 3/10 Understanding Gradient's Platform and Defining Minimum Viable Agents Alessio and Swyx ask foundational questions about Gradient's agentic platform architecture and how to define a minimum viable agent. Mark explains the transition from RPA workflows to non-deterministic probabilistic pipelines.5:52–9:21 · The hosts as informed peer 4/10 Enterprise Tooling Bottlenecks and Out-of-Domain Generalization Swyx frames the distinction between traditional ML and out-of-domain AI generalization. Mark shares the founding story of Gradient and the frustrations with brittle enterprise internal tooling.9:22–14:34 · The hosts as informed peer 4/10 Motivation Behind 1M Context Llama-3 and Crusoe Compute Swyx prompts Mark on why Gradient chose context extension on Llama-3 and asks for clarification on Crusoe Compute's infrastructure. Mark explains the motivation and partnership details.14:34–22:48 · The hosts as informed peer 4/10 Curriculum Learning, RoPE Scaling, and Base Theta Parameters Alessio asks why long context isn't standard and jokes about knowing what theta is. Mark thoroughly educates the hosts on curriculum learning, positional interpolation, RoPE scaling, and base theta parameter mechanics.22:48–28:00 · The hosts as informed peer 7/10 Evaluating Long-Context Techniques and Ring Attention Implementations Swyx demonstrates deep domain familiarity by bringing up ALiBi, YaRN, Pose, and Ring Attention implementations like Zhang Peiyuan's EasyContext repo. Mark details their GPU cluster topology and practical implementation choices.28:00–33:42 · The hosts as informed peer 4/10 Dataset Curation Strategy and Catastrophic Forgetting Prevention Alessio asks about dataset curation and whether models must know the underlying distribution. Mark breaks down their multi-stage continual pretraining with SlimPajama and UltraChat to avoid catastrophic forgetting.33:42–42:41 · The hosts as informed peer 6/10 Multi-Stage Data Mixing, Synthetic Pipelines, and LoRA Alchemy Swyx suggests multi-stage loss adjustment to prevent overfitting, which Mark qualifies as theoretically hard, pointing to DoReMi and Gemini 1.5. They also dive into LoRA merging as singular value decomposition.42:41–49:06 · The hosts as informed peer 6/10 Long-Context Benchmarks: Beyond Needle in a Haystack to 4M Scaling Swyx lists benchmark suites like Ruler, InfiniteBench, and ZeroScrolls, and asks about degradation when scaling to 4M context. Mark explains multi-needle multi-query evaluation and floating point precision issues at extreme theta values.49:06–1:02:26 · The hosts as informed peer 6/10 Session State Management, Multi-Shot Prompting, and Early Multimodal Fusion Swyx challenges whether pushing context length is fighting the last war when the frontier is shifting to native multimodality, referencing GPT-4o and Chameleon. Mark pushes back by arguing context window is critical for state tracking and video framing.1:02:27–1:11:40 · The hosts as informed peer 5/10 Research Routines, Perplexity Signals, and Future Directions Swyx digs into Mark's practical signal detection for filtering AI papers on Twitter and Discord, and asks for exact perplexity targets. Mark details early perplexity drop heuristics and highlights DeepSeek's Multi-head Latent Attention.3:04–5:51 · Guest teaching 4/10 Understanding Gradient's Platform and Defining Minimum Viable Agents Alessio and Swyx ask foundational questions about Gradient's agentic platform architecture and how to define a minimum viable agent. Mark explains the transition from RPA workflows to non-deterministic probabilistic pipelines.5:52–9:21 · Guest teaching 3/10 Enterprise Tooling Bottlenecks and Out-of-Domain Generalization Swyx frames the distinction between traditional ML and out-of-domain AI generalization. Mark shares the founding story of Gradient and the frustrations with brittle enterprise internal tooling.9:22–14:34 · Guest teaching 4/10 Motivation Behind 1M Context Llama-3 and Crusoe Compute Swyx prompts Mark on why Gradient chose context extension on Llama-3 and asks for clarification on Crusoe Compute's infrastructure. Mark explains the motivation and partnership details.14:34–22:48 · Guest teaching 7/10 Curriculum Learning, RoPE Scaling, and Base Theta Parameters Alessio asks why long context isn't standard and jokes about knowing what theta is. Mark thoroughly educates the hosts on curriculum learning, positional interpolation, RoPE scaling, and base theta parameter mechanics.22:48–28:00 · Guest teaching 4/10 Evaluating Long-Context Techniques and Ring Attention Implementations Swyx demonstrates deep domain familiarity by bringing up ALiBi, YaRN, Pose, and Ring Attention implementations like Zhang Peiyuan's EasyContext repo. Mark details their GPU cluster topology and practical implementation choices.28:00–33:42 · Guest teaching 5/10 Dataset Curation Strategy and Catastrophic Forgetting Prevention Alessio asks about dataset curation and whether models must know the underlying distribution. Mark breaks down their multi-stage continual pretraining with SlimPajama and UltraChat to avoid catastrophic forgetting.33:42–42:41 · Guest teaching 5/10 Multi-Stage Data Mixing, Synthetic Pipelines, and LoRA Alchemy Swyx suggests multi-stage loss adjustment to prevent overfitting, which Mark qualifies as theoretically hard, pointing to DoReMi and Gemini 1.5. They also dive into LoRA merging as singular value decomposition.42:41–49:06 · Guest teaching 4/10 Long-Context Benchmarks: Beyond Needle in a Haystack to 4M Scaling Swyx lists benchmark suites like Ruler, InfiniteBench, and ZeroScrolls, and asks about degradation when scaling to 4M context. Mark explains multi-needle multi-query evaluation and floating point precision issues at extreme theta values.49:06–1:02:26 · Guest teaching 4/10 Session State Management, Multi-Shot Prompting, and Early Multimodal Fusion Swyx challenges whether pushing context length is fighting the last war when the frontier is shifting to native multimodality, referencing GPT-4o and Chameleon. Mark pushes back by arguing context window is critical for state tracking and video framing.1:02:27–1:11:40 · Guest teaching 4/10 Research Routines, Perplexity Signals, and Future Directions Swyx digs into Mark's practical signal detection for filtering AI papers on Twitter and Discord, and asks for exact perplexity targets. Mark details early perplexity drop heuristics and highlights DeepSeek's Multi-head Latent Attention.3:04–5:51 · Guest disagreement 1/10 Understanding Gradient's Platform and Defining Minimum Viable Agents Alessio and Swyx ask foundational questions about Gradient's agentic platform architecture and how to define a minimum viable agent. Mark explains the transition from RPA workflows to non-deterministic probabilistic pipelines.5:52–9:21 · Guest disagreement 1/10 Enterprise Tooling Bottlenecks and Out-of-Domain Generalization Swyx frames the distinction between traditional ML and out-of-domain AI generalization. Mark shares the founding story of Gradient and the frustrations with brittle enterprise internal tooling.9:22–14:34 · Guest disagreement 1/10 Motivation Behind 1M Context Llama-3 and Crusoe Compute Swyx prompts Mark on why Gradient chose context extension on Llama-3 and asks for clarification on Crusoe Compute's infrastructure. Mark explains the motivation and partnership details.14:34–22:48 · Guest disagreement 1/10 Curriculum Learning, RoPE Scaling, and Base Theta Parameters Alessio asks why long context isn't standard and jokes about knowing what theta is. Mark thoroughly educates the hosts on curriculum learning, positional interpolation, RoPE scaling, and base theta parameter mechanics.22:48–28:00 · Guest disagreement 1/10 Evaluating Long-Context Techniques and Ring Attention Implementations Swyx demonstrates deep domain familiarity by bringing up ALiBi, YaRN, Pose, and Ring Attention implementations like Zhang Peiyuan's EasyContext repo. Mark details their GPU cluster topology and practical implementation choices.28:00–33:42 · Guest disagreement 1/10 Dataset Curation Strategy and Catastrophic Forgetting Prevention Alessio asks about dataset curation and whether models must know the underlying distribution. Mark breaks down their multi-stage continual pretraining with SlimPajama and UltraChat to avoid catastrophic forgetting.33:42–42:41 · Guest disagreement 2/10 Multi-Stage Data Mixing, Synthetic Pipelines, and LoRA Alchemy Swyx suggests multi-stage loss adjustment to prevent overfitting, which Mark qualifies as theoretically hard, pointing to DoReMi and Gemini 1.5. They also dive into LoRA merging as singular value decomposition.42:41–49:06 · Guest disagreement 1/10 Long-Context Benchmarks: Beyond Needle in a Haystack to 4M Scaling Swyx lists benchmark suites like Ruler, InfiniteBench, and ZeroScrolls, and asks about degradation when scaling to 4M context. Mark explains multi-needle multi-query evaluation and floating point precision issues at extreme theta values.49:06–1:02:26 · Guest disagreement 2/10 Session State Management, Multi-Shot Prompting, and Early Multimodal Fusion Swyx challenges whether pushing context length is fighting the last war when the frontier is shifting to native multimodality, referencing GPT-4o and Chameleon. Mark pushes back by arguing context window is critical for state tracking and video framing.1:02:27–1:11:40 · Guest disagreement 1/10 Research Routines, Perplexity Signals, and Future Directions Swyx digs into Mark's practical signal detection for filtering AI papers on Twitter and Discord, and asks for exact perplexity targets. Mark details early perplexity drop heuristics and highlights DeepSeek's Multi-head Latent Attention.3:04–5:51 · The hosts pushing back 1/10 Understanding Gradient's Platform and Defining Minimum Viable Agents Alessio and Swyx ask foundational questions about Gradient's agentic platform architecture and how to define a minimum viable agent. Mark explains the transition from RPA workflows to non-deterministic probabilistic pipelines.5:52–9:21 · The hosts pushing back 1/10 Enterprise Tooling Bottlenecks and Out-of-Domain Generalization Swyx frames the distinction between traditional ML and out-of-domain AI generalization. Mark shares the founding story of Gradient and the frustrations with brittle enterprise internal tooling.9:22–14:34 · The hosts pushing back 1/10 Motivation Behind 1M Context Llama-3 and Crusoe Compute Swyx prompts Mark on why Gradient chose context extension on Llama-3 and asks for clarification on Crusoe Compute's infrastructure. Mark explains the motivation and partnership details.14:34–22:48 · The hosts pushing back 1/10 Curriculum Learning, RoPE Scaling, and Base Theta Parameters Alessio asks why long context isn't standard and jokes about knowing what theta is. Mark thoroughly educates the hosts on curriculum learning, positional interpolation, RoPE scaling, and base theta parameter mechanics.22:48–28:00 · The hosts pushing back 2/10 Evaluating Long-Context Techniques and Ring Attention Implementations Swyx demonstrates deep domain familiarity by bringing up ALiBi, YaRN, Pose, and Ring Attention implementations like Zhang Peiyuan's EasyContext repo. Mark details their GPU cluster topology and practical implementation choices.28:00–33:42 · The hosts pushing back 1/10 Dataset Curation Strategy and Catastrophic Forgetting Prevention Alessio asks about dataset curation and whether models must know the underlying distribution. Mark breaks down their multi-stage continual pretraining with SlimPajama and UltraChat to avoid catastrophic forgetting.33:42–42:41 · The hosts pushing back 2/10 Multi-Stage Data Mixing, Synthetic Pipelines, and LoRA Alchemy Swyx suggests multi-stage loss adjustment to prevent overfitting, which Mark qualifies as theoretically hard, pointing to DoReMi and Gemini 1.5. They also dive into LoRA merging as singular value decomposition.42:41–49:06 · The hosts pushing back 1/10 Long-Context Benchmarks: Beyond Needle in a Haystack to 4M Scaling Swyx lists benchmark suites like Ruler, InfiniteBench, and ZeroScrolls, and asks about degradation when scaling to 4M context. Mark explains multi-needle multi-query evaluation and floating point precision issues at extreme theta values.49:06–1:02:26 · The hosts pushing back 3/10 Session State Management, Multi-Shot Prompting, and Early Multimodal Fusion Swyx challenges whether pushing context length is fighting the last war when the frontier is shifting to native multimodality, referencing GPT-4o and Chameleon. Mark pushes back by arguing context window is critical for state tracking and video framing.1:02:27–1:11:40 · The hosts pushing back 1/10 Research Routines, Perplexity Signals, and Future Directions Swyx digs into Mark's practical signal detection for filtering AI papers on Twitter and Discord, and asks for exact perplexity targets. Mark details early perplexity drop heuristics and highlights DeepSeek's Multi-head Latent Attention.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 25.6% · guest 74.4%0:00 · the hosts 25.6% · guest 74.4%3:00 · the hosts 6.5% · guest 93.5%3:00 · the hosts 6.5% · guest 93.5%6:00 · the hosts 11.2% · guest 88.8%6:00 · the hosts 11.2% · guest 88.8%9:00 · the hosts 29.5% · guest 70.5%9:00 · the hosts 29.5% · guest 70.5%12:00 · the hosts 11.3% · guest 88.7%12:00 · the hosts 11.3% · guest 88.7%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 15.6% · guest 84.4%21:00 · the hosts 15.6% · guest 84.4%24:00 · the hosts 21.7% · guest 78.3%24:00 · the hosts 21.7% · guest 78.3%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 29.2% · guest 70.8%33:00 · the hosts 29.2% · guest 70.8%36:00 · the hosts 24.7% · guest 75.3%36:00 · the hosts 24.7% · guest 75.3%39:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%42:00 · the hosts 23.6% · guest 76.4%42:00 · the hosts 23.6% · guest 76.4%45:00 · the hosts 11.8% · guest 88.2%45:00 · the hosts 11.8% · guest 88.2%48:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%57:00 · the hosts 36.8% · guest 63.2%57:00 · the hosts 36.8% · guest 63.2%1:00:00 · the hosts 41.1% · guest 58.9%1:00:00 · the hosts 41.1% · guest 58.9%1:03:00 · the hosts 7.2% · guest 92.8%1:03:00 · the hosts 7.2% · guest 92.8%1:06:00 · the hosts 17% · guest 83%1:06:00 · the hosts 17% · guest 83%1:09:00 · the hosts 0.4% · guest 99.6%1:09:00 · the hosts 0.4% · guest 99.6%1:12:00 · the hosts 0% · guest 0%1:12:00 · the hosts 0% · guest 0%
Sharpest disagreement ▶ 49:40 Pushing Back on Pure Context Size as a Metric

Mark counters the assumption that simply increasing raw context length solves problems, asserting that iterative agentic workflows like AlphaCodium often outperform brute force context stuffing.

Hardest push from the hosts ▶ 58:32 Host Challenges Frontier Focus

Swyx directly pushes back against chasing 10x context length as fighting the last war when the industry focus has rapidly pivoted toward early fusion multimodality and GPT-4o style models.

Biggest teaching moment ▶ 19:45 Deconstructing RoPE Theta and Interpolation

Mark delivers an in-depth technical explanation of positional interpolation versus extrapolation and how the base theta parameter governs rotational embedding frequencies.

The host holds their own ▶ 25:06 Swyx Drilling into Ring Attention and Repos

Swyx demonstrates sharp technical literacy by citing Zhang Peiyuan's EasyContext and Lucidrains implementations, directly engaging on GPU cluster attention mechanics.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Understanding Gradient's Platform and Defining Minimum Viable Agents 3411 Alessio and Swyx ask foundational questions about Gradient's agentic platform architecture and how to define a minimum viable agent. Mark explains the transition from RPA workflows to non-deterministic probabilistic pipelines.
Enterprise Tooling Bottlenecks and Out-of-Domain Generalization 4311 Swyx frames the distinction between traditional ML and out-of-domain AI generalization. Mark shares the founding story of Gradient and the frustrations with brittle enterprise internal tooling.
Motivation Behind 1M Context Llama-3 and Crusoe Compute 4411 Swyx prompts Mark on why Gradient chose context extension on Llama-3 and asks for clarification on Crusoe Compute's infrastructure. Mark explains the motivation and partnership details.
Curriculum Learning, RoPE Scaling, and Base Theta Parameters 4711 Alessio asks why long context isn't standard and jokes about knowing what theta is. Mark thoroughly educates the hosts on curriculum learning, positional interpolation, RoPE scaling, and base theta parameter mechanics.
Evaluating Long-Context Techniques and Ring Attention Implementations 7412 Swyx demonstrates deep domain familiarity by bringing up ALiBi, YaRN, Pose, and Ring Attention implementations like Zhang Peiyuan's EasyContext repo. Mark details their GPU cluster topology and practical implementation choices.
Dataset Curation Strategy and Catastrophic Forgetting Prevention 4511 Alessio asks about dataset curation and whether models must know the underlying distribution. Mark breaks down their multi-stage continual pretraining with SlimPajama and UltraChat to avoid catastrophic forgetting.
Multi-Stage Data Mixing, Synthetic Pipelines, and LoRA Alchemy 6522 Swyx suggests multi-stage loss adjustment to prevent overfitting, which Mark qualifies as theoretically hard, pointing to DoReMi and Gemini 1.5. They also dive into LoRA merging as singular value decomposition.
Long-Context Benchmarks: Beyond Needle in a Haystack to 4M Scaling 6411 Swyx lists benchmark suites like Ruler, InfiniteBench, and ZeroScrolls, and asks about degradation when scaling to 4M context. Mark explains multi-needle multi-query evaluation and floating point precision issues at extreme theta values.
Session State Management, Multi-Shot Prompting, and Early Multimodal Fusion 6423 Swyx challenges whether pushing context length is fighting the last war when the frontier is shifting to native multimodality, referencing GPT-4o and Chameleon. Mark pushes back by arguing context window is critical for state tracking and video framing.
Research Routines, Perplexity Signals, and Future Directions 5411 Swyx digs into Mark's practical signal detection for filtering AI papers on Twitter and Discord, and asks for exact perplexity targets. Mark details early perplexity drop heuristics and highlights DeepSeek's Multi-head Latent Attention.

Statements from this episode (23)

Opinion
Huang: AI talent race mirrors past quant trading wars
“Now we intersect again when it kind of feels like more or less the same, right? Like the AI wars, the trading wars back in the day too, to a certain extent, and the grab for talent.”
Mark Huang May 31, 2024 ▶ 0:55
Disclosure
Huang: Gradient aims to transition enterprise RPA to autonomous agents
“Well, quite simply, like, gradient, we're a full stack AI platform, and what we really want to do is we want to enable all of the, you know, RPA workloads or the codified automation workloads that existed in the enterprise before. We really want to enable peop…”
Mark Huang May 31, 2024 ▶ 3:51
Insight
Huang: True AI agents require measurable probability improvements per node
“It's like on each stage of the node, you're gonna have to see a marginal improvement in the probability of success for that particular workload because of non-determinism.”
Mark Huang May 31, 2024 ▶ 5:16
Opinion
Huang: Google's internal AI tooling was far superior to competitors
“Google was using AI for systems before everybody else too, right? They invented a transformer, and their internal set of tooling was just so far superior to everything else. Like, it's really hard for people to go back after seeing that.”
Mark Huang May 31, 2024 ▶ 7:41
Insight
Huang: RAG versus fine-tuning is fundamentally just meta-learning
“And like, at the end of the day, it's just all meta-learning, right? Like, all we want is, like, the best meta learning workflow or meta learning setup possible to be able to adapt the model to do anything.”
Mark Huang May 31, 2024 ▶ 11:03
Disclosure
Gradient trained 1M context Llama-3 on Crusoe's Nvidia L40 GPUs
“It just made it really easy for us to, like, scale up with their L-Forties, and those are the specific GPU instances we used, and coordinating that effort with them to get, you know, that dedicated cluster first to do the project”
Mark Huang May 31, 2024 ▶ 13:47
Assertion Supported
Huang: Curriculum context expansion outperforms full-length training from scratch
“If you train a model on a shorter context and you progressively increase that context to, like, You know, the final limit that you have, like, 32 K is usually the limit of Lama two was that long. It actually performs better than if you try to train 32 K the …”
Mark Huang May 31, 2024 ▶ 15:39
Assertion Supported
Huang: Context scaling requires positional interpolation rather than extrapolation
“There's it's super confusing, but it's like, there's extrapolate positional extrapolation, and then there's interpolation. You want interpolation. It's been shown that just pure extrapolation makes the model a lot worse, and it's harder to attend to stuff, whe…”
Mark Huang May 31, 2024 ▶ 21:04
Assertion Not checkable as stated
Huang: Modern LLMs have shifted from ALiBi to RoPE scaling
“Some of the newer architectures don't actually employ it a lot. I think the last architecture that actually really employed it was the Mosaic MPT model class, and then almost all the models these days are all rope scaling, and then effectively you can use yarn…”
Mark Huang May 31, 2024 ▶ 23:13
Assertion Open · timeframe May 2025
Huang: PoSE breaks down on needle-in-a-haystack at 500k tokens
“It does start to break down a little bit more on the longer, longer context. So, like, 500,000 to a million it appeared that it doesn't hold as well specifically for, like, needle in the haystack.”
Mark Huang May 31, 2024 ▶ 24:11
Assertion Not checkable as stated
Huang: Berkeley's JAX Ring Attention does not work well on GPUs
“The Jaxx implementation just does not work on, on GPUs very well. Like, any naive setup that you do, like, it just won't run out of the box very easily”
Mark Huang May 31, 2024 ▶ 26:02
Assertion Not checkable as stated
Huang: EasyContext was the first functional PyTorch Ring Attention implementation
“Easy context was the first PyTorch implementation that applied it with native libraries that worked pretty well. And then we adapted it ourselves in order to configure it for our cluster network topology.”
Mark Huang May 31, 2024 ▶ 27:08
Insight
Huang: Placing answers at the end of long contexts breaks attention
“You could create, like, a long context dataset where, like, every single time the last 200 tokens can answer the entire question, and that's never gonna make the model attend to anything.”
Mark Huang May 31, 2024 ▶ 30:34
Insight
Huang: Adding one billion tokens cannot teach trillion-token models new knowledge
“All models these days are now double-digit trillions, right? So it's kind of a drop in the bucket if you really think I can just put, you know, a billion tokens in there, and I actually think that the model's gonna truly learn new Information.”
Mark Huang May 31, 2024 ▶ 31:49
Assertion Supported
Huang: Training CodeLlama on Llama 2 caused catastrophic language forgetting
“We do have historical precedent where CodeLlama was, you know, trained further from the original CodeLlama was trained further from Lama II, and it just lost, All its language capabilities, basically, right?”
Mark Huang May 31, 2024 ▶ 33:03
Insight
Huang: LoRA merging succeeds on style but fails on complex capabilities
“Like, I will not lie to say I'm really surprised how effective it is sometimes, but I do notice that for more complex abilities other than, like, more stylistic stuff, it does, it kind of falls through, because maybe it's, it requires a much deeper path in the…”
Mark Huang May 31, 2024 ▶ 40:21
Opinion
Huang: Model merging is polluting open LLM leaderboards
“That is extremely interesting from the developer community, and I want to see more of it except it is, to a certain extent, kind of polluting the leaderboards these days, because it's so targeted, and like, now you can kind of game the metric by just finding a…”
Mark Huang May 31, 2024 ▶ 41:27
Insight
Huang: Needle in a Haystack is a primitive LLM benchmark prerequisite
“I think needle in a haystack is definitely, like, the standard for presenting the work in a way that people can understand and also proving out. I would say, like, I view it as, like, a primitive. That you have to pass in order to give the model any shot of do…”
Mark Huang May 31, 2024 ▶ 42:55
Opinion
Huang: RULER benchmark is more comprehensive than Gemini's multi-needle test
“I would even argue is more comprehensive than the benchmark that, that Gemini released for their, like, multi-needle in the haystack.”
Mark Huang May 31, 2024 ▶ 44:18
Insight
Huang: LLM session state management will require huge context windows
“Making the model track state and have state management over time is really, really hard. And it's an incredibly hard evaluation that will probably only really work when you have a huge context.”
Mark Huang May 31, 2024 ▶ 51:37
Opinion
Huang: Multimodal inputs like video will drive long-context LLM demand
“Multimodality is, in my opinion, going to be, it's going to be pivotal for long context. Just because, like, videos when you're getting into the frames per second and you're getting into lots of images and, like, things that are a lot more, like, embodied You …”
Mark Huang May 31, 2024 ▶ 56:42
Insight
Huang: Correct RoPE theta scaling causes immediate early perplexity drops
“Specifically when the one trick you should pay attention to is you know that your context length and theta scaling is working right if the early steps in the perplexity go straight down. So like when it wasn't correct, it would oscillate a lot in the beginning…”
Mark Huang May 31, 2024 ▶ 1:06:03
Opinion
Huang: DeepSeek's Multi-Head Latent Attention is novel and underrated
“Underrated specific instance would be, like, the deep seek paper. I'd never seen it before, but, like, the multi-head latent attention, like, that was really unexpected to me because, like, I thought I'd seen every Not every type, obviously, but, like, every w…”
Mark Huang May 31, 2024 ▶ 1:08:55
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.