May 22, 2024 · 55m · latent-space

LLM Asia Paper Club Survey Round

0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

The LLM Asia Paper Club survey session convenes researchers to present and evaluate recent developments in large language model reasoning, uncertainty estimation, mechanistic interpretability, and speculative decoding. Through technical presentations and Q&A discussions, participants examine how token manipulation, internal activation probing, and novel architectures enhance model transparency and inference efficiency.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 2.0 Guest teaching 0.8 Guest disagreement 0.1 The hosts pushing back 0.3
05100:0015:0030:0045:000:00–9:17 · The hosts as informed peer 0/10 Presentation on "Let's Think Dot by Thought" Paper This is a solo presentation monologue by Jung-sen detailing the 'Let's Think Dot by Thought' paper. In accordance with monologue rules, host-side metrics and combativeness are scored at zero.9:18–12:18 · The hosts as informed peer 3/10 Q&A on Filler Tokens and Positional Bias The host moderates questions from the chat regarding filler tokens and introduces an observation about how spaces affect token computation. Gabriel chimes in collaboratively regarding positional bias, and the host gently clarifies Gabriel's framing.12:22–26:31 · The hosts as informed peer 0/10 Presentation on Estimating LLM Uncertainty with Secondary Models Gabriel delivers a technical presentation on estimating LLM uncertainty via secondary regression models. Because this is an uninterrupted presentation, host metrics and combativeness are zero.26:34–32:06 · The hosts as informed peer 2/10 Q&A on Model Calibration and Uncertainty Estimation The host reads questions from chat and asks whether a well-calibrated model is simply one that performs well on downstream tasks. Gabriel politely corrects this misconception, explaining that calibration refers specifically to output probabilities matching empirical likelihoods.32:14–42:42 · The hosts as informed peer 0/10 Presentation on Mechanistic Interpretability and Monosemanticity Casper presents Anthropic's 'Towards Monosemanticity' paper covering mechanistic interpretability and sparse autoencoders without host interaction.42:43–44:44 · The hosts as informed peer 1/10 Q&A on Sparse Autoencoders and Dictionary Learning The host facilitates Q&A by reading chat inquiries comparing sparse autoencoders to probing classifiers. Casper answers collaboratively and explains the unsupervised nature and reconstruction metrics of SAEs.44:46–53:22 · The hosts as informed peer 6/10 Presentation on Medusa for Efficient Speculative Decoding The host delivers a technical walkthrough of Medusa speculative decoding, breaking down memory bandwidth bottlenecks, additional MLP heads, and training objectives. There is no guest combativeness or schooling.53:31–55:23 · The hosts as informed peer 4/10 Q&A on Medusa Speculative Decoding and Meeting Conclusion The host addresses a chat question regarding beam search computational overhead versus Medusa speculative decoding before concluding the session.0:00–9:17 · Guest teaching 0/10 Presentation on "Let's Think Dot by Thought" Paper This is a solo presentation monologue by Jung-sen detailing the 'Let's Think Dot by Thought' paper. In accordance with monologue rules, host-side metrics and combativeness are scored at zero.9:18–12:18 · Guest teaching 1/10 Q&A on Filler Tokens and Positional Bias The host moderates questions from the chat regarding filler tokens and introduces an observation about how spaces affect token computation. Gabriel chimes in collaboratively regarding positional bias, and the host gently clarifies Gabriel's framing.12:22–26:31 · Guest teaching 0/10 Presentation on Estimating LLM Uncertainty with Secondary Models Gabriel delivers a technical presentation on estimating LLM uncertainty via secondary regression models. Because this is an uninterrupted presentation, host metrics and combativeness are zero.26:34–32:06 · Guest teaching 3/10 Q&A on Model Calibration and Uncertainty Estimation The host reads questions from chat and asks whether a well-calibrated model is simply one that performs well on downstream tasks. Gabriel politely corrects this misconception, explaining that calibration refers specifically to output probabilities matching empirical likelihoods.32:14–42:42 · Guest teaching 0/10 Presentation on Mechanistic Interpretability and Monosemanticity Casper presents Anthropic's 'Towards Monosemanticity' paper covering mechanistic interpretability and sparse autoencoders without host interaction.42:43–44:44 · Guest teaching 2/10 Q&A on Sparse Autoencoders and Dictionary Learning The host facilitates Q&A by reading chat inquiries comparing sparse autoencoders to probing classifiers. Casper answers collaboratively and explains the unsupervised nature and reconstruction metrics of SAEs.44:46–53:22 · Guest teaching 0/10 Presentation on Medusa for Efficient Speculative Decoding The host delivers a technical walkthrough of Medusa speculative decoding, breaking down memory bandwidth bottlenecks, additional MLP heads, and training objectives. There is no guest combativeness or schooling.53:31–55:23 · Guest teaching 0/10 Q&A on Medusa Speculative Decoding and Meeting Conclusion The host addresses a chat question regarding beam search computational overhead versus Medusa speculative decoding before concluding the session.0:00–9:17 · Guest disagreement 0/10 Presentation on "Let's Think Dot by Thought" Paper This is a solo presentation monologue by Jung-sen detailing the 'Let's Think Dot by Thought' paper. In accordance with monologue rules, host-side metrics and combativeness are scored at zero.9:18–12:18 · Guest disagreement 0/10 Q&A on Filler Tokens and Positional Bias The host moderates questions from the chat regarding filler tokens and introduces an observation about how spaces affect token computation. Gabriel chimes in collaboratively regarding positional bias, and the host gently clarifies Gabriel's framing.12:22–26:31 · Guest disagreement 0/10 Presentation on Estimating LLM Uncertainty with Secondary Models Gabriel delivers a technical presentation on estimating LLM uncertainty via secondary regression models. Because this is an uninterrupted presentation, host metrics and combativeness are zero.26:34–32:06 · Guest disagreement 1/10 Q&A on Model Calibration and Uncertainty Estimation The host reads questions from chat and asks whether a well-calibrated model is simply one that performs well on downstream tasks. Gabriel politely corrects this misconception, explaining that calibration refers specifically to output probabilities matching empirical likelihoods.32:14–42:42 · Guest disagreement 0/10 Presentation on Mechanistic Interpretability and Monosemanticity Casper presents Anthropic's 'Towards Monosemanticity' paper covering mechanistic interpretability and sparse autoencoders without host interaction.42:43–44:44 · Guest disagreement 0/10 Q&A on Sparse Autoencoders and Dictionary Learning The host facilitates Q&A by reading chat inquiries comparing sparse autoencoders to probing classifiers. Casper answers collaboratively and explains the unsupervised nature and reconstruction metrics of SAEs.44:46–53:22 · Guest disagreement 0/10 Presentation on Medusa for Efficient Speculative Decoding The host delivers a technical walkthrough of Medusa speculative decoding, breaking down memory bandwidth bottlenecks, additional MLP heads, and training objectives. There is no guest combativeness or schooling.53:31–55:23 · Guest disagreement 0/10 Q&A on Medusa Speculative Decoding and Meeting Conclusion The host addresses a chat question regarding beam search computational overhead versus Medusa speculative decoding before concluding the session.0:00–9:17 · The hosts pushing back 0/10 Presentation on "Let's Think Dot by Thought" Paper This is a solo presentation monologue by Jung-sen detailing the 'Let's Think Dot by Thought' paper. In accordance with monologue rules, host-side metrics and combativeness are scored at zero.9:18–12:18 · The hosts pushing back 1/10 Q&A on Filler Tokens and Positional Bias The host moderates questions from the chat regarding filler tokens and introduces an observation about how spaces affect token computation. Gabriel chimes in collaboratively regarding positional bias, and the host gently clarifies Gabriel's framing.12:22–26:31 · The hosts pushing back 0/10 Presentation on Estimating LLM Uncertainty with Secondary Models Gabriel delivers a technical presentation on estimating LLM uncertainty via secondary regression models. Because this is an uninterrupted presentation, host metrics and combativeness are zero.26:34–32:06 · The hosts pushing back 1/10 Q&A on Model Calibration and Uncertainty Estimation The host reads questions from chat and asks whether a well-calibrated model is simply one that performs well on downstream tasks. Gabriel politely corrects this misconception, explaining that calibration refers specifically to output probabilities matching empirical likelihoods.32:14–42:42 · The hosts pushing back 0/10 Presentation on Mechanistic Interpretability and Monosemanticity Casper presents Anthropic's 'Towards Monosemanticity' paper covering mechanistic interpretability and sparse autoencoders without host interaction.42:43–44:44 · The hosts pushing back 0/10 Q&A on Sparse Autoencoders and Dictionary Learning The host facilitates Q&A by reading chat inquiries comparing sparse autoencoders to probing classifiers. Casper answers collaboratively and explains the unsupervised nature and reconstruction metrics of SAEs.44:46–53:22 · The hosts pushing back 0/10 Presentation on Medusa for Efficient Speculative Decoding The host delivers a technical walkthrough of Medusa speculative decoding, breaking down memory bandwidth bottlenecks, additional MLP heads, and training objectives. There is no guest combativeness or schooling.53:31–55:23 · The hosts pushing back 0/10 Q&A on Medusa Speculative Decoding and Meeting Conclusion The host addresses a chat question regarding beam search computational overhead versus Medusa speculative decoding before concluding the session.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 28:03 Gabriel rejects conflating calibration with performance

Gabriel directly challenges the host's suggested definition that calibration equals task performance, clarifying that returning values between zero and one does not inherently make them true probabilities.

Hardest push from the hosts ▶ 11:30 Host restructures positional bias explanation

The host pushes back on Gabriel's description of positional encoding changes, refining the argument to state that added whitespace alters prompt alignment rather than the positional bias mechanism itself.

Biggest teaching moment ▶ 27:53 Gabriel educates on probability calibration

Gabriel walks the group through statistical calibration using an empirical rain forecast example to educate the room after the host asks a naive question.

The host holds their own ▶ 47:30 Host details Medusa loss functions and architecture

The host showcases deep domain mastery while explaining how auxiliary MLP heads predict future tokens using frozen base models and modified cross-entropy loss.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Presentation on "Let's Think Dot by Thought" Paper 0000 This is a solo presentation monologue by Jung-sen detailing the 'Let's Think Dot by Thought' paper. In accordance with monologue rules, host-side metrics and combativeness are scored at zero.
Q&A on Filler Tokens and Positional Bias 3101 The host moderates questions from the chat regarding filler tokens and introduces an observation about how spaces affect token computation. Gabriel chimes in collaboratively regarding positional bias, and the host gently clarifies Gabriel's framing.
Presentation on Estimating LLM Uncertainty with Secondary Models 0000 Gabriel delivers a technical presentation on estimating LLM uncertainty via secondary regression models. Because this is an uninterrupted presentation, host metrics and combativeness are zero.
Q&A on Model Calibration and Uncertainty Estimation 2311 The host reads questions from chat and asks whether a well-calibrated model is simply one that performs well on downstream tasks. Gabriel politely corrects this misconception, explaining that calibration refers specifically to output probabilities matching empirical likelihoods.
Presentation on Mechanistic Interpretability and Monosemanticity 0000 Casper presents Anthropic's 'Towards Monosemanticity' paper covering mechanistic interpretability and sparse autoencoders without host interaction.
Q&A on Sparse Autoencoders and Dictionary Learning 1200 The host facilitates Q&A by reading chat inquiries comparing sparse autoencoders to probing classifiers. Casper answers collaboratively and explains the unsupervised nature and reconstruction metrics of SAEs.
Presentation on Medusa for Efficient Speculative Decoding 6000 The host delivers a technical walkthrough of Medusa speculative decoding, breaking down memory bandwidth bottlenecks, additional MLP heads, and training objectives. There is no guest combativeness or schooling.
Q&A on Medusa Speculative Decoding and Meeting Conclusion 4000 The host addresses a chat question regarding beam search computational overhead versus Medusa speculative decoding before concluding the session.

Statements from this episode (0)

Nothing in this episode matches those filters. clear them

Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.