May 22, 2024 · 55m · latent-space
LLM Asia Paper Club Survey Round
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
The LLM Asia Paper Club survey session convenes researchers to present and evaluate recent developments in large language model reasoning, uncertainty estimation, mechanistic interpretability, and speculative decoding. Through technical presentations and Q&A discussions, participants examine how token manipulation, internal activation probing, and novel architectures enhance model transparency and inference efficiency.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Gabriel directly challenges the host's suggested definition that calibration equals task performance, clarifying that returning values between zero and one does not inherently make them true probabilities.
Hardest push from the hosts ▶ 11:30 Host restructures positional bias explanationThe host pushes back on Gabriel's description of positional encoding changes, refining the argument to state that added whitespace alters prompt alignment rather than the positional bias mechanism itself.
Biggest teaching moment ▶ 27:53 Gabriel educates on probability calibrationGabriel walks the group through statistical calibration using an empirical rain forecast example to educate the room after the host asks a naive question.
The host holds their own ▶ 47:30 Host details Medusa loss functions and architectureThe host showcases deep domain mastery while explaining how auxiliary MLP heads predict future tokens using frozen base models and modified cross-entropy loss.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Presentation on "Let's Think Dot by Thought" Paper | 0 | 0 | 0 | 0 | This is a solo presentation monologue by Jung-sen detailing the 'Let's Think Dot by Thought' paper. In accordance with monologue rules, host-side metrics and combativeness are scored at zero. | |
| Q&A on Filler Tokens and Positional Bias | 3 | 1 | 0 | 1 | The host moderates questions from the chat regarding filler tokens and introduces an observation about how spaces affect token computation. Gabriel chimes in collaboratively regarding positional bias, and the host gently clarifies Gabriel's framing. | |
| Presentation on Estimating LLM Uncertainty with Secondary Models | 0 | 0 | 0 | 0 | Gabriel delivers a technical presentation on estimating LLM uncertainty via secondary regression models. Because this is an uninterrupted presentation, host metrics and combativeness are zero. | |
| Q&A on Model Calibration and Uncertainty Estimation | 2 | 3 | 1 | 1 | The host reads questions from chat and asks whether a well-calibrated model is simply one that performs well on downstream tasks. Gabriel politely corrects this misconception, explaining that calibration refers specifically to output probabilities matching empirical likelihoods. | |
| Presentation on Mechanistic Interpretability and Monosemanticity | 0 | 0 | 0 | 0 | Casper presents Anthropic's 'Towards Monosemanticity' paper covering mechanistic interpretability and sparse autoencoders without host interaction. | |
| Q&A on Sparse Autoencoders and Dictionary Learning | 1 | 2 | 0 | 0 | The host facilitates Q&A by reading chat inquiries comparing sparse autoencoders to probing classifiers. Casper answers collaboratively and explains the unsupervised nature and reconstruction metrics of SAEs. | |
| Presentation on Medusa for Efficient Speculative Decoding | 6 | 0 | 0 | 0 | The host delivers a technical walkthrough of Medusa speculative decoding, breaking down memory bandwidth bottlenecks, additional MLP heads, and training objectives. There is no guest combativeness or schooling. | |
| Q&A on Medusa Speculative Decoding and Meeting Conclusion | 4 | 0 | 0 | 0 | The host addresses a chat question regarding beam search computational overhead versus Medusa speculative decoding before concluding the session. |
Statements from this episode (0)
Nothing in this episode matches those filters. clear them