Sep 18, 2024 · 47m · latent-space

[Paper Club] 🍓 On Reasoning: Q-STaR and Friends!

Eugene Yan · 1m spoken Eugene Cheah · 1m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

This Paper Club presentation examines the progression of self-taught reasoning in language models, detailing the bootstrapping mechanics of STaR, the continuous token-level thoughts of Quiet-STaR, and the DPO verifier architecture of V-STaR along with their empirical and computational trade-offs.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 5.7 Guest teaching 2.3 Guest disagreement 1.3 The hosts pushing back 2.2
05100:0015:0030:0045:004:31–13:32 · The hosts as informed peer 6/10 STaR Algorithm Architecture and Experimental Setup The presenter leads the walkthrough of the original STaR paper, explaining rationales versus backward rationalizations. When Eugene Chia asks why they cannot train on all three candidate rationales, the presenter directly explains why flawed or insufficient reasoning degrades the dataset.13:32–20:01 · The hosts as informed peer 6/10 GSM8K Mathematical Reasoning and Human Correlation The presenter runs an interactive exercise on arithmetic reasoning problems from GSM8K to demonstrate GPT-J performance. The participants engage cooperatively to solve multi-step problems, and the presenter highlights the surprising human-machine step correlation.20:01–25:32 · The hosts as informed peer 5/10 Q&A on Synthetic Data and Domain Generalization Participants inquire about generalizing STaR beyond synthetic math and code data. Eugene Chia shares findings from training sub-1B models with digit inversion during calculations, prompting collaborative technical discussion.25:32–34:21 · The hosts as informed peer 5/10 Quiet-STaR and Parallel Token-Level Reasoning The presenter reviews Quiet-STaR and parallel thought generation, expressing skepticism about practical inference economics. RJ and Eugene Chia offer alternative interpretations regarding attention masking and token simulation in production models.34:21–41:28 · The hosts as informed peer 6/10 Quiet-STaR Architecture and Self-Correction Performance The discussion covers Quiet-STaR's mixing head and self-correction abilities, referencing John Schulman's remarks on training models to revise mistakes. The presenter dismisses the overall empirical lift of Quiet-STaR as underwhelming.41:28–47:11 · The hosts as informed peer 6/10 V-STaR and Verifier Training via DPO The presenter details V-STaR's approach of training a DPO verifier on both correct and incorrect rationales. Participants clarify the distinction between outcome reward models and process reward models.4:31–13:32 · Guest teaching 1/10 STaR Algorithm Architecture and Experimental Setup The presenter leads the walkthrough of the original STaR paper, explaining rationales versus backward rationalizations. When Eugene Chia asks why they cannot train on all three candidate rationales, the presenter directly explains why flawed or insufficient reasoning degrades the dataset.13:32–20:01 · Guest teaching 1/10 GSM8K Mathematical Reasoning and Human Correlation The presenter runs an interactive exercise on arithmetic reasoning problems from GSM8K to demonstrate GPT-J performance. The participants engage cooperatively to solve multi-step problems, and the presenter highlights the surprising human-machine step correlation.20:01–25:32 · Guest teaching 4/10 Q&A on Synthetic Data and Domain Generalization Participants inquire about generalizing STaR beyond synthetic math and code data. Eugene Chia shares findings from training sub-1B models with digit inversion during calculations, prompting collaborative technical discussion.25:32–34:21 · Guest teaching 3/10 Quiet-STaR and Parallel Token-Level Reasoning The presenter reviews Quiet-STaR and parallel thought generation, expressing skepticism about practical inference economics. RJ and Eugene Chia offer alternative interpretations regarding attention masking and token simulation in production models.34:21–41:28 · Guest teaching 3/10 Quiet-STaR Architecture and Self-Correction Performance The discussion covers Quiet-STaR's mixing head and self-correction abilities, referencing John Schulman's remarks on training models to revise mistakes. The presenter dismisses the overall empirical lift of Quiet-STaR as underwhelming.41:28–47:11 · Guest teaching 2/10 V-STaR and Verifier Training via DPO The presenter details V-STaR's approach of training a DPO verifier on both correct and incorrect rationales. Participants clarify the distinction between outcome reward models and process reward models.4:31–13:32 · Guest disagreement 1/10 STaR Algorithm Architecture and Experimental Setup The presenter leads the walkthrough of the original STaR paper, explaining rationales versus backward rationalizations. When Eugene Chia asks why they cannot train on all three candidate rationales, the presenter directly explains why flawed or insufficient reasoning degrades the dataset.13:32–20:01 · Guest disagreement 1/10 GSM8K Mathematical Reasoning and Human Correlation The presenter runs an interactive exercise on arithmetic reasoning problems from GSM8K to demonstrate GPT-J performance. The participants engage cooperatively to solve multi-step problems, and the presenter highlights the surprising human-machine step correlation.20:01–25:32 · Guest disagreement 2/10 Q&A on Synthetic Data and Domain Generalization Participants inquire about generalizing STaR beyond synthetic math and code data. Eugene Chia shares findings from training sub-1B models with digit inversion during calculations, prompting collaborative technical discussion.25:32–34:21 · Guest disagreement 2/10 Quiet-STaR and Parallel Token-Level Reasoning The presenter reviews Quiet-STaR and parallel thought generation, expressing skepticism about practical inference economics. RJ and Eugene Chia offer alternative interpretations regarding attention masking and token simulation in production models.34:21–41:28 · Guest disagreement 1/10 Quiet-STaR Architecture and Self-Correction Performance The discussion covers Quiet-STaR's mixing head and self-correction abilities, referencing John Schulman's remarks on training models to revise mistakes. The presenter dismisses the overall empirical lift of Quiet-STaR as underwhelming.41:28–47:11 · Guest disagreement 1/10 V-STaR and Verifier Training via DPO The presenter details V-STaR's approach of training a DPO verifier on both correct and incorrect rationales. Participants clarify the distinction between outcome reward models and process reward models.4:31–13:32 · The hosts pushing back 2/10 STaR Algorithm Architecture and Experimental Setup The presenter leads the walkthrough of the original STaR paper, explaining rationales versus backward rationalizations. When Eugene Chia asks why they cannot train on all three candidate rationales, the presenter directly explains why flawed or insufficient reasoning degrades the dataset.13:32–20:01 · The hosts pushing back 1/10 GSM8K Mathematical Reasoning and Human Correlation The presenter runs an interactive exercise on arithmetic reasoning problems from GSM8K to demonstrate GPT-J performance. The participants engage cooperatively to solve multi-step problems, and the presenter highlights the surprising human-machine step correlation.20:01–25:32 · The hosts pushing back 3/10 Q&A on Synthetic Data and Domain Generalization Participants inquire about generalizing STaR beyond synthetic math and code data. Eugene Chia shares findings from training sub-1B models with digit inversion during calculations, prompting collaborative technical discussion.25:32–34:21 · The hosts pushing back 3/10 Quiet-STaR and Parallel Token-Level Reasoning The presenter reviews Quiet-STaR and parallel thought generation, expressing skepticism about practical inference economics. RJ and Eugene Chia offer alternative interpretations regarding attention masking and token simulation in production models.34:21–41:28 · The hosts pushing back 2/10 Quiet-STaR Architecture and Self-Correction Performance The discussion covers Quiet-STaR's mixing head and self-correction abilities, referencing John Schulman's remarks on training models to revise mistakes. The presenter dismisses the overall empirical lift of Quiet-STaR as underwhelming.41:28–47:11 · The hosts pushing back 2/10 V-STaR and Verifier Training via DPO The presenter details V-STaR's approach of training a DPO verifier on both correct and incorrect rationales. Participants clarify the distinction between outcome reward models and process reward models.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 32:46 Disputing whether Claude uses genuine thinking tokens

Eugene Chia challenges the presenter's assertion that thinking tokens are merely UI prompts rather than architectural tokens.

Hardest push from the hosts ▶ 9:59 Rejecting the idea of training on all rationale candidates

The presenter firmly rejects Eugene Chia's premise of including all candidate rationales, pointing out that flawed logic undermines reasoning fine-tuning.

Biggest teaching moment ▶ 23:19 Explaining digit inversion benefits in transformer arithmetic

Eugene Chia educates the group on how reversing number order during generation aligns with traditional right-to-left math operations.

The host holds their own ▶ 33:00 Differentiating prompt-based tags from internal meta-tokens

The presenter clearly distinguishes un-emitted architectural meta-tokens in research from surface-level XML thinking tags used in commercial prompts.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
STaR Algorithm Architecture and Experimental Setup 6112 The presenter leads the walkthrough of the original STaR paper, explaining rationales versus backward rationalizations. When Eugene Chia asks why they cannot train on all three candidate rationales, the presenter directly explains why flawed or insufficient reasoning degrades the dataset.
GSM8K Mathematical Reasoning and Human Correlation 6111 The presenter runs an interactive exercise on arithmetic reasoning problems from GSM8K to demonstrate GPT-J performance. The participants engage cooperatively to solve multi-step problems, and the presenter highlights the surprising human-machine step correlation.
Q&A on Synthetic Data and Domain Generalization 5423 Participants inquire about generalizing STaR beyond synthetic math and code data. Eugene Chia shares findings from training sub-1B models with digit inversion during calculations, prompting collaborative technical discussion.
Quiet-STaR and Parallel Token-Level Reasoning 5323 The presenter reviews Quiet-STaR and parallel thought generation, expressing skepticism about practical inference economics. RJ and Eugene Chia offer alternative interpretations regarding attention masking and token simulation in production models.
Quiet-STaR Architecture and Self-Correction Performance 6312 The discussion covers Quiet-STaR's mixing head and self-correction abilities, referencing John Schulman's remarks on training models to revise mistakes. The presenter dismisses the overall empirical lift of Quiet-STaR as underwhelming.
V-STaR and Verifier Training via DPO 6212 The presenter details V-STaR's approach of training a DPO verifier on both correct and incorrect rationales. Participants clarify the distinction between outcome reward models and process reward models.

Statements from this episode (1)

Assertion Supported
Eugene Chia: Inverting numbers in reasoning traces improves math model performance
“The crazy one, the crazy thing that we did was that we inverted the numbers during the calculation and it seems to work better.”
Eugene Cheah Sep 18, 2024 ▶ 23:13
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.