Sep 18, 2024 · 47m · latent-space
[Paper Club] 🍓 On Reasoning: Q-STaR and Friends!
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
This Paper Club presentation examines the progression of self-taught reasoning in language models, detailing the bootstrapping mechanics of STaR, the continuous token-level thoughts of Quiet-STaR, and the DPO verifier architecture of V-STaR along with their empirical and computational trade-offs.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Eugene Chia challenges the presenter's assertion that thinking tokens are merely UI prompts rather than architectural tokens.
Hardest push from the hosts ▶ 9:59 Rejecting the idea of training on all rationale candidatesThe presenter firmly rejects Eugene Chia's premise of including all candidate rationales, pointing out that flawed logic undermines reasoning fine-tuning.
Biggest teaching moment ▶ 23:19 Explaining digit inversion benefits in transformer arithmeticEugene Chia educates the group on how reversing number order during generation aligns with traditional right-to-left math operations.
The host holds their own ▶ 33:00 Differentiating prompt-based tags from internal meta-tokensThe presenter clearly distinguishes un-emitted architectural meta-tokens in research from surface-level XML thinking tags used in commercial prompts.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| STaR Algorithm Architecture and Experimental Setup | 6 | 1 | 1 | 2 | The presenter leads the walkthrough of the original STaR paper, explaining rationales versus backward rationalizations. When Eugene Chia asks why they cannot train on all three candidate rationales, the presenter directly explains why flawed or insufficient reasoning degrades the dataset. | |
| GSM8K Mathematical Reasoning and Human Correlation | 6 | 1 | 1 | 1 | The presenter runs an interactive exercise on arithmetic reasoning problems from GSM8K to demonstrate GPT-J performance. The participants engage cooperatively to solve multi-step problems, and the presenter highlights the surprising human-machine step correlation. | |
| Q&A on Synthetic Data and Domain Generalization | 5 | 4 | 2 | 3 | Participants inquire about generalizing STaR beyond synthetic math and code data. Eugene Chia shares findings from training sub-1B models with digit inversion during calculations, prompting collaborative technical discussion. | |
| Quiet-STaR and Parallel Token-Level Reasoning | 5 | 3 | 2 | 3 | The presenter reviews Quiet-STaR and parallel thought generation, expressing skepticism about practical inference economics. RJ and Eugene Chia offer alternative interpretations regarding attention masking and token simulation in production models. | |
| Quiet-STaR Architecture and Self-Correction Performance | 6 | 3 | 1 | 2 | The discussion covers Quiet-STaR's mixing head and self-correction abilities, referencing John Schulman's remarks on training models to revise mistakes. The presenter dismisses the overall empirical lift of Quiet-STaR as underwhelming. | |
| V-STaR and Verifier Training via DPO | 6 | 2 | 1 | 2 | The presenter details V-STaR's approach of training a DPO verifier on both correct and incorrect rationales. Participants clarify the distinction between outcome reward models and process reward models. |