Dec 31, 2025 · 34m · latent-space

[State of AI Papers 2025] Fixing Research with Social Signals, OCR & Implementation — Team AlphaXiv

0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of the Latent Space podcast, the founders of AlphaXiv discuss transforming static academic papers into interactive, executable research artifacts while analyzing standout NeurIPS architectures and practical AI solutions for the peer review and reproducibility crisis.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 4.9 Guest teaching 3.9 Guest disagreement 1.8 The hosts pushing back 3.1
05100:0010:0020:0030:000:00–3:07 · The hosts as informed peer 2/10 Origins of AlphaXiv at Stanford Host opens by inquiring about AlphaXiv's origins. The founders share their Stanford dorm background and the student research frustrations that led to building paper annotations.3:08–5:59 · The hosts as informed peer 6/10 Differentiating from Hugging Face and Exploring OCR Models Host probes why AlphaXiv won over Hugging Face and drills into PDF parsing and OCR model options. Rayhan details DeepSeek OCR's price-to-performance on A100s versus API alternatives.6:00–10:35 · The hosts as informed peer 5/10 Expanding Beyond Papers to Interactive Research Artifacts Host asks if AlphaXiv aims to replace Papers with Code or act as a recommendation system. The team outlines expanding from PDFs to executable Docker implementations.10:36–13:14 · The hosts as informed peer 5/10 Favorite Papers at NeurIPS: Compute-Efficient Architectures Rayhan shares compute-efficient NeurIPS papers including Tiny Recursive Models and low-rank evolutionary strategies. Host asks clarifying questions on search versus gradient descent.13:14–18:41 · The hosts as informed peer 6/10 AI for Science: Agent Laboratory and Virtual Labs Host pushes back against grouping all automated research under 'science' and questions recurring automated scientist claims. Guests reframe AI agents as pragmatic virtual lab assistants.18:42–22:12 · The hosts as informed peer 6/10 Reinforcement Learning for Reasoning and the Rise of Qwen Co-founders discuss Agent R-One and sample-efficient Qwen fine-tuning. Host references the presence of the Qwen team at NeurIPS and compares Qwen's trajectory with DeepSeek.22:12–25:04 · The hosts as informed peer 6/10 The Academic Peer Review Crisis and AI Linters Host raises the breakdown of academic peer review at conferences like ICLR. Founders discuss their refusal to build paper-writing bots while exploring pre-submission AI linters.25:04–28:55 · The hosts as informed peer 5/10 Overcoming Semantic Search with Social Signals and Dynamic Artifacts Rayhan explains why naive semantic search fails across 3 million arXiv papers without social signals. Host brings up Emergent Mind using YouTube traction as an alternate signal.28:55–34:21 · The hosts as informed peer 7/10 Empowering Applied Engineers and Ranking Implementation Ease Host gives direct pushback warning that automating paper execution environments is an impossible task, citing Replicate's pivot. Rayhan counters that they only focus on power-law papers and author-assisted setup.34:21–34:31 · The hosts as informed peer 1/10 Conclusion and Future Outlook Host wraps up the interview, commends AlphaXiv's mission, and gives closing well-wishes.0:00–3:07 · Guest teaching 2/10 Origins of AlphaXiv at Stanford Host opens by inquiring about AlphaXiv's origins. The founders share their Stanford dorm background and the student research frustrations that led to building paper annotations.3:08–5:59 · Guest teaching 4/10 Differentiating from Hugging Face and Exploring OCR Models Host probes why AlphaXiv won over Hugging Face and drills into PDF parsing and OCR model options. Rayhan details DeepSeek OCR's price-to-performance on A100s versus API alternatives.6:00–10:35 · Guest teaching 4/10 Expanding Beyond Papers to Interactive Research Artifacts Host asks if AlphaXiv aims to replace Papers with Code or act as a recommendation system. The team outlines expanding from PDFs to executable Docker implementations.10:36–13:14 · Guest teaching 6/10 Favorite Papers at NeurIPS: Compute-Efficient Architectures Rayhan shares compute-efficient NeurIPS papers including Tiny Recursive Models and low-rank evolutionary strategies. Host asks clarifying questions on search versus gradient descent.13:14–18:41 · Guest teaching 4/10 AI for Science: Agent Laboratory and Virtual Labs Host pushes back against grouping all automated research under 'science' and questions recurring automated scientist claims. Guests reframe AI agents as pragmatic virtual lab assistants.18:42–22:12 · Guest teaching 5/10 Reinforcement Learning for Reasoning and the Rise of Qwen Co-founders discuss Agent R-One and sample-efficient Qwen fine-tuning. Host references the presence of the Qwen team at NeurIPS and compares Qwen's trajectory with DeepSeek.22:12–25:04 · Guest teaching 4/10 The Academic Peer Review Crisis and AI Linters Host raises the breakdown of academic peer review at conferences like ICLR. Founders discuss their refusal to build paper-writing bots while exploring pre-submission AI linters.25:04–28:55 · Guest teaching 5/10 Overcoming Semantic Search with Social Signals and Dynamic Artifacts Rayhan explains why naive semantic search fails across 3 million arXiv papers without social signals. Host brings up Emergent Mind using YouTube traction as an alternate signal.28:55–34:21 · Guest teaching 5/10 Empowering Applied Engineers and Ranking Implementation Ease Host gives direct pushback warning that automating paper execution environments is an impossible task, citing Replicate's pivot. Rayhan counters that they only focus on power-law papers and author-assisted setup.34:21–34:31 · Guest teaching 0/10 Conclusion and Future Outlook Host wraps up the interview, commends AlphaXiv's mission, and gives closing well-wishes.0:00–3:07 · Guest disagreement 1/10 Origins of AlphaXiv at Stanford Host opens by inquiring about AlphaXiv's origins. The founders share their Stanford dorm background and the student research frustrations that led to building paper annotations.3:08–5:59 · Guest disagreement 2/10 Differentiating from Hugging Face and Exploring OCR Models Host probes why AlphaXiv won over Hugging Face and drills into PDF parsing and OCR model options. Rayhan details DeepSeek OCR's price-to-performance on A100s versus API alternatives.6:00–10:35 · Guest disagreement 2/10 Expanding Beyond Papers to Interactive Research Artifacts Host asks if AlphaXiv aims to replace Papers with Code or act as a recommendation system. The team outlines expanding from PDFs to executable Docker implementations.10:36–13:14 · Guest disagreement 1/10 Favorite Papers at NeurIPS: Compute-Efficient Architectures Rayhan shares compute-efficient NeurIPS papers including Tiny Recursive Models and low-rank evolutionary strategies. Host asks clarifying questions on search versus gradient descent.13:14–18:41 · Guest disagreement 3/10 AI for Science: Agent Laboratory and Virtual Labs Host pushes back against grouping all automated research under 'science' and questions recurring automated scientist claims. Guests reframe AI agents as pragmatic virtual lab assistants.18:42–22:12 · Guest disagreement 2/10 Reinforcement Learning for Reasoning and the Rise of Qwen Co-founders discuss Agent R-One and sample-efficient Qwen fine-tuning. Host references the presence of the Qwen team at NeurIPS and compares Qwen's trajectory with DeepSeek.22:12–25:04 · Guest disagreement 2/10 The Academic Peer Review Crisis and AI Linters Host raises the breakdown of academic peer review at conferences like ICLR. Founders discuss their refusal to build paper-writing bots while exploring pre-submission AI linters.25:04–28:55 · Guest disagreement 2/10 Overcoming Semantic Search with Social Signals and Dynamic Artifacts Rayhan explains why naive semantic search fails across 3 million arXiv papers without social signals. Host brings up Emergent Mind using YouTube traction as an alternate signal.28:55–34:21 · Guest disagreement 3/10 Empowering Applied Engineers and Ranking Implementation Ease Host gives direct pushback warning that automating paper execution environments is an impossible task, citing Replicate's pivot. Rayhan counters that they only focus on power-law papers and author-assisted setup.34:21–34:31 · Guest disagreement 0/10 Conclusion and Future Outlook Host wraps up the interview, commends AlphaXiv's mission, and gives closing well-wishes.0:00–3:07 · The hosts pushing back 1/10 Origins of AlphaXiv at Stanford Host opens by inquiring about AlphaXiv's origins. The founders share their Stanford dorm background and the student research frustrations that led to building paper annotations.3:08–5:59 · The hosts pushing back 3/10 Differentiating from Hugging Face and Exploring OCR Models Host probes why AlphaXiv won over Hugging Face and drills into PDF parsing and OCR model options. Rayhan details DeepSeek OCR's price-to-performance on A100s versus API alternatives.6:00–10:35 · The hosts pushing back 3/10 Expanding Beyond Papers to Interactive Research Artifacts Host asks if AlphaXiv aims to replace Papers with Code or act as a recommendation system. The team outlines expanding from PDFs to executable Docker implementations.10:36–13:14 · The hosts pushing back 2/10 Favorite Papers at NeurIPS: Compute-Efficient Architectures Rayhan shares compute-efficient NeurIPS papers including Tiny Recursive Models and low-rank evolutionary strategies. Host asks clarifying questions on search versus gradient descent.13:14–18:41 · The hosts pushing back 6/10 AI for Science: Agent Laboratory and Virtual Labs Host pushes back against grouping all automated research under 'science' and questions recurring automated scientist claims. Guests reframe AI agents as pragmatic virtual lab assistants.18:42–22:12 · The hosts pushing back 4/10 Reinforcement Learning for Reasoning and the Rise of Qwen Co-founders discuss Agent R-One and sample-efficient Qwen fine-tuning. Host references the presence of the Qwen team at NeurIPS and compares Qwen's trajectory with DeepSeek.22:12–25:04 · The hosts pushing back 3/10 The Academic Peer Review Crisis and AI Linters Host raises the breakdown of academic peer review at conferences like ICLR. Founders discuss their refusal to build paper-writing bots while exploring pre-submission AI linters.25:04–28:55 · The hosts pushing back 2/10 Overcoming Semantic Search with Social Signals and Dynamic Artifacts Rayhan explains why naive semantic search fails across 3 million arXiv papers without social signals. Host brings up Emergent Mind using YouTube traction as an alternate signal.28:55–34:21 · The hosts pushing back 7/10 Empowering Applied Engineers and Ranking Implementation Ease Host gives direct pushback warning that automating paper execution environments is an impossible task, citing Replicate's pivot. Rayhan counters that they only focus on power-law papers and author-assisted setup.34:21–34:31 · The hosts pushing back 0/10 Conclusion and Future Outlook Host wraps up the interview, commends AlphaXiv's mission, and gives closing well-wishes.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 31:30 Rayhan rejects host's framing on reproducibility sandboxes

Rayhan rejects the host's critique by arguing AlphaXiv is not trying to solve universal academic reproducibility, but rather catering strictly to applied engineers via power-law implementation ease.

Hardest push from the hosts ▶ 31:19 Host warns founders they are biting off more than they can chew

Host explicitly pushes back against the company's product direction, warning that building executable sandboxes for papers is an impossible task that caused Replicate to pivot.

Biggest teaching moment ▶ 11:36 Rayhan breaks down Tiny Recursive Models

Rayhan explains the mechanics of 7-million parameter recursive latent transformers achieving high ARC-AGI scores with minimal compute compared to massive frontier reasoning models.

The host holds their own ▶ 32:24 Host cites Brev and NVIDIA Launchables infrastructure

Host demonstrates deep domain expertise by citing NVIDIA Launchables and his background as an investor in Brev, explaining domain-specific solutions for hosting GPU environments.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Origins of AlphaXiv at Stanford 2211 Host opens by inquiring about AlphaXiv's origins. The founders share their Stanford dorm background and the student research frustrations that led to building paper annotations.
Differentiating from Hugging Face and Exploring OCR Models 6423 Host probes why AlphaXiv won over Hugging Face and drills into PDF parsing and OCR model options. Rayhan details DeepSeek OCR's price-to-performance on A100s versus API alternatives.
Expanding Beyond Papers to Interactive Research Artifacts 5423 Host asks if AlphaXiv aims to replace Papers with Code or act as a recommendation system. The team outlines expanding from PDFs to executable Docker implementations.
Favorite Papers at NeurIPS: Compute-Efficient Architectures 5612 Rayhan shares compute-efficient NeurIPS papers including Tiny Recursive Models and low-rank evolutionary strategies. Host asks clarifying questions on search versus gradient descent.
AI for Science: Agent Laboratory and Virtual Labs 6436 Host pushes back against grouping all automated research under 'science' and questions recurring automated scientist claims. Guests reframe AI agents as pragmatic virtual lab assistants.
Reinforcement Learning for Reasoning and the Rise of Qwen 6524 Co-founders discuss Agent R-One and sample-efficient Qwen fine-tuning. Host references the presence of the Qwen team at NeurIPS and compares Qwen's trajectory with DeepSeek.
The Academic Peer Review Crisis and AI Linters 6423 Host raises the breakdown of academic peer review at conferences like ICLR. Founders discuss their refusal to build paper-writing bots while exploring pre-submission AI linters.
Overcoming Semantic Search with Social Signals and Dynamic Artifacts 5522 Rayhan explains why naive semantic search fails across 3 million arXiv papers without social signals. Host brings up Emergent Mind using YouTube traction as an alternate signal.
Empowering Applied Engineers and Ranking Implementation Ease 7537 Host gives direct pushback warning that automating paper execution environments is an impossible task, citing Replicate's pivot. Rayhan counters that they only focus on power-law papers and author-assisted setup.
Conclusion and Future Outlook 1000 Host wraps up the interview, commends AlphaXiv's mission, and gives closing well-wishes.

Statements from this episode (0)

Nothing in this episode matches those filters. clear them

Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.