May 21, 2025 · 32m · latent-space

DeepWiki: The GitHub Encyclopedia

Silas Alberti · 14m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Silas Alberti of Cognition joins the Latent Space podcast to discuss DeepWiki, an AI-powered documentation and deep research tool that generates high-level architectural knowledge graphs for GitHub codebases. The conversation covers Cognition's custom infrastructure, graph-based code analysis, the future of cross-repository search, and their newly released open-source CUDA model trained with multi-turn reinforcement learning.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 5.5 Guest teaching 4.6 Guest disagreement 1.0 The hosts pushing back 1.4
05100:0010:0020:0030:001:07–5:17 · The hosts as informed peer 4/10 DeepWiki Origin and GitHub Research Concept Swix and Alessio probe the origins of DeepWiki and its compute costs. Silas explains their recency-weighted repository selection and indexing economics without conflict.5:17–10:01 · The hosts as informed peer 6/10 Security, Rate Limiting, and Frictionless Access Swix discusses VS Code extension architecture, ASTs, and dynamic linking challenges. Silas details how graph algorithms built by competitive programmer Gennady Korotkevich parse commit history and LSP graphs.10:01–14:30 · The hosts as informed peer 5/10 In-House Infrastructure and Indexing Queues The hosts inquire about orchestration tools, leading Silas to explain Cognition's in-house engineering culture and custom vector database implementation. Alessio suggests consumption patterns via MCP and deep links.14:30–19:22 · The hosts as informed peer 5/10 Browser Integration and DeepWiki as Dynamic Documentation Swix brings up OpenAI's GitHub deep research feature and discusses code embedding models and eval limitations. Silas describes why DeepWiki's high-level system mapping provides more grounded code exploration.19:24–21:24 · The hosts as informed peer 6/10 Live Demonstration with smol-podcaster Codebase Alessio shares a live walkthrough of DeepWiki indexing his smol-podcaster repository. Silas highlights how automated architectural summaries solve common developer onboarding friction.21:24–24:42 · The hosts as informed peer 6/10 Interactive Docs, Sandbox Execution, and Onboarding Friction Alessio argues for interactive code execution within documentation, leading into a technical discussion about dev container limitations and sandbox configuration hurdles during Devin onboarding.24:42–28:04 · The hosts as informed peer 6/10 Cross-Repository Deep Research across GitHub Silas proposes cross-repository deep research across all of GitHub. Alessio and Swix counter with real-world edge cases like stale forks and semantic search limits previously encountered by Sourcegraph.28:05–31:51 · The hosts as informed peer 6/10 Open Source CUDA Model Release and Multi-Turn RL Silas details Cognition's open-source CUDA model release trained with multi-turn RL. Swix connects the verifier mechanics directly to OpenAI's reinforcement fine-tuning paradigm.1:07–5:17 · Guest teaching 5/10 DeepWiki Origin and GitHub Research Concept Swix and Alessio probe the origins of DeepWiki and its compute costs. Silas explains their recency-weighted repository selection and indexing economics without conflict.5:17–10:01 · Guest teaching 6/10 Security, Rate Limiting, and Frictionless Access Swix discusses VS Code extension architecture, ASTs, and dynamic linking challenges. Silas details how graph algorithms built by competitive programmer Gennady Korotkevich parse commit history and LSP graphs.10:01–14:30 · Guest teaching 5/10 In-House Infrastructure and Indexing Queues The hosts inquire about orchestration tools, leading Silas to explain Cognition's in-house engineering culture and custom vector database implementation. Alessio suggests consumption patterns via MCP and deep links.14:30–19:22 · Guest teaching 4/10 Browser Integration and DeepWiki as Dynamic Documentation Swix brings up OpenAI's GitHub deep research feature and discusses code embedding models and eval limitations. Silas describes why DeepWiki's high-level system mapping provides more grounded code exploration.19:24–21:24 · Guest teaching 3/10 Live Demonstration with smol-podcaster Codebase Alessio shares a live walkthrough of DeepWiki indexing his smol-podcaster repository. Silas highlights how automated architectural summaries solve common developer onboarding friction.21:24–24:42 · Guest teaching 4/10 Interactive Docs, Sandbox Execution, and Onboarding Friction Alessio argues for interactive code execution within documentation, leading into a technical discussion about dev container limitations and sandbox configuration hurdles during Devin onboarding.24:42–28:04 · Guest teaching 4/10 Cross-Repository Deep Research across GitHub Silas proposes cross-repository deep research across all of GitHub. Alessio and Swix counter with real-world edge cases like stale forks and semantic search limits previously encountered by Sourcegraph.28:05–31:51 · Guest teaching 6/10 Open Source CUDA Model Release and Multi-Turn RL Silas details Cognition's open-source CUDA model release trained with multi-turn RL. Swix connects the verifier mechanics directly to OpenAI's reinforcement fine-tuning paradigm.1:07–5:17 · Guest disagreement 1/10 DeepWiki Origin and GitHub Research Concept Swix and Alessio probe the origins of DeepWiki and its compute costs. Silas explains their recency-weighted repository selection and indexing economics without conflict.5:17–10:01 · Guest disagreement 1/10 Security, Rate Limiting, and Frictionless Access Swix discusses VS Code extension architecture, ASTs, and dynamic linking challenges. Silas details how graph algorithms built by competitive programmer Gennady Korotkevich parse commit history and LSP graphs.10:01–14:30 · Guest disagreement 2/10 In-House Infrastructure and Indexing Queues The hosts inquire about orchestration tools, leading Silas to explain Cognition's in-house engineering culture and custom vector database implementation. Alessio suggests consumption patterns via MCP and deep links.14:30–19:22 · Guest disagreement 1/10 Browser Integration and DeepWiki as Dynamic Documentation Swix brings up OpenAI's GitHub deep research feature and discusses code embedding models and eval limitations. Silas describes why DeepWiki's high-level system mapping provides more grounded code exploration.19:24–21:24 · Guest disagreement 0/10 Live Demonstration with smol-podcaster Codebase Alessio shares a live walkthrough of DeepWiki indexing his smol-podcaster repository. Silas highlights how automated architectural summaries solve common developer onboarding friction.21:24–24:42 · Guest disagreement 1/10 Interactive Docs, Sandbox Execution, and Onboarding Friction Alessio argues for interactive code execution within documentation, leading into a technical discussion about dev container limitations and sandbox configuration hurdles during Devin onboarding.24:42–28:04 · Guest disagreement 1/10 Cross-Repository Deep Research across GitHub Silas proposes cross-repository deep research across all of GitHub. Alessio and Swix counter with real-world edge cases like stale forks and semantic search limits previously encountered by Sourcegraph.28:05–31:51 · Guest disagreement 1/10 Open Source CUDA Model Release and Multi-Turn RL Silas details Cognition's open-source CUDA model release trained with multi-turn RL. Swix connects the verifier mechanics directly to OpenAI's reinforcement fine-tuning paradigm.1:07–5:17 · The hosts pushing back 1/10 DeepWiki Origin and GitHub Research Concept Swix and Alessio probe the origins of DeepWiki and its compute costs. Silas explains their recency-weighted repository selection and indexing economics without conflict.5:17–10:01 · The hosts pushing back 2/10 Security, Rate Limiting, and Frictionless Access Swix discusses VS Code extension architecture, ASTs, and dynamic linking challenges. Silas details how graph algorithms built by competitive programmer Gennady Korotkevich parse commit history and LSP graphs.10:01–14:30 · The hosts pushing back 1/10 In-House Infrastructure and Indexing Queues The hosts inquire about orchestration tools, leading Silas to explain Cognition's in-house engineering culture and custom vector database implementation. Alessio suggests consumption patterns via MCP and deep links.14:30–19:22 · The hosts pushing back 2/10 Browser Integration and DeepWiki as Dynamic Documentation Swix brings up OpenAI's GitHub deep research feature and discusses code embedding models and eval limitations. Silas describes why DeepWiki's high-level system mapping provides more grounded code exploration.19:24–21:24 · The hosts pushing back 0/10 Live Demonstration with smol-podcaster Codebase Alessio shares a live walkthrough of DeepWiki indexing his smol-podcaster repository. Silas highlights how automated architectural summaries solve common developer onboarding friction.21:24–24:42 · The hosts pushing back 2/10 Interactive Docs, Sandbox Execution, and Onboarding Friction Alessio argues for interactive code execution within documentation, leading into a technical discussion about dev container limitations and sandbox configuration hurdles during Devin onboarding.24:42–28:04 · The hosts pushing back 2/10 Cross-Repository Deep Research across GitHub Silas proposes cross-repository deep research across all of GitHub. Alessio and Swix counter with real-world edge cases like stale forks and semantic search limits previously encountered by Sourcegraph.28:05–31:51 · The hosts pushing back 1/10 Open Source CUDA Model Release and Multi-Turn RL Silas details Cognition's open-source CUDA model release trained with multi-turn RL. Swix connects the verifier mechanics directly to OpenAI's reinforcement fine-tuning paradigm.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 10:44 Silas Defends Building Custom In-House Tooling

Silas playfully rejects the conventional SaaS approach of using off-the-shelf workflow orchestrators and vector databases, defending their culture of rolling custom infrastructure.

Hardest push from the hosts ▶ 15:28 Swix Challenges Deep Research Differentiation

Swix pushes Silas on whether DeepWiki's deep research mode provides measurable quality improvements over competing tools like OpenAI's GitHub research.

Biggest teaching moment ▶ 29:07 Silas Explains Multi-Turn RL Compiler Dynamics

Silas breaks down why single-turn models become overly conservative while multi-turn models take aggressive optimization risks because they have execution steps to self-correct compiler errors.

The host holds their own ▶ 25:42 Alessio Details Real-World Code Search Failure Modes

Alessio leverages his deep framework experience with Rails gems to demonstrate why naive cross-repo search fails on abandoned forks and outdated packages.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
DeepWiki Origin and GitHub Research Concept 4511 Swix and Alessio probe the origins of DeepWiki and its compute costs. Silas explains their recency-weighted repository selection and indexing economics without conflict.
Security, Rate Limiting, and Frictionless Access 6612 Swix discusses VS Code extension architecture, ASTs, and dynamic linking challenges. Silas details how graph algorithms built by competitive programmer Gennady Korotkevich parse commit history and LSP graphs.
In-House Infrastructure and Indexing Queues 5521 The hosts inquire about orchestration tools, leading Silas to explain Cognition's in-house engineering culture and custom vector database implementation. Alessio suggests consumption patterns via MCP and deep links.
Browser Integration and DeepWiki as Dynamic Documentation 5412 Swix brings up OpenAI's GitHub deep research feature and discusses code embedding models and eval limitations. Silas describes why DeepWiki's high-level system mapping provides more grounded code exploration.
Live Demonstration with smol-podcaster Codebase 6300 Alessio shares a live walkthrough of DeepWiki indexing his smol-podcaster repository. Silas highlights how automated architectural summaries solve common developer onboarding friction.
Interactive Docs, Sandbox Execution, and Onboarding Friction 6412 Alessio argues for interactive code execution within documentation, leading into a technical discussion about dev container limitations and sandbox configuration hurdles during Devin onboarding.
Cross-Repository Deep Research across GitHub 6412 Silas proposes cross-repository deep research across all of GitHub. Alessio and Swix counter with real-world edge cases like stale forks and semantic search limits previously encountered by Sourcegraph.
Open Source CUDA Model Release and Multi-Turn RL 6611 Silas details Cognition's open-source CUDA model release trained with multi-turn RL. Swix connects the verifier mechanics directly to OpenAI's reinforcement fine-tuning paradigm.

Statements from this episode (17)

Assertion Supported
Alberti: Cognition is based in SF after early nomadic hacker houses
“We're actually mainly based here in SF. I mean, yeah, the founding mythology is, I think we were a nomadic company in the beginning. So we had hacker house in different places. I think the first one was actually in Burning Game. Then it was in New York, then b…”
Silas Alberti May 21, 2025 ▶ 0:40
Disclosure
Silas Alberti: DeepWiki began as a free open-source side project
“I mean, it was supposed to be this like open source focus, like free site project almost, but it became bigger than expected.”
Silas Alberti May 21, 2025 ▶ 1:34
Disclosure
Silas Alberti: DeepWiki pairs AI codebase documentation with a research agent
“And I think it has like two components. It has the wiki itself, which is basically like AI generated Got documentation for any code base on GitHub, and then there's, like, the deep research agent that leverages both this, like, wiki, but also, like, any code f…”
Silas Alberti May 21, 2025 ▶ 1:54
Disclosure
Silas Alberti: DeepWiki scaled infrastructure for tens of thousands of developers
“I think the particular thing that we had to make happen for DeepWiki was, like, scaling up the infra so we could support, like, tens of thousands of open source developers at the same time. And I think also just, like, Making it be, there's like no sign up req…”
Silas Alberti May 21, 2025 ▶ 2:30
Prediction Not checkable as stated
Alberti: DeepWiki compute spend may break $1M soon
“So we think we should be now at like the high, like tens of thousands of repos. So I guess you can like do the rough map. Like we might be approaching quite significant numbers of compute spend here. I mean maybe we'll like break the million soon.”
Silas Alberti May 21, 2025 ▶ 3:46
Assertion Not checkable as stated
Alberti: 80% to 90% of DeepWiki users view pre-indexed popular repos
“So I would say, like, 80%, 90% of users just, like, check out one of the popular repos.”
Silas Alberti May 21, 2025 ▶ 5:13
Insight
Alberti: Folder structure graphs are poor representations of codebase architecture
“There are some, the simplest one is a folder structure graph, but unfortunately, it's, like, actually a pretty bad one because people sometimes just, you know, like, put all the components in this folder and all the servers in this folder.”
Silas Alberti May 21, 2025 ▶ 8:28
Disclosure
Alberti: Cognition built its own in-house vector database
“This culture of like building everything in-house. It's also kind of funny, like we, there's like a rack component to this. And I've been actually, I personally advocate, I can just like use a vector database and make life easy. But then there's a really stron…”
Silas Alberti May 21, 2025 ▶ 10:46
Assertion Not checkable as stated
Alberti: DeepWiki experienced spikes of 2,000 indexing jobs queued
“Sometimes it's like random spikes. I don't know how they happen. Like 2000 indexing drops in the queue. Suddenly it was like the other day.”
Silas Alberti May 21, 2025 ▶ 11:35
Disclosure
Alberti: Devin updates codebase indexes on every commit for paid users
“If you like pay for Devin, for example, like we actually update it on literally every commit incrementally.”
Silas Alberti May 21, 2025 ▶ 13:42
Assertion Supported
Alberti: Over 1,000 GitHub projects added DeepWiki badges to their repositories
“There was more than a thousand projects already had The like deep wiki batch, which I thought was like super awesome to see like way more than I expected.”
Silas Alberti May 21, 2025 ▶ 14:02
Disclosure
Alberti: DeepWiki auto-updates documentation for free repositories displaying its badge
“So we just decided, okay, how about we just detect whether you have the wiki batch and then we'll keep it updated for you.”
Silas Alberti May 21, 2025 ▶ 14:19
Insight
Alberti: Small numbers of manually curated evals beat large eval sets
“I think usually like small numbers of very high quality evals are the way to go and like we curate them manually and make sure they're really good.”
Silas Alberti May 21, 2025 ▶ 17:39
Disclosure
Alberti: Setting up dev environments is Devin's highest onboarding hurdle
“So Devon actually needs, like, a full dev environment to be set up. And I think you guys probably both went through this when you onboarded with Devon. It's, like, probably, like, the highest friction onboarding point that we still need to, like,”
Silas Alberti May 21, 2025 ▶ 23:25
Insight
Alberti: Wiki pages are a better abstraction than pure RAG for codebase search
“If you just do like pure, like rack on like such a big base code files, it'll just be like pretty bad at a certain point, you know, on a single code base. Sure. I can see it, but Tens of thousands of code bases. It's tougher. But I think actually the wiki page…”
Silas Alberti May 21, 2025 ▶ 26:29
Insight
Alberti: Multi-Turn RL Enables Aggressive Code Optimization Over Single-Turn Models
“Basically the single-turn model that was just trained on, like, getting the best result after one turn. It would basically be a little bit, like, too careful, because it couldn't risk writing, like, non-compiling code, whereas, like, the multi-turn model would…”
Silas Alberti May 21, 2025 ▶ 30:19
Prediction Not checkable as stated
Alberti: AI App Companies Must Encode Product Needs Into RL Feedback
“I think that's almost how I view like the future of application layer companies. Cause I mean, yeah, you see like the different, the labs are also now creating these like RL platforms and you can soon like customize models with RL on your personal, like on you…”
Silas Alberti May 21, 2025 ▶ 31:16
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.