Nov 3, 2023 · 1h 18m · latent-space
Beating GPT-4 with Open Source Models - with Michael Royzen of Phind
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this interview, Phind founder Michael Royzen discusses his journey from building on-device computer vision apps to creating an industry-leading AI developer search engine. He details how Phind fine-tuned open-source models to outperform GPT-4 on coding benchmarks, scaled through Y Combinator, and designed specialized workflows for programmers.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 7.7% of the talking time here. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Michael politely but directly rejects the core premise presented by Cursor's Aman, arguing that features like diffs work fine in extensions and that developers need tools covering stages beyond the IDE.
Hardest push from the hosts ▶ 1:13:49 Swyx questions formal grammars for hallucination preventionSwyx pushes back against Michael's proposed direction of using formal grammars and LMQL, pointing out that LMQL is overly rigid for the vague goal of hallucination avoidance.
Biggest teaching moment ▶ 13:36 Early history of instruction tuning and Chinchilla optimalityMichael educates Swyx on the timeline of LLM development, explaining why Bloom underperformed due to suboptimal data ratios and detailing how BigScience T0 predated InstructGPT and Flan-T5 in multi-task instruction tuning.
The host holds their own ▶ 44:34 Swyx breaks down Code Llama pretraining parametersSwyx demonstrates precise technical mastery by reciting Code Llama's exact pretraining token count, extended context token thresholds, and modified RoPE theta hyperparameter settings.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Smart Lens Origins and Early Computer Vision Work | 4 | 3 | 1 | 1 | Swyx demonstrates domain awareness by comparing Smart Lens to Be My Eyes and citing usage statistics from the GPT-4 Vision system card. Michael shares the origins of fine-tuning Inception on ImageNet locally on iOS. | |
| Transitioning to NLP and Building the First Hello Search Engine | 5 | 6 | 1 | 2 | Michael walks through early LLM search history with BART, ELI5, and BigScience T0. Swyx actively probes whether T0 was a precursor to Bloom and clarifies Common Crawl filtering strategies. | |
| Defining Phind and Developer Workflow Integration | 3 | 4 | 4 | 1 | Michael explicitly disagrees with Cursor CEO Aman's stance on having to own the IDE, arguing that extensions and web-based conceptual planning are sufficient for developer workflows. | |
| Viral Growth, Paul Graham's Rebranding, and GPT-4 Integration | 4 | 4 | 1 | 2 | Swyx pushes on the non-standard spelling of Phind and questions potential friction with OpenAI over offering free GPT-4 access before official rollouts. Michael recounts Paul Graham rebranding the company. | |
| Training the Phind Model and Beating GPT-4 on Code Benchmarks | 6 | 6 | 2 | 2 | Swyx demonstrates technical depth by reciting the Code Llama 34B architecture specs, RoPE theta modifications, and context extension methods. Michael explains why HumanEval is contaminated and how Phind tops the BigCode leaderboard. | |
| Phind Pair Programmer, Message Pinning, and Replit Sandboxes | 5 | 3 | 1 | 2 | Swyx brings inside knowledge from conversations with Replit CEO Amjad Masad regarding sandbox APIs and upcoming competitive overlap. Michael outlines conversational state management via pinned messages. | |
| The YC Experience: Paul Graham, Ron Conway, and NVIDIA Compute | 3 | 4 | 1 | 1 | Michael shares the narrative of pitching Paul Graham, meeting Ron Conway, and getting Jensen Huang to allocate GPUs and write custom streaming features in FasterTransformer. The hosts guide the storytelling. | |
| Local Models, Quantization Formats, and Engineering Methodology | 4 | 5 | 1 | 1 | Michael breaks down the performance and precision differences between integer quantization and NVIDIA FP8 mixed precision in Transformer Engine. Swyx questions the minimum limits of quantization formats. | |
| Applied AI Hiring and Future Directions in AI Reasoning | 6 | 4 | 2 | 3 | Swyx debates the emerging Applied AI Engineer job title adopted by OpenAI and questions whether formal grammars like LMQL are appropriate for preventing general hallucinations. Michael details internal evaluation methodologies. |