Nov 3, 2023 · 1h 18m · latent-space

Beating GPT-4 with Open Source Models - with Michael Royzen of Phind

Michael Royzen · 1h 2m spoken Shawn Wang · 5m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this interview, Phind founder Michael Royzen discusses his journey from building on-device computer vision apps to creating an industry-leading AI developer search engine. He details how Phind fine-tuned open-source models to outperform GPT-4 on coding benchmarks, scaled through Y Combinator, and designed specialized workflows for programmers.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 7.7% of the talking time here. How this is scored →

The hosts as informed peer 4.4 Guest teaching 4.3 Guest disagreement 1.6 The hosts pushing back 1.7
05100:0020:0040:001:00:000:07–5:10 · The hosts as informed peer 4/10 Smart Lens Origins and Early Computer Vision Work Swyx demonstrates domain awareness by comparing Smart Lens to Be My Eyes and citing usage statistics from the GPT-4 Vision system card. Michael shares the origins of fine-tuning Inception on ImageNet locally on iOS.5:10–20:57 · The hosts as informed peer 5/10 Transitioning to NLP and Building the First Hello Search Engine Michael walks through early LLM search history with BART, ELI5, and BigScience T0. Swyx actively probes whether T0 was a precursor to Bloom and clarifies Common Crawl filtering strategies.20:57–28:21 · The hosts as informed peer 3/10 Defining Phind and Developer Workflow Integration Michael explicitly disagrees with Cursor CEO Aman's stance on having to own the IDE, arguing that extensions and web-based conceptual planning are sufficient for developer workflows.28:21–37:24 · The hosts as informed peer 4/10 Viral Growth, Paul Graham's Rebranding, and GPT-4 Integration Swyx pushes on the non-standard spelling of Phind and questions potential friction with OpenAI over offering free GPT-4 access before official rollouts. Michael recounts Paul Graham rebranding the company.37:24–47:03 · The hosts as informed peer 6/10 Training the Phind Model and Beating GPT-4 on Code Benchmarks Swyx demonstrates technical depth by reciting the Code Llama 34B architecture specs, RoPE theta modifications, and context extension methods. Michael explains why HumanEval is contaminated and how Phind tops the BigCode leaderboard.47:03–51:28 · The hosts as informed peer 5/10 Phind Pair Programmer, Message Pinning, and Replit Sandboxes Swyx brings inside knowledge from conversations with Replit CEO Amjad Masad regarding sandbox APIs and upcoming competitive overlap. Michael outlines conversational state management via pinned messages.51:28–1:04:47 · The hosts as informed peer 3/10 The YC Experience: Paul Graham, Ron Conway, and NVIDIA Compute Michael shares the narrative of pitching Paul Graham, meeting Ron Conway, and getting Jensen Huang to allocate GPUs and write custom streaming features in FasterTransformer. The hosts guide the storytelling.1:04:47–1:10:08 · The hosts as informed peer 4/10 Local Models, Quantization Formats, and Engineering Methodology Michael breaks down the performance and precision differences between integer quantization and NVIDIA FP8 mixed precision in Transformer Engine. Swyx questions the minimum limits of quantization formats.1:10:08–1:18:37 · The hosts as informed peer 6/10 Applied AI Hiring and Future Directions in AI Reasoning Swyx debates the emerging Applied AI Engineer job title adopted by OpenAI and questions whether formal grammars like LMQL are appropriate for preventing general hallucinations. Michael details internal evaluation methodologies.0:07–5:10 · Guest teaching 3/10 Smart Lens Origins and Early Computer Vision Work Swyx demonstrates domain awareness by comparing Smart Lens to Be My Eyes and citing usage statistics from the GPT-4 Vision system card. Michael shares the origins of fine-tuning Inception on ImageNet locally on iOS.5:10–20:57 · Guest teaching 6/10 Transitioning to NLP and Building the First Hello Search Engine Michael walks through early LLM search history with BART, ELI5, and BigScience T0. Swyx actively probes whether T0 was a precursor to Bloom and clarifies Common Crawl filtering strategies.20:57–28:21 · Guest teaching 4/10 Defining Phind and Developer Workflow Integration Michael explicitly disagrees with Cursor CEO Aman's stance on having to own the IDE, arguing that extensions and web-based conceptual planning are sufficient for developer workflows.28:21–37:24 · Guest teaching 4/10 Viral Growth, Paul Graham's Rebranding, and GPT-4 Integration Swyx pushes on the non-standard spelling of Phind and questions potential friction with OpenAI over offering free GPT-4 access before official rollouts. Michael recounts Paul Graham rebranding the company.37:24–47:03 · Guest teaching 6/10 Training the Phind Model and Beating GPT-4 on Code Benchmarks Swyx demonstrates technical depth by reciting the Code Llama 34B architecture specs, RoPE theta modifications, and context extension methods. Michael explains why HumanEval is contaminated and how Phind tops the BigCode leaderboard.47:03–51:28 · Guest teaching 3/10 Phind Pair Programmer, Message Pinning, and Replit Sandboxes Swyx brings inside knowledge from conversations with Replit CEO Amjad Masad regarding sandbox APIs and upcoming competitive overlap. Michael outlines conversational state management via pinned messages.51:28–1:04:47 · Guest teaching 4/10 The YC Experience: Paul Graham, Ron Conway, and NVIDIA Compute Michael shares the narrative of pitching Paul Graham, meeting Ron Conway, and getting Jensen Huang to allocate GPUs and write custom streaming features in FasterTransformer. The hosts guide the storytelling.1:04:47–1:10:08 · Guest teaching 5/10 Local Models, Quantization Formats, and Engineering Methodology Michael breaks down the performance and precision differences between integer quantization and NVIDIA FP8 mixed precision in Transformer Engine. Swyx questions the minimum limits of quantization formats.1:10:08–1:18:37 · Guest teaching 4/10 Applied AI Hiring and Future Directions in AI Reasoning Swyx debates the emerging Applied AI Engineer job title adopted by OpenAI and questions whether formal grammars like LMQL are appropriate for preventing general hallucinations. Michael details internal evaluation methodologies.0:07–5:10 · Guest disagreement 1/10 Smart Lens Origins and Early Computer Vision Work Swyx demonstrates domain awareness by comparing Smart Lens to Be My Eyes and citing usage statistics from the GPT-4 Vision system card. Michael shares the origins of fine-tuning Inception on ImageNet locally on iOS.5:10–20:57 · Guest disagreement 1/10 Transitioning to NLP and Building the First Hello Search Engine Michael walks through early LLM search history with BART, ELI5, and BigScience T0. Swyx actively probes whether T0 was a precursor to Bloom and clarifies Common Crawl filtering strategies.20:57–28:21 · Guest disagreement 4/10 Defining Phind and Developer Workflow Integration Michael explicitly disagrees with Cursor CEO Aman's stance on having to own the IDE, arguing that extensions and web-based conceptual planning are sufficient for developer workflows.28:21–37:24 · Guest disagreement 1/10 Viral Growth, Paul Graham's Rebranding, and GPT-4 Integration Swyx pushes on the non-standard spelling of Phind and questions potential friction with OpenAI over offering free GPT-4 access before official rollouts. Michael recounts Paul Graham rebranding the company.37:24–47:03 · Guest disagreement 2/10 Training the Phind Model and Beating GPT-4 on Code Benchmarks Swyx demonstrates technical depth by reciting the Code Llama 34B architecture specs, RoPE theta modifications, and context extension methods. Michael explains why HumanEval is contaminated and how Phind tops the BigCode leaderboard.47:03–51:28 · Guest disagreement 1/10 Phind Pair Programmer, Message Pinning, and Replit Sandboxes Swyx brings inside knowledge from conversations with Replit CEO Amjad Masad regarding sandbox APIs and upcoming competitive overlap. Michael outlines conversational state management via pinned messages.51:28–1:04:47 · Guest disagreement 1/10 The YC Experience: Paul Graham, Ron Conway, and NVIDIA Compute Michael shares the narrative of pitching Paul Graham, meeting Ron Conway, and getting Jensen Huang to allocate GPUs and write custom streaming features in FasterTransformer. The hosts guide the storytelling.1:04:47–1:10:08 · Guest disagreement 1/10 Local Models, Quantization Formats, and Engineering Methodology Michael breaks down the performance and precision differences between integer quantization and NVIDIA FP8 mixed precision in Transformer Engine. Swyx questions the minimum limits of quantization formats.1:10:08–1:18:37 · Guest disagreement 2/10 Applied AI Hiring and Future Directions in AI Reasoning Swyx debates the emerging Applied AI Engineer job title adopted by OpenAI and questions whether formal grammars like LMQL are appropriate for preventing general hallucinations. Michael details internal evaluation methodologies.0:07–5:10 · The hosts pushing back 1/10 Smart Lens Origins and Early Computer Vision Work Swyx demonstrates domain awareness by comparing Smart Lens to Be My Eyes and citing usage statistics from the GPT-4 Vision system card. Michael shares the origins of fine-tuning Inception on ImageNet locally on iOS.5:10–20:57 · The hosts pushing back 2/10 Transitioning to NLP and Building the First Hello Search Engine Michael walks through early LLM search history with BART, ELI5, and BigScience T0. Swyx actively probes whether T0 was a precursor to Bloom and clarifies Common Crawl filtering strategies.20:57–28:21 · The hosts pushing back 1/10 Defining Phind and Developer Workflow Integration Michael explicitly disagrees with Cursor CEO Aman's stance on having to own the IDE, arguing that extensions and web-based conceptual planning are sufficient for developer workflows.28:21–37:24 · The hosts pushing back 2/10 Viral Growth, Paul Graham's Rebranding, and GPT-4 Integration Swyx pushes on the non-standard spelling of Phind and questions potential friction with OpenAI over offering free GPT-4 access before official rollouts. Michael recounts Paul Graham rebranding the company.37:24–47:03 · The hosts pushing back 2/10 Training the Phind Model and Beating GPT-4 on Code Benchmarks Swyx demonstrates technical depth by reciting the Code Llama 34B architecture specs, RoPE theta modifications, and context extension methods. Michael explains why HumanEval is contaminated and how Phind tops the BigCode leaderboard.47:03–51:28 · The hosts pushing back 2/10 Phind Pair Programmer, Message Pinning, and Replit Sandboxes Swyx brings inside knowledge from conversations with Replit CEO Amjad Masad regarding sandbox APIs and upcoming competitive overlap. Michael outlines conversational state management via pinned messages.51:28–1:04:47 · The hosts pushing back 1/10 The YC Experience: Paul Graham, Ron Conway, and NVIDIA Compute Michael shares the narrative of pitching Paul Graham, meeting Ron Conway, and getting Jensen Huang to allocate GPUs and write custom streaming features in FasterTransformer. The hosts guide the storytelling.1:04:47–1:10:08 · The hosts pushing back 1/10 Local Models, Quantization Formats, and Engineering Methodology Michael breaks down the performance and precision differences between integer quantization and NVIDIA FP8 mixed precision in Transformer Engine. Swyx questions the minimum limits of quantization formats.1:10:08–1:18:37 · The hosts pushing back 3/10 Applied AI Hiring and Future Directions in AI Reasoning Swyx debates the emerging Applied AI Engineer job title adopted by OpenAI and questions whether formal grammars like LMQL are appropriate for preventing general hallucinations. Michael details internal evaluation methodologies.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 24% · guest 76%0:00 · the hosts 24% · guest 76%3:00 · the hosts 14.7% · guest 85.3%3:00 · the hosts 14.7% · guest 85.3%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 3.6% · guest 96.4%9:00 · the hosts 3.6% · guest 96.4%12:00 · the hosts 1.7% · guest 98.3%12:00 · the hosts 1.7% · guest 98.3%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0.3% · guest 99.7%21:00 · the hosts 0.3% · guest 99.7%24:00 · the hosts 11.9% · guest 88.1%24:00 · the hosts 11.9% · guest 88.1%27:00 · the hosts 22.1% · guest 77.9%27:00 · the hosts 22.1% · guest 77.9%30:00 · the hosts 13.8% · guest 86.2%30:00 · the hosts 13.8% · guest 86.2%33:00 · the hosts 6.2% · guest 93.8%33:00 · the hosts 6.2% · guest 93.8%36:00 · the hosts 8% · guest 92%36:00 · the hosts 8% · guest 92%39:00 · the hosts 0.2% · guest 99.8%39:00 · the hosts 0.2% · guest 99.8%42:00 · the hosts 10.2% · guest 89.8%42:00 · the hosts 10.2% · guest 89.8%45:00 · the hosts 1.7% · guest 98.3%45:00 · the hosts 1.7% · guest 98.3%48:00 · the hosts 9.1% · guest 90.9%48:00 · the hosts 9.1% · guest 90.9%51:00 · the hosts 9.5% · guest 90.5%51:00 · the hosts 9.5% · guest 90.5%54:00 · the hosts 1.6% · guest 98.4%54:00 · the hosts 1.6% · guest 98.4%57:00 · the hosts 12.7% · guest 87.3%57:00 · the hosts 12.7% · guest 87.3%1:00:00 · the hosts 5.5% · guest 94.5%1:00:00 · the hosts 5.5% · guest 94.5%1:03:00 · the hosts 11.4% · guest 88.6%1:03:00 · the hosts 11.4% · guest 88.6%1:06:00 · the hosts 5.6% · guest 94.4%1:06:00 · the hosts 5.6% · guest 94.4%1:09:00 · the hosts 9.6% · guest 90.4%1:09:00 · the hosts 9.6% · guest 90.4%1:12:00 · the hosts 14.6% · guest 85.4%1:12:00 · the hosts 14.6% · guest 85.4%1:15:00 · the hosts 4.9% · guest 95.1%1:15:00 · the hosts 4.9% · guest 95.1%1:18:00 · the hosts 7.3% · guest 92.7%1:18:00 · the hosts 7.3% · guest 92.7%
Sharpest disagreement ▶ 23:42 Pushback against Cursor's IDE ownership thesis

Michael politely but directly rejects the core premise presented by Cursor's Aman, arguing that features like diffs work fine in extensions and that developers need tools covering stages beyond the IDE.

Hardest push from the hosts ▶ 1:13:49 Swyx questions formal grammars for hallucination prevention

Swyx pushes back against Michael's proposed direction of using formal grammars and LMQL, pointing out that LMQL is overly rigid for the vague goal of hallucination avoidance.

Biggest teaching moment ▶ 13:36 Early history of instruction tuning and Chinchilla optimality

Michael educates Swyx on the timeline of LLM development, explaining why Bloom underperformed due to suboptimal data ratios and detailing how BigScience T0 predated InstructGPT and Flan-T5 in multi-task instruction tuning.

The host holds their own ▶ 44:34 Swyx breaks down Code Llama pretraining parameters

Swyx demonstrates precise technical mastery by reciting Code Llama's exact pretraining token count, extended context token thresholds, and modified RoPE theta hyperparameter settings.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Smart Lens Origins and Early Computer Vision Work 4311 Swyx demonstrates domain awareness by comparing Smart Lens to Be My Eyes and citing usage statistics from the GPT-4 Vision system card. Michael shares the origins of fine-tuning Inception on ImageNet locally on iOS.
Transitioning to NLP and Building the First Hello Search Engine 5612 Michael walks through early LLM search history with BART, ELI5, and BigScience T0. Swyx actively probes whether T0 was a precursor to Bloom and clarifies Common Crawl filtering strategies.
Defining Phind and Developer Workflow Integration 3441 Michael explicitly disagrees with Cursor CEO Aman's stance on having to own the IDE, arguing that extensions and web-based conceptual planning are sufficient for developer workflows.
Viral Growth, Paul Graham's Rebranding, and GPT-4 Integration 4412 Swyx pushes on the non-standard spelling of Phind and questions potential friction with OpenAI over offering free GPT-4 access before official rollouts. Michael recounts Paul Graham rebranding the company.
Training the Phind Model and Beating GPT-4 on Code Benchmarks 6622 Swyx demonstrates technical depth by reciting the Code Llama 34B architecture specs, RoPE theta modifications, and context extension methods. Michael explains why HumanEval is contaminated and how Phind tops the BigCode leaderboard.
Phind Pair Programmer, Message Pinning, and Replit Sandboxes 5312 Swyx brings inside knowledge from conversations with Replit CEO Amjad Masad regarding sandbox APIs and upcoming competitive overlap. Michael outlines conversational state management via pinned messages.
The YC Experience: Paul Graham, Ron Conway, and NVIDIA Compute 3411 Michael shares the narrative of pitching Paul Graham, meeting Ron Conway, and getting Jensen Huang to allocate GPUs and write custom streaming features in FasterTransformer. The hosts guide the storytelling.
Local Models, Quantization Formats, and Engineering Methodology 4511 Michael breaks down the performance and precision differences between integer quantization and NVIDIA FP8 mixed precision in Transformer Engine. Swyx questions the minimum limits of quantization formats.
Applied AI Hiring and Future Directions in AI Reasoning 6423 Swyx debates the emerging Applied AI Engineer job title adopted by OpenAI and questions whether formal grammars like LMQL are appropriate for preventing general hallucinations. Michael details internal evaluation methodologies.

Statements from this episode (24)

Opinion
Royzen: Smart Lens worked better than Google Lens v1 at launch
“It worked so well that it actually worked better than Google Lens which released its V-One around the same time.”
Michael Royzen Nov 3, 2023 ▶ 3:01
Assertion Supported
Swix: Be My Eyes users only use the app 1.5 times per day
“The average usage of BMIs per day is 1.5 times.”
Shawn Wang Nov 3, 2023 ▶ 4:15
Insight
Royzen: Users prefer text-to-image generation over image captioning
“I was also looking into image captioning, where like you give a model an image, and then it tells you what's in the image, but it turns out that what people want is the exact opposite. People want to give a description of an image, and then have the AI generat…”
Michael Royzen Nov 3, 2023 ▶ 4:20
Assertion Partly supported
Royzen: BigScience's T-Zero predated InstructGPT in large-scale instruction tuning
“I think T-Zero is the first model that did large-scale instruction tuning from diverse data sources in the fall of twenty-twenty-one. This is before InstructGPT. This is before Flan T-Five, which came out in twenty-twenty-two. This is, I think, the very, very …”
Michael Royzen Nov 3, 2023 ▶ 14:42
Assertion Supported
Royzen: Phind built the first internet-scale LLM RAG search in 2022
“And to the best of my knowledge, I think that's the first example that I'm aware of a LLM search engine model that's effectively connected to, like, a large enough index that I would consider, like, an internet scale. So, so I think we were the first to releas…”
Michael Royzen Nov 3, 2023 ▶ 16:07
Disclosure
Royzen: Phind focuses on developer reasoning engines over general search
“So as, I think there's always an opportunity for us to become more general if we wanted. But We've been along this path of, like, what is the best, most advanced reasoning engine that's connected to your code base, that's connected to the internet, that we can…”
Michael Royzen Nov 3, 2023 ▶ 20:37
Opinion
Royzen: AI code startups are ultimately competing with ChatGPT
“Really who everyone's competing with is ChatGPT which only has, like, that one web interface, and, like, ChatGPT is really the bar.”
Michael Royzen Nov 3, 2023 ▶ 23:08
Opinion
Royzen: AI dev tools do not need to own the IDE
“Somewhere where I disagree with him is that you need to own the IDE. I think like he made kind of some good points about, you know, not having platform risk in the long term, but some of the, you know, features that were mentioned, like suggesting diffs, for e…”
Michael Royzen Nov 3, 2023 ▶ 24:29
Prediction Not checkable as stated
Phind CEO: Future Programming Will Just Be Problem Solving, Delegating Implementation to AI
“In the future, you know, in the future, I think programming is just going to be really just the problem solving. Like you come up with an idea, you come up with like the basic design for the algorithm in your head, and you just tell the AI, hey, just like, jus…”
Michael Royzen Nov 3, 2023 ▶ 25:41
Assertion Not checkable as stated
Royzen: Users switch to Phind when ChatGPT-4 fails on code
“What really shocks us is that a lot of the people who do that they're coming from ChatGPT. So they tried it in ChatGPT with ChatGPT-IV. It didn't work. Maybe it required like some multi-step reasoning. Maybe it required to like, Some internet context or someth…”
Michael Royzen Nov 3, 2023 ▶ 27:07
Assertion Not checkable as stated
Royzen: Paul Graham personally chose the company name 'Phind'
“Paul Graham actually picked it for us.”
Michael Royzen Nov 3, 2023 ▶ 30:01
Prediction Not checkable as stated
Royzen: The leap from GPT-4 to GPT-5 will be smaller
“I think that GPT-IV, my hypothesis is that the jump from four to 4.5, or four to five, will be smaller than the jump from Three to four.”
Michael Royzen Nov 3, 2023 ▶ 37:38
Prediction Not checkable as stated
Royzen: Fine-tuned open source models will beat proprietary in 2024
“So I think that even if a delta exists, in twenty-twenty-four, the delta between proprietary and open source won't be large enough that a startup like us, with a lot of data that we've collected, can take the data that we have, fine-tune an open source model, …”
Michael Royzen Nov 3, 2023 ▶ 38:25
Insight
Royzen: Large context windows outperform RAG chunking for code
“Like, I think it's generally been shown that if you have the space to just put The raw files inside of a big context window. That is still better than chunking and retrieval. It just is.”
Michael Royzen Nov 3, 2023 ▶ 39:24
Assertion Partly supported
Royzen: Phind Leads BigCode Leaderboard by 10 Points in Multi-Language Code
“All of our models are at the top of the big code leaderboard by far. It's not close, particularly in languages other than Python. We have a 10 point gap between us and the next best model on Java, JavaScript, I think C-sharp multilingual.”
Michael Royzen Nov 3, 2023 ▶ 41:03
Assertion Not checkable as stated
Royzen: GPT-4 was trained on HumanEval, proving data contamination
“GPT-IV itself has been trained on human eval, and we know this because GPT-IV is able to predict the exact doc string in many of the problems. I've seen it predict, like, the specific example values in the doc string, which is extremely improbable for it to ju…”
Michael Royzen Nov 3, 2023 ▶ 41:31
Assertion Not checkable as stated
Royzen: Training on code unlocked general spatial and temporal reasoning
“We've seen emerging capabilities in the find model, whereby training it on high quality code, it can actually, like, reason better. It went from not being able to solve like, World problems where like riddles where like with like temporal and like low, like pl…”
Michael Royzen Nov 3, 2023 ▶ 43:51
Disclosure
Phind Plans Native Code Interpreter Features to Recursively Iterate on Code
“And Replit is great, and people use that feature. But yeah, I think there's more we can do in terms of, like, having something a bit closer to code interpreter where it's able to run the code and then, like, recursively iterate on it.”
Michael Royzen Nov 3, 2023 ▶ 49:32
Assertion Supported
Swix: Amjad Masad is Building Replit APIs While Competing in AI Models
“Amjad has specifically told me in person that he's, he wants to enable that for people. At the same time, he's also working on his own models. And Ghost Rider and, you know, all the other stuff. Yeah. So it's gonna get interesting, like, he wants to power you,…”
Shawn Wang Nov 3, 2023 ▶ 49:50
Insight
Royzen: NVIDIA Remains Cloud-Agnostic Because It Wins Regardless
“At NVIDIA, They know that they're going to win regardless. So they don't care where you get the GPUs from. They're like, they're truly neutral, unlike various sales reps that you might encounter at various like clouds and, you know, hardware companies, et cete…”
Michael Royzen Nov 3, 2023 ▶ 1:02:39
Assertion Not publicly verifiable
Royzen: NVIDIA Built Custom FasterTransformer Feature for Phind
“They actually implemented a custom feature for us in Faster Transformer which is one of their libraries... They implemented streaming generation for T-Five-based models, which we were running at the time up until we switched to GPT in In February, March of thi…”
Michael Royzen Nov 3, 2023 ▶ 1:04:04
Assertion Supported
Royzen: INT8 quantization offers storage optimization without guaranteed inference speedups
“But with int eight, there's not necessarily a Speed increase. It's just the storage optimization.”
Michael Royzen Nov 3, 2023 ▶ 1:05:46
Assertion Supported
Royzen: Quantized LLMs currently underperform unquantized baselines in quality
“So we have these great quantization libraries that, you know, for the most part are able to get the size down with not that much quality loss, but there is some, like the quantized models currently are actually worse than the non-quantized ones.”
Michael Royzen Nov 3, 2023 ▶ 1:05:53
Assertion Supported
Swix: OpenAI is adopting the 'Applied AI Engineer' title
“Well, for what it's worth, OpenAI is adopting Applied AI Engineer.”
Shawn Wang Nov 3, 2023 ▶ 1:11:42
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.