Aug 17, 2024 · 1h 10m · latent-space

Answer.ai & AI Magic with Jeremy Howard

Jeremy Howard · 52m spoken Swix (Shawn) · 4m spoken
0:00 / 0:00

⌖ your search result is the highlighted band (46:00–46:26). Playback starts there

▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this Latent Space podcast episode, Jeremy Howard discusses Answer.ai's Public Benefit Corporation mission, breaks down systems engineering breakthroughs for local fine-tuning, challenges decoder-only model dogmas, and introduces new developer tools like FastHTML and Dialogue Engineering.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 5.9 Guest teaching 6.1 Guest disagreement 3.6 The hosts pushing back 2.0
05100:0015:0030:0045:001:00:000:43–3:19 · The hosts as informed peer 6/10 Rethinking Fine-Tuning as a Pre-Training Continuum Swix references Jeremy's past hot take on the death of fine-tuning and cites an ICLR paper on data-driven priors. Jeremy clarifies that fine-tuning is part of a pre-training continuum and rejects starting from random weights.3:20–6:06 · The hosts as informed peer 7/10 Multi-Phase Pre-Training and Optimization Schedules Swix introduces Snowflake Arctic's multi-phase pre-training and Meta's schedule-free optimizers. Jeremy bluntly dismisses schedule-free approaches, advocating for functional learning rate schedules and maximum parameter control.6:09–14:06 · The hosts as informed peer 6/10 AI Governance, OpenAI Dynamics, and Public Benefit Corporations Alessio and Jeremy explore OpenAI's governance implosion and the mechanics of Public Benefit Corporations. Jeremy breaks down why non-profit boards overseeing equity-driven commercial entities inevitably fail.14:07–27:04 · The hosts as informed peer 5/10 Answer.ai's Organizational Structure and Self-Directed Projects Jeremy explains Answer.ai's flat structure and organic project origination, citing spontaneous developments like BERT-24 and WebGPU tooling.27:04–32:02 · The hosts as informed peer 5/10 Evaluating Practical Talent Over Academic Credentials Alessio asks about screening practical talent versus traditional credentials. Jeremy highlights hiring unconventional problem-solvers based on code quality and proven tenacity rather than academic pedigree.32:03–37:08 · The hosts as informed peer 6/10 Architectural Bets: The Resurgence of Encoder-Decoder Models Swix queries the non-consensus bet on encoder-decoder models. Jeremy explains why decoder-only models are over-indexed and structurally inefficient for classification and fixed representation tasks.37:10–43:41 · The hosts as informed peer 6/10 Engineering Challenges in FSDP, QLoRA, and QDoRA Implementation Alessio asks about fine-tuning 70B models on consumer hardware. Jeremy details the grueling systems engineering required to unify FSDP, bitsandbytes, and PEFT into a coherent workflow.43:42–47:26 · The hosts as informed peer 7/10 Local Inference Optimization, Quantization, and Merged Adapters Swix pushes back on Jeremy's assertion that nobody should merge models by citing zero-data merging benefits. Jeremy forcefully clarifies that distributing merged full weights rather than merged adapters is the anti-pattern.47:27–56:44 · The hosts as informed peer 5/10 FastHTML: Bringing Web Fundamentals to Pure Python Development Jeremy unveils FastHTML, critiquing modern JavaScript SPA complexity and advocating for pure Python web development built directly on web standards and HTMX.56:44–1:04:13 · The hosts as informed peer 6/10 Dialogue Engineering, AI Magic, and Practical Developer Tooling Swix asks about AI-assisted development tools. Jeremy introduces dialogue engineering and AI Magic as an interactive middle ground between basic chat teletypes and complex IDEs.0:43–3:19 · Guest teaching 6/10 Rethinking Fine-Tuning as a Pre-Training Continuum Swix references Jeremy's past hot take on the death of fine-tuning and cites an ICLR paper on data-driven priors. Jeremy clarifies that fine-tuning is part of a pre-training continuum and rejects starting from random weights.3:20–6:06 · Guest teaching 6/10 Multi-Phase Pre-Training and Optimization Schedules Swix introduces Snowflake Arctic's multi-phase pre-training and Meta's schedule-free optimizers. Jeremy bluntly dismisses schedule-free approaches, advocating for functional learning rate schedules and maximum parameter control.6:09–14:06 · Guest teaching 7/10 AI Governance, OpenAI Dynamics, and Public Benefit Corporations Alessio and Jeremy explore OpenAI's governance implosion and the mechanics of Public Benefit Corporations. Jeremy breaks down why non-profit boards overseeing equity-driven commercial entities inevitably fail.14:07–27:04 · Guest teaching 5/10 Answer.ai's Organizational Structure and Self-Directed Projects Jeremy explains Answer.ai's flat structure and organic project origination, citing spontaneous developments like BERT-24 and WebGPU tooling.27:04–32:02 · Guest teaching 5/10 Evaluating Practical Talent Over Academic Credentials Alessio asks about screening practical talent versus traditional credentials. Jeremy highlights hiring unconventional problem-solvers based on code quality and proven tenacity rather than academic pedigree.32:03–37:08 · Guest teaching 7/10 Architectural Bets: The Resurgence of Encoder-Decoder Models Swix queries the non-consensus bet on encoder-decoder models. Jeremy explains why decoder-only models are over-indexed and structurally inefficient for classification and fixed representation tasks.37:10–43:41 · Guest teaching 6/10 Engineering Challenges in FSDP, QLoRA, and QDoRA Implementation Alessio asks about fine-tuning 70B models on consumer hardware. Jeremy details the grueling systems engineering required to unify FSDP, bitsandbytes, and PEFT into a coherent workflow.43:42–47:26 · Guest teaching 7/10 Local Inference Optimization, Quantization, and Merged Adapters Swix pushes back on Jeremy's assertion that nobody should merge models by citing zero-data merging benefits. Jeremy forcefully clarifies that distributing merged full weights rather than merged adapters is the anti-pattern.47:27–56:44 · Guest teaching 6/10 FastHTML: Bringing Web Fundamentals to Pure Python Development Jeremy unveils FastHTML, critiquing modern JavaScript SPA complexity and advocating for pure Python web development built directly on web standards and HTMX.56:44–1:04:13 · Guest teaching 6/10 Dialogue Engineering, AI Magic, and Practical Developer Tooling Swix asks about AI-assisted development tools. Jeremy introduces dialogue engineering and AI Magic as an interactive middle ground between basic chat teletypes and complex IDEs.0:43–3:19 · Guest disagreement 3/10 Rethinking Fine-Tuning as a Pre-Training Continuum Swix references Jeremy's past hot take on the death of fine-tuning and cites an ICLR paper on data-driven priors. Jeremy clarifies that fine-tuning is part of a pre-training continuum and rejects starting from random weights.3:20–6:06 · Guest disagreement 5/10 Multi-Phase Pre-Training and Optimization Schedules Swix introduces Snowflake Arctic's multi-phase pre-training and Meta's schedule-free optimizers. Jeremy bluntly dismisses schedule-free approaches, advocating for functional learning rate schedules and maximum parameter control.6:09–14:06 · Guest disagreement 4/10 AI Governance, OpenAI Dynamics, and Public Benefit Corporations Alessio and Jeremy explore OpenAI's governance implosion and the mechanics of Public Benefit Corporations. Jeremy breaks down why non-profit boards overseeing equity-driven commercial entities inevitably fail.14:07–27:04 · Guest disagreement 2/10 Answer.ai's Organizational Structure and Self-Directed Projects Jeremy explains Answer.ai's flat structure and organic project origination, citing spontaneous developments like BERT-24 and WebGPU tooling.27:04–32:02 · Guest disagreement 3/10 Evaluating Practical Talent Over Academic Credentials Alessio asks about screening practical talent versus traditional credentials. Jeremy highlights hiring unconventional problem-solvers based on code quality and proven tenacity rather than academic pedigree.32:03–37:08 · Guest disagreement 4/10 Architectural Bets: The Resurgence of Encoder-Decoder Models Swix queries the non-consensus bet on encoder-decoder models. Jeremy explains why decoder-only models are over-indexed and structurally inefficient for classification and fixed representation tasks.37:10–43:41 · Guest disagreement 3/10 Engineering Challenges in FSDP, QLoRA, and QDoRA Implementation Alessio asks about fine-tuning 70B models on consumer hardware. Jeremy details the grueling systems engineering required to unify FSDP, bitsandbytes, and PEFT into a coherent workflow.43:42–47:26 · Guest disagreement 6/10 Local Inference Optimization, Quantization, and Merged Adapters Swix pushes back on Jeremy's assertion that nobody should merge models by citing zero-data merging benefits. Jeremy forcefully clarifies that distributing merged full weights rather than merged adapters is the anti-pattern.47:27–56:44 · Guest disagreement 3/10 FastHTML: Bringing Web Fundamentals to Pure Python Development Jeremy unveils FastHTML, critiquing modern JavaScript SPA complexity and advocating for pure Python web development built directly on web standards and HTMX.56:44–1:04:13 · Guest disagreement 3/10 Dialogue Engineering, AI Magic, and Practical Developer Tooling Swix asks about AI-assisted development tools. Jeremy introduces dialogue engineering and AI Magic as an interactive middle ground between basic chat teletypes and complex IDEs.0:43–3:19 · The hosts pushing back 2/10 Rethinking Fine-Tuning as a Pre-Training Continuum Swix references Jeremy's past hot take on the death of fine-tuning and cites an ICLR paper on data-driven priors. Jeremy clarifies that fine-tuning is part of a pre-training continuum and rejects starting from random weights.3:20–6:06 · The hosts pushing back 3/10 Multi-Phase Pre-Training and Optimization Schedules Swix introduces Snowflake Arctic's multi-phase pre-training and Meta's schedule-free optimizers. Jeremy bluntly dismisses schedule-free approaches, advocating for functional learning rate schedules and maximum parameter control.6:09–14:06 · The hosts pushing back 1/10 AI Governance, OpenAI Dynamics, and Public Benefit Corporations Alessio and Jeremy explore OpenAI's governance implosion and the mechanics of Public Benefit Corporations. Jeremy breaks down why non-profit boards overseeing equity-driven commercial entities inevitably fail.14:07–27:04 · The hosts pushing back 1/10 Answer.ai's Organizational Structure and Self-Directed Projects Jeremy explains Answer.ai's flat structure and organic project origination, citing spontaneous developments like BERT-24 and WebGPU tooling.27:04–32:02 · The hosts pushing back 1/10 Evaluating Practical Talent Over Academic Credentials Alessio asks about screening practical talent versus traditional credentials. Jeremy highlights hiring unconventional problem-solvers based on code quality and proven tenacity rather than academic pedigree.32:03–37:08 · The hosts pushing back 2/10 Architectural Bets: The Resurgence of Encoder-Decoder Models Swix queries the non-consensus bet on encoder-decoder models. Jeremy explains why decoder-only models are over-indexed and structurally inefficient for classification and fixed representation tasks.37:10–43:41 · The hosts pushing back 1/10 Engineering Challenges in FSDP, QLoRA, and QDoRA Implementation Alessio asks about fine-tuning 70B models on consumer hardware. Jeremy details the grueling systems engineering required to unify FSDP, bitsandbytes, and PEFT into a coherent workflow.43:42–47:26 · The hosts pushing back 6/10 Local Inference Optimization, Quantization, and Merged Adapters Swix pushes back on Jeremy's assertion that nobody should merge models by citing zero-data merging benefits. Jeremy forcefully clarifies that distributing merged full weights rather than merged adapters is the anti-pattern.47:27–56:44 · The hosts pushing back 1/10 FastHTML: Bringing Web Fundamentals to Pure Python Development Jeremy unveils FastHTML, critiquing modern JavaScript SPA complexity and advocating for pure Python web development built directly on web standards and HTMX.56:44–1:04:13 · The hosts pushing back 2/10 Dialogue Engineering, AI Magic, and Practical Developer Tooling Swix asks about AI-assisted development tools. Jeremy introduces dialogue engineering and AI Magic as an interactive middle ground between basic chat teletypes and complex IDEs.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:03:00 · the hosts 0% · guest 100%1:03:00 · the hosts 0% · guest 100%1:06:00 · the hosts 0% · guest 100%1:06:00 · the hosts 0% · guest 100%1:09:00 · the hosts 0% · guest 100%1:09:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 46:25 Jeremy's sharp correction on distributing merged models

Jeremy directly rejects Swix's premise regarding model merging, firmly clarifying that distributing merged full weight models rather than quantized bases with merged adapters is inefficient.

Hardest push from the hosts ▶ 46:02 Swix challenges Jeremy's dismissal of model merging

Swix openly defends model merging against Jeremy's categorical assertion, pointing out that merging models provides unique value when working without training data.

Biggest teaching moment ▶ 7:28 Structural breakdown of non-profit board misalignments

Jeremy methodically dissects why academic non-profit boards cannot effectively govern commercial entities when employee compensation is tied directly to equity and profit.

The host holds their own ▶ 3:20 Swix frames multi-phase pre-training trends

Swix demonstrates deep industry monitoring by citing Snowflake Arctic's exact web-to-code dataset mix transitions across training phases.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Rethinking Fine-Tuning as a Pre-Training Continuum 6632 Swix references Jeremy's past hot take on the death of fine-tuning and cites an ICLR paper on data-driven priors. Jeremy clarifies that fine-tuning is part of a pre-training continuum and rejects starting from random weights.
Multi-Phase Pre-Training and Optimization Schedules 7653 Swix introduces Snowflake Arctic's multi-phase pre-training and Meta's schedule-free optimizers. Jeremy bluntly dismisses schedule-free approaches, advocating for functional learning rate schedules and maximum parameter control.
AI Governance, OpenAI Dynamics, and Public Benefit Corporations 6741 Alessio and Jeremy explore OpenAI's governance implosion and the mechanics of Public Benefit Corporations. Jeremy breaks down why non-profit boards overseeing equity-driven commercial entities inevitably fail.
Answer.ai's Organizational Structure and Self-Directed Projects 5521 Jeremy explains Answer.ai's flat structure and organic project origination, citing spontaneous developments like BERT-24 and WebGPU tooling.
Evaluating Practical Talent Over Academic Credentials 5531 Alessio asks about screening practical talent versus traditional credentials. Jeremy highlights hiring unconventional problem-solvers based on code quality and proven tenacity rather than academic pedigree.
Architectural Bets: The Resurgence of Encoder-Decoder Models 6742 Swix queries the non-consensus bet on encoder-decoder models. Jeremy explains why decoder-only models are over-indexed and structurally inefficient for classification and fixed representation tasks.
Engineering Challenges in FSDP, QLoRA, and QDoRA Implementation 6631 Alessio asks about fine-tuning 70B models on consumer hardware. Jeremy details the grueling systems engineering required to unify FSDP, bitsandbytes, and PEFT into a coherent workflow.
Local Inference Optimization, Quantization, and Merged Adapters 7766 Swix pushes back on Jeremy's assertion that nobody should merge models by citing zero-data merging benefits. Jeremy forcefully clarifies that distributing merged full weights rather than merged adapters is the anti-pattern.
FastHTML: Bringing Web Fundamentals to Pure Python Development 5631 Jeremy unveils FastHTML, critiquing modern JavaScript SPA complexity and advocating for pure Python web development built directly on web standards and HTMX.
Dialogue Engineering, AI Magic, and Practical Developer Tooling 6632 Swix asks about AI-assisted development tools. Jeremy introduces dialogue engineering and AI Magic as an interactive middle ground between basic chat teletypes and complex IDEs.

Statements from this episode (25)

Insight
Howard: Training stages form a continuum allowing deep modification of pre-trained models
“Sorry, it wasn't the end of fine-tuning, but more that we should treat it as a continuum, and we should have much higher expectations of how much you can do with an already trained model. You can really add a lot of behavior to it. You can change its behavior.…”
Jeremy Howard Aug 17, 2024 ▶ 2:04
Insight
Howard: Training AI models from random weights is almost never justified
“If you're training for random weights, you better have a really good reason, you know, because it seems so unlikely to me that nobody has ever trained on data that has any similarity whatsoever to the general class of data you're working with, and that's the o…”
Jeremy Howard Aug 17, 2024 ▶ 2:54
Insight
Howard: Pre-training data mixes should be continuous per-batch functions, not discrete phases
“So the point at which they're doing proper continued pre-training is the point at which that becomes a continuum rather than a phase. So the only difference with what I was describing last time is to say, like, oh, they should, you know, There's a function or …”
Jeremy Howard Aug 17, 2024 ▶ 4:08
Opinion
Howard: Meta's schedule-free optimizers are not that exciting
“I mean, I don't care very much, honestly. Like, I don't think that schedule-free optimizer's that exciting. It's fine.”
Jeremy Howard Aug 17, 2024 ▶ 5:01
Insight
Howard: Non-profit boards cannot control commercial entities with equity-compensated staff
“This didn't make sense to have like a so-called non-profit where then there are people working at a commercial company that's owned by or controlled nominally by the non-profit where the people in the company are being given the equivalent of stock options. Li…”
Jeremy Howard Aug 17, 2024 ▶ 7:55
Insight
Howard: Corporations are sociopathic by design due to fiduciary duty
“Companies are sociopathic, like, by design. And so the alignment problem, as it relates to companies, has not been solved. Like, companies become huge, they devour their founders, they devour their communities, and they do things where even the CEOs, you know,…”
Jeremy Howard Aug 17, 2024 ▶ 9:09
Disclosure
Howard: Answer.ai required angel investors to commit to long-term voting principles
“When we did our second round, which was an angel round, we had everybody invest through a long-term SBV, which we set up. Where everybody had to agree to vote in line with long-term value principles.”
Jeremy Howard Aug 17, 2024 ▶ 11:05
Insight
Howard: Public Benefit Corporation status allows founders to reject hostile buyouts
“If you're not a public benefit corporation, Then somebody can come along and offer to buy you with a stated description of, like, turning your company into the thing you most hate, right? And if they offer you more than the market value of your company and you…”
Jeremy Howard Aug 17, 2024 ▶ 12:20
Assertion Not checkable as stated
Howard: 80% of top unique creators have unconventional or non-mainstream backgrounds
“Like, 80% of the time, I find out the person has a really unusual background. So, like, often they'll have, like, either they, like, came from poverty and, like, didn't get an opportunity to go to good school, or they, like, you know, had dyslexia and, you kno…”
Jeremy Howard Aug 17, 2024 ▶ 15:43
Disclosure
Howard: Answer.ai operates with no managers and zero corporate hierarchy
“We don't have any managers. We don't have any hierarchy from that point of view. So, for example, I'm not a manager, which means I don't get to tell people what to do or how to do it or when to do it.”
Jeremy Howard Aug 17, 2024 ▶ 19:28
Opinion
Howard: Tech builds too many vanity foundation models over fine-tuning
“People are building too many vanity foundation models rather than taking better advantage of fine-tuning”
Jeremy Howard Aug 17, 2024 ▶ 25:06
Opinion
Howard: Decoder models make little sense for non-generative tasks
“So anytime you're not generating an arbitrary length sequence of tokens decoder models don't seem to make much sense to me.”
Jeremy Howard Aug 17, 2024 ▶ 35:23
Assertion Not checkable as stated
Howard: Decoder models must be far larger to match DeBERTa
“Now, the interesting thing is, you see, unlike Kaggle competitions, that decoder models still Are at least competitive with things like DiBerta VIII. But they have to be way bigger to be competitive with things like DiBerta VIII. And the only reason they are c…”
Jeremy Howard Aug 17, 2024 ▶ 35:33
Prediction Not checkable as stated
Howard: Reka's model is probably superior to GPT and Claude for certain tasks
“There's a whole model that's been trained in a different way. So there's probably a whole lot of tasks it's probably better at than you know, GPT and Gemini and Claude.”
Jeremy Howard Aug 17, 2024 ▶ 36:42
Opinion
Jeremy Howard: Hugging Face libraries suffer from excessive coupling
“The hugging face library in peft doesn't really work in practice unless you use it with other things. And there's a lot of coupling in the hugging face ecosystem where, like, none of it works separately. You have to use it all together, which I don't love.”
Jeremy Howard Aug 17, 2024 ▶ 39:55
Opinion
Jeremy Howard: Open source AI lacks rigorous performance evaluations
“There's not a lot of, ah, really good Performance type evals going on in the open source ecosystem, so there's an extraordinary amount of, like, things where people say, like, oh, we built this thing, and it has this result, and when you actually check it does…”
Jeremy Howard Aug 17, 2024 ▶ 42:35
Insight
Howard: Developers should distribute merged adapters rather than merged models
“To explain, it's not that you shouldn't merge models, it's that you shouldn't be distributing a merged model. You should distribute it a merged adapter. 99% of the time. And actually often, one of the best things happening in the model merging world is actuall…”
Jeremy Howard Aug 17, 2024 ▶ 46:26
Disclosure
Howard: Answer.ai aims to build thousands of products with 12 people
“We want to create thousands of Commercially successful products at Answer.ai. And we want to do that with like, 12 people.”
Jeremy Howard Aug 17, 2024 ▶ 48:07
Opinion
Howard: Building web apps is much worse now than 15 years ago
“Much to my, you know, horror, the story around creating web applications is much worse now than it was 10 or 15 years ago, in terms of, like, if I say to a data scientist, here's how to create and deploy a web application, You know, either you have to learn Ja…”
Jeremy Howard Aug 17, 2024 ▶ 48:50
Assertion Supported
Howard: FastHTML uses web foundations directly unlike Streamlit or Gradio
“Unlike excellent projects like Streamlit and Gradio, you're not working on top of a highly abstracted thing that's got nothing to do with web foundations. You're working with web foundations directly, but you're able to do it by using pure Python.”
Jeremy Howard Aug 17, 2024 ▶ 50:40
Disclosure
Howard previews AI Magic, his upcoming dialogue engineering system
“So, I've created a new approach. It's not called prompt engineering. It's called dialogue engineering. And I'm creating a system for doing dialogue engineering. It's currently called AI Magic. I'm doing most of my work in this system, and it's making me much m…”
Jeremy Howard Aug 17, 2024 ▶ 57:10
Opinion
Howard: Cursor and VS Code shoehorn AI into legacy software paradigms
“It's like a convenience over the top of this incredibly complicated system that full-time, sophisticated software engineers have designed over the past few decades in a totally different environment as a way to build software, you know. And so we're trying to,…”
Jeremy Howard Aug 17, 2024 ▶ 59:58
Assertion Supported
Howard: Google Gemini is about to release KV caching support
“Gemini is about to finally come out with KV caching, and this is something that Austin actually and Gemma.cpp had had on his roadmap for years well not years, months, long time is, is that.”
Jeremy Howard Aug 17, 2024 ▶ 1:06:34
Prediction Not checkable as stated
Howard: AI developers will spend 12 months mapping RAG, fine-tuning, and KV caching
“Something over the next 12 months people will be spending time thinking about is how to, like, where to use RAG, where to use fine-tuning, where to use KV cache storage, you know, and how to use state.”
Jeremy Howard Aug 17, 2024 ▶ 1:08:05
Insight
Howard: Diffusion should be used to sketch answers before generating tokens
“The idea of, like, there should be a piece of the generative pipeline which is, like, thinking about the answer and coming up with a sketch of what the answer looks like before you start out putting tokens. That's where it kind of feels like diffusion ought to…”
Jeremy Howard Aug 17, 2024 ▶ 1:09:11
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.