Nov 8, 2023 · 42m · mad

Custom LLMs at Scale: Lamini CEO Sharon Zhou’s Playbook for Enterprise AI

Sharon Zhou · 30m spoken Matt Turck · 8m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of The MAD Podcast, host Matt Turck interviews Sharon Zhou, Co-Founder and CEO of Lamini, about fine-tuning large language models for enterprise deployment. Sharon discusses technical breakthroughs in model convergence, Parameter Efficient Fine-Tuning (PEFT), enterprise reliability, and Lamini's hardware partnership with AMD.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 21% of the talking time here. How this is scored →

Matt as informed peer 3.3 Guest teaching 4.6 Guest disagreement 0.9 Matt pushing back 1.5
05100:0015:0030:000:08–3:53 · Matt as informed peer 2/10 Sharon Zhou's Background and Journey to AI Matt warmly introduces Sharon Zhou and asks about her unique academic journey combining classics and computer science. Sharon explains her background at Google, Harvard, and Stanford under Andrew Ng, establishing her pedigree.3:53–8:09 · Matt as informed peer 4/10 Defining Pre-Training, Prompt Engineering, and Fine-Tuning Matt asks Sharon to differentiate pre-training, prompt engineering, and fine-tuning. Matt jumps in to clearly articulate pre-training as the multi-billion dollar base model creation phase before fine-tuning.8:09–10:15 · Matt as informed peer 2/10 Why Fine-Tuning is Difficult and Technical Solutions Matt asks why fine-tuning is technically challenging. Sharon educates him on model convergence, hyperparameter tuning, and reducing training duration from months to milliseconds.10:15–13:08 · Matt as informed peer 2/10 Lamini Platform Architecture and AMD GPU Integration Matt asks for a product tour of Lamini. Sharon outlines their software stack, data ingestion partners, SDKs, and their exclusive capability to run on AMD GPUs.13:08–22:01 · Matt as informed peer 5/10 RAG vs. RAFT and Model-Assisted RLHF in the Enterprise Matt probes the distinction between RAG and Retrieval-Augmented Fine-Tuning (RAFT), offering a layman summary. He also questions whether enterprise users must manually click thumbs up/down for RLHF, prompting Sharon to advocate for model-assisted feedback loops.22:01–28:52 · Matt as informed peer 5/10 Enterprise Model Scoping and Eliminating Hallucinations Matt pushes on whether fine-tuning eliminates hallucinations and asks if Sharon is solely in the small model camp. Sharon reframes model sizing and introduces parameter-efficient fine-tuning (PEFT), explaining how Lamini reduces model switching latency from 3 months to 3 milliseconds.28:52–32:54 · Matt as informed peer 5/10 AI Agents, Workflows, and Multi-Step Model Chains Matt asks about AI agent adoption in enterprises and specifically challenges whether chaining multiple models creates compounding error and latency. Sharon agrees and explains how fine-tuning a single unified model can consolidate complex chains.32:54–37:25 · Matt as informed peer 4/10 The AMD Partnership, Superstation, and Scaling Laws Matt asks about the AMD partnership and questions whether AMD hardware truly achieves parity with Nvidia CUDA. Sharon details her co-founder Greg Diamos's background in CUDA architecture and explains the core principles behind LLM scaling laws.37:25–41:05 · Matt as informed peer 3/10 Enterprise Go-to-Market, Customer Maturity, and Deployment Speed Matt inquires about customer maturity, go-to-market friction, and emerging enterprise use cases. Sharon notes rapid customer learning curves and emphasizes that companies must deploy quickly despite transitioning from deterministic to probabilistic software.41:05–42:20 · Matt as informed peer 1/10 Conclusion and Resource Links Matt wraps up the interview and asks where listeners can find Sharon and her online courses. Sharon provides her contact details and Coursera course links.0:08–3:53 · Guest teaching 2/10 Sharon Zhou's Background and Journey to AI Matt warmly introduces Sharon Zhou and asks about her unique academic journey combining classics and computer science. Sharon explains her background at Google, Harvard, and Stanford under Andrew Ng, establishing her pedigree.3:53–8:09 · Guest teaching 5/10 Defining Pre-Training, Prompt Engineering, and Fine-Tuning Matt asks Sharon to differentiate pre-training, prompt engineering, and fine-tuning. Matt jumps in to clearly articulate pre-training as the multi-billion dollar base model creation phase before fine-tuning.8:09–10:15 · Guest teaching 6/10 Why Fine-Tuning is Difficult and Technical Solutions Matt asks why fine-tuning is technically challenging. Sharon educates him on model convergence, hyperparameter tuning, and reducing training duration from months to milliseconds.10:15–13:08 · Guest teaching 5/10 Lamini Platform Architecture and AMD GPU Integration Matt asks for a product tour of Lamini. Sharon outlines their software stack, data ingestion partners, SDKs, and their exclusive capability to run on AMD GPUs.13:08–22:01 · Guest teaching 6/10 RAG vs. RAFT and Model-Assisted RLHF in the Enterprise Matt probes the distinction between RAG and Retrieval-Augmented Fine-Tuning (RAFT), offering a layman summary. He also questions whether enterprise users must manually click thumbs up/down for RLHF, prompting Sharon to advocate for model-assisted feedback loops.22:01–28:52 · Guest teaching 6/10 Enterprise Model Scoping and Eliminating Hallucinations Matt pushes on whether fine-tuning eliminates hallucinations and asks if Sharon is solely in the small model camp. Sharon reframes model sizing and introduces parameter-efficient fine-tuning (PEFT), explaining how Lamini reduces model switching latency from 3 months to 3 milliseconds.28:52–32:54 · Guest teaching 5/10 AI Agents, Workflows, and Multi-Step Model Chains Matt asks about AI agent adoption in enterprises and specifically challenges whether chaining multiple models creates compounding error and latency. Sharon agrees and explains how fine-tuning a single unified model can consolidate complex chains.32:54–37:25 · Guest teaching 6/10 The AMD Partnership, Superstation, and Scaling Laws Matt asks about the AMD partnership and questions whether AMD hardware truly achieves parity with Nvidia CUDA. Sharon details her co-founder Greg Diamos's background in CUDA architecture and explains the core principles behind LLM scaling laws.37:25–41:05 · Guest teaching 4/10 Enterprise Go-to-Market, Customer Maturity, and Deployment Speed Matt inquires about customer maturity, go-to-market friction, and emerging enterprise use cases. Sharon notes rapid customer learning curves and emphasizes that companies must deploy quickly despite transitioning from deterministic to probabilistic software.41:05–42:20 · Guest teaching 1/10 Conclusion and Resource Links Matt wraps up the interview and asks where listeners can find Sharon and her online courses. Sharon provides her contact details and Coursera course links.0:08–3:53 · Guest disagreement 0/10 Sharon Zhou's Background and Journey to AI Matt warmly introduces Sharon Zhou and asks about her unique academic journey combining classics and computer science. Sharon explains her background at Google, Harvard, and Stanford under Andrew Ng, establishing her pedigree.3:53–8:09 · Guest disagreement 1/10 Defining Pre-Training, Prompt Engineering, and Fine-Tuning Matt asks Sharon to differentiate pre-training, prompt engineering, and fine-tuning. Matt jumps in to clearly articulate pre-training as the multi-billion dollar base model creation phase before fine-tuning.8:09–10:15 · Guest disagreement 1/10 Why Fine-Tuning is Difficult and Technical Solutions Matt asks why fine-tuning is technically challenging. Sharon educates him on model convergence, hyperparameter tuning, and reducing training duration from months to milliseconds.10:15–13:08 · Guest disagreement 0/10 Lamini Platform Architecture and AMD GPU Integration Matt asks for a product tour of Lamini. Sharon outlines their software stack, data ingestion partners, SDKs, and their exclusive capability to run on AMD GPUs.13:08–22:01 · Guest disagreement 2/10 RAG vs. RAFT and Model-Assisted RLHF in the Enterprise Matt probes the distinction between RAG and Retrieval-Augmented Fine-Tuning (RAFT), offering a layman summary. He also questions whether enterprise users must manually click thumbs up/down for RLHF, prompting Sharon to advocate for model-assisted feedback loops.22:01–28:52 · Guest disagreement 2/10 Enterprise Model Scoping and Eliminating Hallucinations Matt pushes on whether fine-tuning eliminates hallucinations and asks if Sharon is solely in the small model camp. Sharon reframes model sizing and introduces parameter-efficient fine-tuning (PEFT), explaining how Lamini reduces model switching latency from 3 months to 3 milliseconds.28:52–32:54 · Guest disagreement 1/10 AI Agents, Workflows, and Multi-Step Model Chains Matt asks about AI agent adoption in enterprises and specifically challenges whether chaining multiple models creates compounding error and latency. Sharon agrees and explains how fine-tuning a single unified model can consolidate complex chains.32:54–37:25 · Guest disagreement 1/10 The AMD Partnership, Superstation, and Scaling Laws Matt asks about the AMD partnership and questions whether AMD hardware truly achieves parity with Nvidia CUDA. Sharon details her co-founder Greg Diamos's background in CUDA architecture and explains the core principles behind LLM scaling laws.37:25–41:05 · Guest disagreement 1/10 Enterprise Go-to-Market, Customer Maturity, and Deployment Speed Matt inquires about customer maturity, go-to-market friction, and emerging enterprise use cases. Sharon notes rapid customer learning curves and emphasizes that companies must deploy quickly despite transitioning from deterministic to probabilistic software.41:05–42:20 · Guest disagreement 0/10 Conclusion and Resource Links Matt wraps up the interview and asks where listeners can find Sharon and her online courses. Sharon provides her contact details and Coursera course links.0:08–3:53 · Matt pushing back 0/10 Sharon Zhou's Background and Journey to AI Matt warmly introduces Sharon Zhou and asks about her unique academic journey combining classics and computer science. Sharon explains her background at Google, Harvard, and Stanford under Andrew Ng, establishing her pedigree.3:53–8:09 · Matt pushing back 1/10 Defining Pre-Training, Prompt Engineering, and Fine-Tuning Matt asks Sharon to differentiate pre-training, prompt engineering, and fine-tuning. Matt jumps in to clearly articulate pre-training as the multi-billion dollar base model creation phase before fine-tuning.8:09–10:15 · Matt pushing back 0/10 Why Fine-Tuning is Difficult and Technical Solutions Matt asks why fine-tuning is technically challenging. Sharon educates him on model convergence, hyperparameter tuning, and reducing training duration from months to milliseconds.10:15–13:08 · Matt pushing back 0/10 Lamini Platform Architecture and AMD GPU Integration Matt asks for a product tour of Lamini. Sharon outlines their software stack, data ingestion partners, SDKs, and their exclusive capability to run on AMD GPUs.13:08–22:01 · Matt pushing back 4/10 RAG vs. RAFT and Model-Assisted RLHF in the Enterprise Matt probes the distinction between RAG and Retrieval-Augmented Fine-Tuning (RAFT), offering a layman summary. He also questions whether enterprise users must manually click thumbs up/down for RLHF, prompting Sharon to advocate for model-assisted feedback loops.22:01–28:52 · Matt pushing back 3/10 Enterprise Model Scoping and Eliminating Hallucinations Matt pushes on whether fine-tuning eliminates hallucinations and asks if Sharon is solely in the small model camp. Sharon reframes model sizing and introduces parameter-efficient fine-tuning (PEFT), explaining how Lamini reduces model switching latency from 3 months to 3 milliseconds.28:52–32:54 · Matt pushing back 3/10 AI Agents, Workflows, and Multi-Step Model Chains Matt asks about AI agent adoption in enterprises and specifically challenges whether chaining multiple models creates compounding error and latency. Sharon agrees and explains how fine-tuning a single unified model can consolidate complex chains.32:54–37:25 · Matt pushing back 2/10 The AMD Partnership, Superstation, and Scaling Laws Matt asks about the AMD partnership and questions whether AMD hardware truly achieves parity with Nvidia CUDA. Sharon details her co-founder Greg Diamos's background in CUDA architecture and explains the core principles behind LLM scaling laws.37:25–41:05 · Matt pushing back 2/10 Enterprise Go-to-Market, Customer Maturity, and Deployment Speed Matt inquires about customer maturity, go-to-market friction, and emerging enterprise use cases. Sharon notes rapid customer learning curves and emphasizes that companies must deploy quickly despite transitioning from deterministic to probabilistic software.41:05–42:20 · Matt pushing back 0/10 Conclusion and Resource Links Matt wraps up the interview and asks where listeners can find Sharon and her online courses. Sharon provides her contact details and Coursera course links.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 28.6% · guest 71.4%0:00 · Matt 28.6% · guest 71.4%3:00 · Matt 21% · guest 79%3:00 · Matt 21% · guest 79%6:00 · Matt 16.5% · guest 83.5%6:00 · Matt 16.5% · guest 83.5%9:00 · Matt 20.2% · guest 79.8%9:00 · Matt 20.2% · guest 79.8%12:00 · Matt 24.3% · guest 75.7%12:00 · Matt 24.3% · guest 75.7%15:00 · Matt 37.5% · guest 62.5%15:00 · Matt 37.5% · guest 62.5%18:00 · Matt 4.8% · guest 95.2%18:00 · Matt 4.8% · guest 95.2%21:00 · Matt 19.7% · guest 80.3%21:00 · Matt 19.7% · guest 80.3%24:00 · Matt 27.6% · guest 72.4%24:00 · Matt 27.6% · guest 72.4%27:00 · Matt 14.5% · guest 85.5%27:00 · Matt 14.5% · guest 85.5%30:00 · Matt 17.8% · guest 82.2%30:00 · Matt 17.8% · guest 82.2%33:00 · Matt 9.3% · guest 90.7%33:00 · Matt 9.3% · guest 90.7%36:00 · Matt 31.6% · guest 68.4%36:00 · Matt 31.6% · guest 68.4%39:00 · Matt 22.3% · guest 77.7%39:00 · Matt 22.3% · guest 77.7%42:00 · Matt 14.4% · guest 85.6%42:00 · Matt 14.4% · guest 85.6%
Sharpest disagreement ▶ 15:55 Sharon rejects human-centric RLHF framing

Sharon directly challenges the standard industry term RLHF by questioning the 'H', arguing that human feedback is tedious and unnecessary when model pipelines can generate superior feedback.

Hardest push from Matt ▶ 17:25 Matt questions manual RLHF enterprise workflow

Matt refuses to accept the premise that enterprise employees will sit at desks giving thumbs up and down to models, pressing Sharon on the practical reality of UI/UX in corporate RLHF.

Biggest teaching moment ▶ 27:10 Sharon breaks down 3-month vs 3-millisecond model switching

Sharon educates the host on GPU switching overhead, demonstrating how naive multi-tenant model switching takes three months while parameter-efficient fine-tuning cuts it down to three milliseconds.

Matt holds his own ▶ 31:35 Matt points out compounding error in LLM chains

Matt demonstrates sharp technical intuition by asking whether chaining model to model introduces compounding errors and latency issues compared to standard API calls.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Sharon Zhou's Background and Journey to AI 2200 Matt warmly introduces Sharon Zhou and asks about her unique academic journey combining classics and computer science. Sharon explains her background at Google, Harvard, and Stanford under Andrew Ng, establishing her pedigree.
Defining Pre-Training, Prompt Engineering, and Fine-Tuning 4511 Matt asks Sharon to differentiate pre-training, prompt engineering, and fine-tuning. Matt jumps in to clearly articulate pre-training as the multi-billion dollar base model creation phase before fine-tuning.
Why Fine-Tuning is Difficult and Technical Solutions 2610 Matt asks why fine-tuning is technically challenging. Sharon educates him on model convergence, hyperparameter tuning, and reducing training duration from months to milliseconds.
Lamini Platform Architecture and AMD GPU Integration 2500 Matt asks for a product tour of Lamini. Sharon outlines their software stack, data ingestion partners, SDKs, and their exclusive capability to run on AMD GPUs.
RAG vs. RAFT and Model-Assisted RLHF in the Enterprise 5624 Matt probes the distinction between RAG and Retrieval-Augmented Fine-Tuning (RAFT), offering a layman summary. He also questions whether enterprise users must manually click thumbs up/down for RLHF, prompting Sharon to advocate for model-assisted feedback loops.
Enterprise Model Scoping and Eliminating Hallucinations 5623 Matt pushes on whether fine-tuning eliminates hallucinations and asks if Sharon is solely in the small model camp. Sharon reframes model sizing and introduces parameter-efficient fine-tuning (PEFT), explaining how Lamini reduces model switching latency from 3 months to 3 milliseconds.
AI Agents, Workflows, and Multi-Step Model Chains 5513 Matt asks about AI agent adoption in enterprises and specifically challenges whether chaining multiple models creates compounding error and latency. Sharon agrees and explains how fine-tuning a single unified model can consolidate complex chains.
The AMD Partnership, Superstation, and Scaling Laws 4612 Matt asks about the AMD partnership and questions whether AMD hardware truly achieves parity with Nvidia CUDA. Sharon details her co-founder Greg Diamos's background in CUDA architecture and explains the core principles behind LLM scaling laws.
Enterprise Go-to-Market, Customer Maturity, and Deployment Speed 3412 Matt inquires about customer maturity, go-to-market friction, and emerging enterprise use cases. Sharon notes rapid customer learning curves and emphasizes that companies must deploy quickly despite transitioning from deterministic to probabilistic software.
Conclusion and Resource Links 1100 Matt wraps up the interview and asks where listeners can find Sharon and her online courses. Sharon provides her contact details and Coursera course links.

Statements from this episode (15)

Opinion
Sharon Zhou: LLMs are the new IP
“LLMs are, I believe, the new IP.”
Sharon Zhou Nov 8, 2023 ▶ 3:45
Assertion Supported
Zhou: Fine-tuning transformed GPT-3 into ChatGPT
“Fine tuning is the technology that got from a research project in 2020 called GPT-III and turned that into ChatGPT, a billion dollar app, right?”
Sharon Zhou Nov 8, 2023 ▶ 4:30
Insight
Zhou: Editing search queries on Google is essentially prompt engineering
“I think Google is essentially prompt engineering. When you edit your query to get the results that you want, that's prompt engineering.”
Sharon Zhou Nov 8, 2023 ▶ 5:29
Assertion Not checkable as stated
Zhou: Lamini cuts LLM fine-tuning time from months to milliseconds
“And by efficiency, I mean, you know, it's instead of something that might take weeks or even months that's bringing it down to even like the millisecond level.”
Sharon Zhou Nov 8, 2023 ▶ 9:56
Assertion Contradicted
Zhou: Lamini is the only platform running LLMs on AMD GPUs
“We are the only folks who can actually run your language models on top of AMD AMD GPUs.”
Sharon Zhou Nov 8, 2023 ▶ 12:37
Assertion Contradicted
Lamini's AMD support unlocks 20,000 GPUs, enough to train GPT-4
“So that unlocks, what that means is that unlocks about 20,000 GPUs readily available today for enterprises to be able to use, and to get a sense of what that means you can train GBD-IV.”
Sharon Zhou Nov 8, 2023 ▶ 12:55
Insight
Zhou: Zero-shot LLMs can replace manual human labeling in RLHF workflows
“Which is that why can't it be another LLM or a pipeline of LLMs that can help with that feedback? I think manual labeling is very tedious, especially for our target user, which is a software engineer. And I don't think people should necessarily have to do all …”
Sharon Zhou Nov 8, 2023 ▶ 16:43
Assertion Not checkable as stated
Zhou: Outsourcing specialized medical AI data labeling to Scale AI failed
“We tried outsourcing actually to like scale AI, et cetera. None of that worked. It had to basically be me.”
Sharon Zhou Nov 8, 2023 ▶ 21:38
Assertion Not checkable as stated
Zhou: LLMs can reach 99% accuracy today with narrow scoping
“I think we can get to that performance today, but it's based on how you scope out the problem. So if it's a very narrow scope, of course you can get that.”
Sharon Zhou Nov 8, 2023 ▶ 22:33
Insight
Zhou: ChatGPT guardrails are difficult due to broad scope
“The reason why ChatGPT is hard to put guardrails on is because they're trying to go after every possible use case, right?”
Sharon Zhou Nov 8, 2023 ▶ 22:40
Disclosure
Zhou: Lamini can train models up to 100 billion parameters
“We can train up to a hundred billion parameters.”
Sharon Zhou Nov 8, 2023 ▶ 24:34
Assertion Open · timeframe Nov 2026
Lamini switches across 1,000 fine-tuned models in three milliseconds
“With the technology that we've used with parameter efficient fine tuning and just like efficiency, different efficiency methods, that time to switch across a thousand models is three milliseconds”
Sharon Zhou Nov 8, 2023 ▶ 27:59
Insight
Sharon Zhou: Compounding chain errors can be solved by fine-tuning one model
“The ways to reduce it are you take the input of the first thing into the chain, you take the output of the last thing, and you fine tune only one model to do the whole thing, for example.”
Sharon Zhou Nov 8, 2023 ▶ 32:05
Disclosure
Lamini's hosted service ran exclusively on AMD GPUs for a year
“The Lamini hosted service over the past year has been running on AMD GPUs only. We haven't been running on NVIDIA chips.”
Sharon Zhou Nov 8, 2023 ▶ 33:01
Assertion Not checkable as stated
Lamini has achieved software parity on AMD GPUs with CUDA
“We have reached software parity with essentially CUDA.”
Sharon Zhou Nov 8, 2023 ▶ 37:03
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.