Nov 8, 2023 · 42m · mad
Custom LLMs at Scale: Lamini CEO Sharon Zhou’s Playbook for Enterprise AI
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this episode of The MAD Podcast, host Matt Turck interviews Sharon Zhou, Co-Founder and CEO of Lamini, about fine-tuning large language models for enterprise deployment. Sharon discusses technical breakthroughs in model convergence, Parameter Efficient Fine-Tuning (PEFT), enterprise reliability, and Lamini's hardware partnership with AMD.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 21% of the talking time here. How this is scored →
speaking balance: gold is Matt, purple is the guest (3 minute bins)
Sharon directly challenges the standard industry term RLHF by questioning the 'H', arguing that human feedback is tedious and unnecessary when model pipelines can generate superior feedback.
Hardest push from Matt ▶ 17:25 Matt questions manual RLHF enterprise workflowMatt refuses to accept the premise that enterprise employees will sit at desks giving thumbs up and down to models, pressing Sharon on the practical reality of UI/UX in corporate RLHF.
Biggest teaching moment ▶ 27:10 Sharon breaks down 3-month vs 3-millisecond model switchingSharon educates the host on GPU switching overhead, demonstrating how naive multi-tenant model switching takes three months while parameter-efficient fine-tuning cuts it down to three milliseconds.
Matt holds his own ▶ 31:35 Matt points out compounding error in LLM chainsMatt demonstrates sharp technical intuition by asking whether chaining model to model introduces compounding errors and latency issues compared to standard API calls.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | Matt as informed peer | Guest teaching | Guest disagreement | Matt pushing back | Why |
|---|---|---|---|---|---|---|
| Sharon Zhou's Background and Journey to AI | 2 | 2 | 0 | 0 | Matt warmly introduces Sharon Zhou and asks about her unique academic journey combining classics and computer science. Sharon explains her background at Google, Harvard, and Stanford under Andrew Ng, establishing her pedigree. | |
| Defining Pre-Training, Prompt Engineering, and Fine-Tuning | 4 | 5 | 1 | 1 | Matt asks Sharon to differentiate pre-training, prompt engineering, and fine-tuning. Matt jumps in to clearly articulate pre-training as the multi-billion dollar base model creation phase before fine-tuning. | |
| Why Fine-Tuning is Difficult and Technical Solutions | 2 | 6 | 1 | 0 | Matt asks why fine-tuning is technically challenging. Sharon educates him on model convergence, hyperparameter tuning, and reducing training duration from months to milliseconds. | |
| Lamini Platform Architecture and AMD GPU Integration | 2 | 5 | 0 | 0 | Matt asks for a product tour of Lamini. Sharon outlines their software stack, data ingestion partners, SDKs, and their exclusive capability to run on AMD GPUs. | |
| RAG vs. RAFT and Model-Assisted RLHF in the Enterprise | 5 | 6 | 2 | 4 | Matt probes the distinction between RAG and Retrieval-Augmented Fine-Tuning (RAFT), offering a layman summary. He also questions whether enterprise users must manually click thumbs up/down for RLHF, prompting Sharon to advocate for model-assisted feedback loops. | |
| Enterprise Model Scoping and Eliminating Hallucinations | 5 | 6 | 2 | 3 | Matt pushes on whether fine-tuning eliminates hallucinations and asks if Sharon is solely in the small model camp. Sharon reframes model sizing and introduces parameter-efficient fine-tuning (PEFT), explaining how Lamini reduces model switching latency from 3 months to 3 milliseconds. | |
| AI Agents, Workflows, and Multi-Step Model Chains | 5 | 5 | 1 | 3 | Matt asks about AI agent adoption in enterprises and specifically challenges whether chaining multiple models creates compounding error and latency. Sharon agrees and explains how fine-tuning a single unified model can consolidate complex chains. | |
| The AMD Partnership, Superstation, and Scaling Laws | 4 | 6 | 1 | 2 | Matt asks about the AMD partnership and questions whether AMD hardware truly achieves parity with Nvidia CUDA. Sharon details her co-founder Greg Diamos's background in CUDA architecture and explains the core principles behind LLM scaling laws. | |
| Enterprise Go-to-Market, Customer Maturity, and Deployment Speed | 3 | 4 | 1 | 2 | Matt inquires about customer maturity, go-to-market friction, and emerging enterprise use cases. Sharon notes rapid customer learning curves and emphasizes that companies must deploy quickly despite transitioning from deterministic to probabilistic software. | |
| Conclusion and Resource Links | 1 | 1 | 0 | 0 | Matt wraps up the interview and asks where listeners can find Sharon and her online courses. Sharon provides her contact details and Coursera course links. |