Aug 9, 2023 · 50m · mad

Startup to Industry Standard: Lukas Biewald Explains How W&B Scaled MLOps for OpenAI, NVIDIA & More

Lukas Biewald · 37m spoken Matt Turck · 8m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of The MAD Podcast, FirstMark Partner Matt Turck interviews Lukas Biewald, Co-Founder and CEO of Weights & Biases, about building a developer-first MLOps platform. Biewald discusses W&B's evolution, strategies for converting bottom-up developer adoption into enterprise contracts, and adapting ML infrastructure for Large Language Models.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 18.2% of the talking time here. How this is scored →

Matt as informed peer 3.0 Guest teaching 2.9 Guest disagreement 0.5 Matt pushing back 0.5
05100:0015:0030:0045:000:10–4:22 · Matt as informed peer 2/10 Lukas Biewald's Journey from CrowdFlower to Weights & Biases Matt introduces Lukas and references their prior 2015 meeting alongside basic W&B funding numbers. Lukas provides a friendly historical summary of CrowdFlower and how developer frustration led to founding Weights & Biases.4:22–13:00 · Matt as informed peer 3/10 Defining the Evolving Machine Learning Developer Persona Matt asks clarifying questions about what constitutes an ML developer and what W&B actually tracks. Lukas educates the host on the complexity of hyperparameters, pre-processing, artifacts, and model registries.13:00–15:46 · Matt as informed peer 2/10 Adapting MLOps Workflows for Large Language Models Matt prompts Lukas on whether LLMs fundamentally change MLOps platforms. Lukas explains that while workflows remain exploratory, LLMs shift developer attention from complex loss graphs to qualitative text analysis and anecdotes.15:46–19:44 · Matt as informed peer 3/10 Challenges in Model Monitoring and Enterprise LLM Adoption Matt asks if enterprises are deploying LLM applications into production at scale. Lukas offers a reality check, noting VCs are often surprised by how few enterprises have LLMs in true live production due to monitoring challenges.19:44–25:02 · Matt as informed peer 3/10 Real-World Enterprise ML Use Cases and Market Maturity Matt asks for production model count estimates from sophisticated enterprise customers. Lukas reframes the premise by explaining that enterprise applications consist of deeply nested multi-model pipelines where counting individual models is misleading.25:02–37:43 · Matt as informed peer 4/10 Building a Developer-Centric Bottom-Up Go-To-Market Strategy Matt demonstrates knowledge of developer GTM tactics by referencing W&B's open blog publishing platform and courseware. Lukas outlines their slow exponential growth, focus on NPS, and third-party integrations.37:43–43:41 · Matt as informed peer 4/10 Enterprise Sales Execution: Bridging Developer Usage to Organization Standardization Matt asks how developer usage transitions into enterprise standardization and asks about optimal AE hiring profiles. Lukas details how credit card usage inside large companies leads to enterprise security and compliance deals.43:41–49:16 · Matt as informed peer 3/10 Executive Leadership Practices, Staying Technical, and Industry Tooling Analysis Matt asks Lukas how he stays technical as CEO and requests reviews of popular developer tools. Lukas breaks down why PyTorch beat TensorFlow due to developer empathy rather than just feature sets.0:10–4:22 · Guest teaching 1/10 Lukas Biewald's Journey from CrowdFlower to Weights & Biases Matt introduces Lukas and references their prior 2015 meeting alongside basic W&B funding numbers. Lukas provides a friendly historical summary of CrowdFlower and how developer frustration led to founding Weights & Biases.4:22–13:00 · Guest teaching 3/10 Defining the Evolving Machine Learning Developer Persona Matt asks clarifying questions about what constitutes an ML developer and what W&B actually tracks. Lukas educates the host on the complexity of hyperparameters, pre-processing, artifacts, and model registries.13:00–15:46 · Guest teaching 4/10 Adapting MLOps Workflows for Large Language Models Matt prompts Lukas on whether LLMs fundamentally change MLOps platforms. Lukas explains that while workflows remain exploratory, LLMs shift developer attention from complex loss graphs to qualitative text analysis and anecdotes.15:46–19:44 · Guest teaching 4/10 Challenges in Model Monitoring and Enterprise LLM Adoption Matt asks if enterprises are deploying LLM applications into production at scale. Lukas offers a reality check, noting VCs are often surprised by how few enterprises have LLMs in true live production due to monitoring challenges.19:44–25:02 · Guest teaching 4/10 Real-World Enterprise ML Use Cases and Market Maturity Matt asks for production model count estimates from sophisticated enterprise customers. Lukas reframes the premise by explaining that enterprise applications consist of deeply nested multi-model pipelines where counting individual models is misleading.25:02–37:43 · Guest teaching 2/10 Building a Developer-Centric Bottom-Up Go-To-Market Strategy Matt demonstrates knowledge of developer GTM tactics by referencing W&B's open blog publishing platform and courseware. Lukas outlines their slow exponential growth, focus on NPS, and third-party integrations.37:43–43:41 · Guest teaching 2/10 Enterprise Sales Execution: Bridging Developer Usage to Organization Standardization Matt asks how developer usage transitions into enterprise standardization and asks about optimal AE hiring profiles. Lukas details how credit card usage inside large companies leads to enterprise security and compliance deals.43:41–49:16 · Guest teaching 3/10 Executive Leadership Practices, Staying Technical, and Industry Tooling Analysis Matt asks Lukas how he stays technical as CEO and requests reviews of popular developer tools. Lukas breaks down why PyTorch beat TensorFlow due to developer empathy rather than just feature sets.0:10–4:22 · Guest disagreement 0/10 Lukas Biewald's Journey from CrowdFlower to Weights & Biases Matt introduces Lukas and references their prior 2015 meeting alongside basic W&B funding numbers. Lukas provides a friendly historical summary of CrowdFlower and how developer frustration led to founding Weights & Biases.4:22–13:00 · Guest disagreement 0/10 Defining the Evolving Machine Learning Developer Persona Matt asks clarifying questions about what constitutes an ML developer and what W&B actually tracks. Lukas educates the host on the complexity of hyperparameters, pre-processing, artifacts, and model registries.13:00–15:46 · Guest disagreement 1/10 Adapting MLOps Workflows for Large Language Models Matt prompts Lukas on whether LLMs fundamentally change MLOps platforms. Lukas explains that while workflows remain exploratory, LLMs shift developer attention from complex loss graphs to qualitative text analysis and anecdotes.15:46–19:44 · Guest disagreement 2/10 Challenges in Model Monitoring and Enterprise LLM Adoption Matt asks if enterprises are deploying LLM applications into production at scale. Lukas offers a reality check, noting VCs are often surprised by how few enterprises have LLMs in true live production due to monitoring challenges.19:44–25:02 · Guest disagreement 1/10 Real-World Enterprise ML Use Cases and Market Maturity Matt asks for production model count estimates from sophisticated enterprise customers. Lukas reframes the premise by explaining that enterprise applications consist of deeply nested multi-model pipelines where counting individual models is misleading.25:02–37:43 · Guest disagreement 0/10 Building a Developer-Centric Bottom-Up Go-To-Market Strategy Matt demonstrates knowledge of developer GTM tactics by referencing W&B's open blog publishing platform and courseware. Lukas outlines their slow exponential growth, focus on NPS, and third-party integrations.37:43–43:41 · Guest disagreement 0/10 Enterprise Sales Execution: Bridging Developer Usage to Organization Standardization Matt asks how developer usage transitions into enterprise standardization and asks about optimal AE hiring profiles. Lukas details how credit card usage inside large companies leads to enterprise security and compliance deals.43:41–49:16 · Guest disagreement 0/10 Executive Leadership Practices, Staying Technical, and Industry Tooling Analysis Matt asks Lukas how he stays technical as CEO and requests reviews of popular developer tools. Lukas breaks down why PyTorch beat TensorFlow due to developer empathy rather than just feature sets.0:10–4:22 · Matt pushing back 0/10 Lukas Biewald's Journey from CrowdFlower to Weights & Biases Matt introduces Lukas and references their prior 2015 meeting alongside basic W&B funding numbers. Lukas provides a friendly historical summary of CrowdFlower and how developer frustration led to founding Weights & Biases.4:22–13:00 · Matt pushing back 0/10 Defining the Evolving Machine Learning Developer Persona Matt asks clarifying questions about what constitutes an ML developer and what W&B actually tracks. Lukas educates the host on the complexity of hyperparameters, pre-processing, artifacts, and model registries.13:00–15:46 · Matt pushing back 0/10 Adapting MLOps Workflows for Large Language Models Matt prompts Lukas on whether LLMs fundamentally change MLOps platforms. Lukas explains that while workflows remain exploratory, LLMs shift developer attention from complex loss graphs to qualitative text analysis and anecdotes.15:46–19:44 · Matt pushing back 1/10 Challenges in Model Monitoring and Enterprise LLM Adoption Matt asks if enterprises are deploying LLM applications into production at scale. Lukas offers a reality check, noting VCs are often surprised by how few enterprises have LLMs in true live production due to monitoring challenges.19:44–25:02 · Matt pushing back 1/10 Real-World Enterprise ML Use Cases and Market Maturity Matt asks for production model count estimates from sophisticated enterprise customers. Lukas reframes the premise by explaining that enterprise applications consist of deeply nested multi-model pipelines where counting individual models is misleading.25:02–37:43 · Matt pushing back 1/10 Building a Developer-Centric Bottom-Up Go-To-Market Strategy Matt demonstrates knowledge of developer GTM tactics by referencing W&B's open blog publishing platform and courseware. Lukas outlines their slow exponential growth, focus on NPS, and third-party integrations.37:43–43:41 · Matt pushing back 1/10 Enterprise Sales Execution: Bridging Developer Usage to Organization Standardization Matt asks how developer usage transitions into enterprise standardization and asks about optimal AE hiring profiles. Lukas details how credit card usage inside large companies leads to enterprise security and compliance deals.43:41–49:16 · Matt pushing back 0/10 Executive Leadership Practices, Staying Technical, and Industry Tooling Analysis Matt asks Lukas how he stays technical as CEO and requests reviews of popular developer tools. Lukas breaks down why PyTorch beat TensorFlow due to developer empathy rather than just feature sets.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 45% · guest 55%0:00 · Matt 45% · guest 55%3:00 · Matt 13.2% · guest 86.8%3:00 · Matt 13.2% · guest 86.8%6:00 · Matt 24.2% · guest 75.8%6:00 · Matt 24.2% · guest 75.8%9:00 · Matt 12.3% · guest 87.7%9:00 · Matt 12.3% · guest 87.7%12:00 · Matt 20.7% · guest 79.3%12:00 · Matt 20.7% · guest 79.3%15:00 · Matt 13.2% · guest 86.8%15:00 · Matt 13.2% · guest 86.8%18:00 · Matt 11.7% · guest 88.3%18:00 · Matt 11.7% · guest 88.3%21:00 · Matt 10.4% · guest 89.6%21:00 · Matt 10.4% · guest 89.6%24:00 · Matt 23.3% · guest 76.7%24:00 · Matt 23.3% · guest 76.7%27:00 · Matt 16% · guest 84%27:00 · Matt 16% · guest 84%30:00 · Matt 21.2% · guest 78.8%30:00 · Matt 21.2% · guest 78.8%33:00 · Matt 7.9% · guest 92.1%33:00 · Matt 7.9% · guest 92.1%36:00 · Matt 15.3% · guest 84.7%36:00 · Matt 15.3% · guest 84.7%39:00 · Matt 11% · guest 89%39:00 · Matt 11% · guest 89%42:00 · Matt 21.8% · guest 78.2%42:00 · Matt 21.8% · guest 78.2%45:00 · Matt 15.3% · guest 84.7%45:00 · Matt 15.3% · guest 84.7%48:00 · Matt 29.5% · guest 70.5%48:00 · Matt 29.5% · guest 70.5%
Sharpest disagreement ▶ 17:49 Debunking VC expectations on enterprise LLM adoption

Lukas directly pushes back on venture capital hype, explaining that despite the industry excitement, very few enterprises have successfully put LLMs into actual production.

Hardest push from Matt ▶ 37:44 Probing top-down enterprise execution behind bottom-up developer GTM

Matt presses Lukas on the difficulty of scaling beyond bottom-up adoption, pointing out that developer virality alone rarely achieves full enterprise standardization without a structured sales force.

Biggest teaching moment ▶ 20:05 Correcting the concept of discrete model counts

Lukas educates Matt on modern ML architecture, explaining that asking how many models a company has in production misses the point since modern applications use interconnected webs of upstream models.

Matt holds his own ▶ 30:35 Highlighting specific developer marketing nuances

Matt displays deep familiarity with W&B's specific growth tactics by bringing up their open-author blog platform and how it engages community developers.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Lukas Biewald's Journey from CrowdFlower to Weights & Biases 2100 Matt introduces Lukas and references their prior 2015 meeting alongside basic W&B funding numbers. Lukas provides a friendly historical summary of CrowdFlower and how developer frustration led to founding Weights & Biases.
Defining the Evolving Machine Learning Developer Persona 3300 Matt asks clarifying questions about what constitutes an ML developer and what W&B actually tracks. Lukas educates the host on the complexity of hyperparameters, pre-processing, artifacts, and model registries.
Adapting MLOps Workflows for Large Language Models 2410 Matt prompts Lukas on whether LLMs fundamentally change MLOps platforms. Lukas explains that while workflows remain exploratory, LLMs shift developer attention from complex loss graphs to qualitative text analysis and anecdotes.
Challenges in Model Monitoring and Enterprise LLM Adoption 3421 Matt asks if enterprises are deploying LLM applications into production at scale. Lukas offers a reality check, noting VCs are often surprised by how few enterprises have LLMs in true live production due to monitoring challenges.
Real-World Enterprise ML Use Cases and Market Maturity 3411 Matt asks for production model count estimates from sophisticated enterprise customers. Lukas reframes the premise by explaining that enterprise applications consist of deeply nested multi-model pipelines where counting individual models is misleading.
Building a Developer-Centric Bottom-Up Go-To-Market Strategy 4201 Matt demonstrates knowledge of developer GTM tactics by referencing W&B's open blog publishing platform and courseware. Lukas outlines their slow exponential growth, focus on NPS, and third-party integrations.
Enterprise Sales Execution: Bridging Developer Usage to Organization Standardization 4201 Matt asks how developer usage transitions into enterprise standardization and asks about optimal AE hiring profiles. Lukas details how credit card usage inside large companies leads to enterprise security and compliance deals.
Executive Leadership Practices, Staying Technical, and Industry Tooling Analysis 3300 Matt asks Lukas how he stays technical as CEO and requests reviews of popular developer tools. Lukas breaks down why PyTorch beat TensorFlow due to developer empathy rather than just feature sets.

Statements from this episode (13)

Insight
Biewald: Data labeling software requires a top-down sales model
“Data labeling, I think really wants to be a top down sale”
Lukas Biewald Aug 9, 2023 ▶ 3:09
Assertion Not checkable as stated
Biewald: Almost all major LLMs were trained using Weights & Biases
“I think all of the major LLMs out there, almost all were trained using weights and biases.”
Lukas Biewald Aug 9, 2023 ▶ 12:17
Insight
Biewald: LLM API developers are less mathematically specialized than traditional ML engineers
“Even the people Working with a lot of these APIs you know, are, like, less huge math nerds than, you know, some of the people that have been training, you know, models for a long time.”
Lukas Biewald Aug 9, 2023 ▶ 14:22
Insight
Biewald: Developers building with LLMs focus heavily on qualitative anecdotes over metrics
“In the LL world, like the anecdote is something people really pay attention to.”
Lukas Biewald Aug 9, 2023 ▶ 15:00
Insight
Biewald: Simple operational errors cause more model failures than data drift
“People talk a lot about data drift in the industry. And that's this idea that like, you know, like language changes over time and you want to know that it's changing and sort of like have your model you know, notice that and update it. But I guess like what I …”
Lukas Biewald Aug 9, 2023 ▶ 16:21
Assertion Supported
Biewald: Most enterprises have not deployed LLMs into production yet
“I think that LLMs in particular, we talk to a lot of the people and we don't see a ton of people getting them into production yet. And I think it's funny, like VCs are always surprised, like when we tell them that I think that I don't know. I'm bullish on LMS,…”
Lukas Biewald Aug 9, 2023 ▶ 18:09
Assertion Not checkable as stated
Biewald: OpenAI is a W&B customer with a small number of production models
“OpenAI has been, like, a longtime customer. I mean, I consider them, like, extraordinarily sophisticated, and they have a pretty small number of models in, in production, so.”
Lukas Biewald Aug 9, 2023 ▶ 21:23
Disclosure
Biewald: No single industry accounts for over 15% of W&B revenue
“No one vertical is more than like, you know, 15% of our usage or revenue or anything like that.”
Lukas Biewald Aug 9, 2023 ▶ 22:18
Disclosure
Biewald: Traditional ML methods like boosted trees remain a large share of W&B usage
“A lot of people still running, you know, boosted trees or, you know, random forests inside of weights and biases. So we, you know, it's like actually huge. We should probably do a block, but it's still a big fraction of our You know, of our user base.”
Lukas Biewald Aug 9, 2023 ▶ 23:10
Assertion Not checkable as stated
Biewald estimates 99% of Global 2000 use ML for core operations
“I bet 99% of the global 2000 is using machine learning for something that they actually really care about.”
Lukas Biewald Aug 9, 2023 ▶ 25:37
Assertion Not checkable as stated
Biewald: Pharma is investing far more in deep learning than realized
“I think pharma is investing way more in deep learning than people realize.”
Lukas Biewald Aug 9, 2023 ▶ 26:42
Insight
Biewald: Technical buyers don't want sales dinners, they just want facts
“Nobody wants to golf or anything. I mean, that's for sure, right? Like, I mean, there's people like a super aggressive salesperson. Some of my salespeople are really competitive. I am actually really competitive myself, but they kind of like suppress it in a w…”
Lukas Biewald Aug 9, 2023 ▶ 42:31
Insight
Biewald: PyTorch beat TensorFlow through developer empathy, not eager execution
“I don't think they really, I think people tell this, the story of sort of the silver bullet. Of like you know, the eager execution model. But I think the reality is they just built a product with so much more empathy.”
Lukas Biewald Aug 9, 2023 ▶ 47:35
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.