Sep 25, 2023 · 24m · a16z

AI Food Fights in the Enterprise with Databricks' Ali Ghodsi

Ali Ghodsi · 15m spoken Ben Horowitz · 5m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this a16z interview, Databricks CEO Ali Ghodsi and Ben Horowitz examine enterprise AI adoption friction, the shift toward custom domain-specific models, open source dynamics, and realistic perspectives on AI risk and evaluation.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The host as informed peer 4.0 Guest teaching 4.9 Guest disagreement 2.7 The host pushing back 3.1
05100:0010:0020:000:10–3:08 · The host as informed peer 3/10 Title Sequence and Important Disclosures Ben asks why enterprises struggle to adopt generative AI, pushing slightly on whether accuracy is truly required for all enterprise use cases. Ali explains enterprise inertia, data security fears, and internal corporate politics around AI ownership.3:08–7:55 · The host as informed peer 4/10 Enterprise Data Security and Proprietary AI Strategy Ben asks sharp strategic questions regarding parameter scaling versus specialized model fine-tuning. Ali details how Mosaic allows enterprises to keep proprietary IP and build cost-efficient task-specific models rather than paying high inference costs for giant models.7:55–10:33 · The host as informed peer 4/10 Fine-Tuning Techniques and the GPU Shortage Ben asks Ali to specify exact fine-tuning methods and why Databricks needs a large model. Ali explains techniques like LoRa, QLoRa, and prefix tuning, while noting extreme GPU scarcity prevents Databricks from unleashing its full sales force.10:33–13:52 · The host as informed peer 3/10 Specialization in AI Applications and the Cisco Analogy Ben asks whether use cases will heavily fragment or consolidate around base models like cloud providers. Ali reframes the entire topic by comparing current LLM obsession to the Cisco router bubble in 2000, asserting application-layer value will far surpass model-layer value.13:52–17:20 · The host as informed peer 4/10 The Role and Future of Open Source AI Ben questions the open-source debate and notes that weights are required alongside code. Ali describes open-source dynamics, weight leaks, university research crises, and the historical catch-up loop between open-source and proprietary software.17:20–19:29 · The host as informed peer 5/10 Flaws in AI Benchmarks and Human-in-the-Loop Necessity When Ben cites AI passing medical exams, Ali strongly rejects the premise, declaring popular benchmarks like MMLU 'bullshit' due to data contamination and test memorization. Ben counters with a knowledgeable comparison to legacy fake database benchmarks.19:29–24:13 · The host as informed peer 5/10 AI Ethics, Automation, and Existential Risk Debates Ben opens by calling out Ali for dodging a previous question on ethics and open-source threats. Ali rejects simple toaster analogies but outlines why existential risk is distant, citing asymmetric GPU costs and lack of self-replication capability.0:10–3:08 · Guest teaching 3/10 Title Sequence and Important Disclosures Ben asks why enterprises struggle to adopt generative AI, pushing slightly on whether accuracy is truly required for all enterprise use cases. Ali explains enterprise inertia, data security fears, and internal corporate politics around AI ownership.3:08–7:55 · Guest teaching 4/10 Enterprise Data Security and Proprietary AI Strategy Ben asks sharp strategic questions regarding parameter scaling versus specialized model fine-tuning. Ali details how Mosaic allows enterprises to keep proprietary IP and build cost-efficient task-specific models rather than paying high inference costs for giant models.7:55–10:33 · Guest teaching 5/10 Fine-Tuning Techniques and the GPU Shortage Ben asks Ali to specify exact fine-tuning methods and why Databricks needs a large model. Ali explains techniques like LoRa, QLoRa, and prefix tuning, while noting extreme GPU scarcity prevents Databricks from unleashing its full sales force.10:33–13:52 · Guest teaching 6/10 Specialization in AI Applications and the Cisco Analogy Ben asks whether use cases will heavily fragment or consolidate around base models like cloud providers. Ali reframes the entire topic by comparing current LLM obsession to the Cisco router bubble in 2000, asserting application-layer value will far surpass model-layer value.13:52–17:20 · Guest teaching 5/10 The Role and Future of Open Source AI Ben questions the open-source debate and notes that weights are required alongside code. Ali describes open-source dynamics, weight leaks, university research crises, and the historical catch-up loop between open-source and proprietary software.17:20–19:29 · Guest teaching 6/10 Flaws in AI Benchmarks and Human-in-the-Loop Necessity When Ben cites AI passing medical exams, Ali strongly rejects the premise, declaring popular benchmarks like MMLU 'bullshit' due to data contamination and test memorization. Ben counters with a knowledgeable comparison to legacy fake database benchmarks.19:29–24:13 · Guest teaching 5/10 AI Ethics, Automation, and Existential Risk Debates Ben opens by calling out Ali for dodging a previous question on ethics and open-source threats. Ali rejects simple toaster analogies but outlines why existential risk is distant, citing asymmetric GPU costs and lack of self-replication capability.0:10–3:08 · Guest disagreement 1/10 Title Sequence and Important Disclosures Ben asks why enterprises struggle to adopt generative AI, pushing slightly on whether accuracy is truly required for all enterprise use cases. Ali explains enterprise inertia, data security fears, and internal corporate politics around AI ownership.3:08–7:55 · Guest disagreement 1/10 Enterprise Data Security and Proprietary AI Strategy Ben asks sharp strategic questions regarding parameter scaling versus specialized model fine-tuning. Ali details how Mosaic allows enterprises to keep proprietary IP and build cost-efficient task-specific models rather than paying high inference costs for giant models.7:55–10:33 · Guest disagreement 2/10 Fine-Tuning Techniques and the GPU Shortage Ben asks Ali to specify exact fine-tuning methods and why Databricks needs a large model. Ali explains techniques like LoRa, QLoRa, and prefix tuning, while noting extreme GPU scarcity prevents Databricks from unleashing its full sales force.10:33–13:52 · Guest disagreement 3/10 Specialization in AI Applications and the Cisco Analogy Ben asks whether use cases will heavily fragment or consolidate around base models like cloud providers. Ali reframes the entire topic by comparing current LLM obsession to the Cisco router bubble in 2000, asserting application-layer value will far surpass model-layer value.13:52–17:20 · Guest disagreement 2/10 The Role and Future of Open Source AI Ben questions the open-source debate and notes that weights are required alongside code. Ali describes open-source dynamics, weight leaks, university research crises, and the historical catch-up loop between open-source and proprietary software.17:20–19:29 · Guest disagreement 6/10 Flaws in AI Benchmarks and Human-in-the-Loop Necessity When Ben cites AI passing medical exams, Ali strongly rejects the premise, declaring popular benchmarks like MMLU 'bullshit' due to data contamination and test memorization. Ben counters with a knowledgeable comparison to legacy fake database benchmarks.19:29–24:13 · Guest disagreement 4/10 AI Ethics, Automation, and Existential Risk Debates Ben opens by calling out Ali for dodging a previous question on ethics and open-source threats. Ali rejects simple toaster analogies but outlines why existential risk is distant, citing asymmetric GPU costs and lack of self-replication capability.0:10–3:08 · The host pushing back 3/10 Title Sequence and Important Disclosures Ben asks why enterprises struggle to adopt generative AI, pushing slightly on whether accuracy is truly required for all enterprise use cases. Ali explains enterprise inertia, data security fears, and internal corporate politics around AI ownership.3:08–7:55 · The host pushing back 2/10 Enterprise Data Security and Proprietary AI Strategy Ben asks sharp strategic questions regarding parameter scaling versus specialized model fine-tuning. Ali details how Mosaic allows enterprises to keep proprietary IP and build cost-efficient task-specific models rather than paying high inference costs for giant models.7:55–10:33 · The host pushing back 2/10 Fine-Tuning Techniques and the GPU Shortage Ben asks Ali to specify exact fine-tuning methods and why Databricks needs a large model. Ali explains techniques like LoRa, QLoRa, and prefix tuning, while noting extreme GPU scarcity prevents Databricks from unleashing its full sales force.10:33–13:52 · The host pushing back 2/10 Specialization in AI Applications and the Cisco Analogy Ben asks whether use cases will heavily fragment or consolidate around base models like cloud providers. Ali reframes the entire topic by comparing current LLM obsession to the Cisco router bubble in 2000, asserting application-layer value will far surpass model-layer value.13:52–17:20 · The host pushing back 3/10 The Role and Future of Open Source AI Ben questions the open-source debate and notes that weights are required alongside code. Ali describes open-source dynamics, weight leaks, university research crises, and the historical catch-up loop between open-source and proprietary software.17:20–19:29 · The host pushing back 4/10 Flaws in AI Benchmarks and Human-in-the-Loop Necessity When Ben cites AI passing medical exams, Ali strongly rejects the premise, declaring popular benchmarks like MMLU 'bullshit' due to data contamination and test memorization. Ben counters with a knowledgeable comparison to legacy fake database benchmarks.19:29–24:13 · The host pushing back 6/10 AI Ethics, Automation, and Existential Risk Debates Ben opens by calling out Ali for dodging a previous question on ethics and open-source threats. Ali rejects simple toaster analogies but outlines why existential risk is distant, citing asymmetric GPU costs and lack of self-replication capability.

speaking balance: gold is the host, purple is the guest (3 minute bins)

0:00 · the host 0% · guest 100%0:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%18:00 · the host 0% · guest 100%18:00 · the host 0% · guest 100%21:00 · the host 0% · guest 100%21:00 · the host 0% · guest 100%24:00 · the host 0% · guest 100%24:00 · the host 0% · guest 100%
Sharpest disagreement ▶ 17:51 Ali calls LLM benchmarks bullshit

Ali forcefully rejects the host's point about medical exam scores, calling popular AI benchmarks bullshit due to dataset contamination and answer memorization.

Hardest push from the host ▶ 19:29 Ben confronts guest on dodging ethics

Ben directly refuses Ali's previous evasion by opening the segment with 'So then let me go to the question that you dodged', forcing him to address AI risk.

Biggest teaching moment ▶ 12:01 Historical reframe via Cisco router boom

Ali educates the host on market dynamics by comparing current LLM infrastructure hype to the 2000 Cisco router bubble, explaining why downstream applications hold the real long-term value.

The host holds their own ▶ 19:11 Ben connects critique to fake database benchmarks

Ben demonstrates his own technical background by backing up Ali's critique with an insightful parallel to legacy fake database performance benchmarks.

the scores for every segment, with the reasoning behind each
ChapterTopicThe host as informed peerGuest teachingGuest disagreementThe host pushing backWhy
Title Sequence and Important Disclosures 3313 Ben asks why enterprises struggle to adopt generative AI, pushing slightly on whether accuracy is truly required for all enterprise use cases. Ali explains enterprise inertia, data security fears, and internal corporate politics around AI ownership.
Enterprise Data Security and Proprietary AI Strategy 4412 Ben asks sharp strategic questions regarding parameter scaling versus specialized model fine-tuning. Ali details how Mosaic allows enterprises to keep proprietary IP and build cost-efficient task-specific models rather than paying high inference costs for giant models.
Fine-Tuning Techniques and the GPU Shortage 4522 Ben asks Ali to specify exact fine-tuning methods and why Databricks needs a large model. Ali explains techniques like LoRa, QLoRa, and prefix tuning, while noting extreme GPU scarcity prevents Databricks from unleashing its full sales force.
Specialization in AI Applications and the Cisco Analogy 3632 Ben asks whether use cases will heavily fragment or consolidate around base models like cloud providers. Ali reframes the entire topic by comparing current LLM obsession to the Cisco router bubble in 2000, asserting application-layer value will far surpass model-layer value.
The Role and Future of Open Source AI 4523 Ben questions the open-source debate and notes that weights are required alongside code. Ali describes open-source dynamics, weight leaks, university research crises, and the historical catch-up loop between open-source and proprietary software.
Flaws in AI Benchmarks and Human-in-the-Loop Necessity 5664 When Ben cites AI passing medical exams, Ali strongly rejects the premise, declaring popular benchmarks like MMLU 'bullshit' due to data contamination and test memorization. Ben counters with a knowledgeable comparison to legacy fake database benchmarks.
AI Ethics, Automation, and Existential Risk Debates 5546 Ben opens by calling out Ali for dodging a previous question on ethics and open-source threats. Ali rejects simple toaster analogies but outlines why existential risk is distant, citing asymmetric GPU costs and lack of self-replication capability.

Statements from this episode (12)

Assertion Not checkable as stated
Horowitz: a16z sees zero generative AI traction in large enterprises
“Every company that has traction is in a category like selling to developers, or selling to consumers, or maybe selling to, like, small kinds of, you know, law firms or these. Kinds of things, but we haven't seen anybody with any traction in the enterprise.”
Ben Horowitz Sep 25, 2023 ▶ 0:37
Insight
Ghodsi: Internal enterprise turf wars are delaying generative AI adoption
“There's like a food fight internally at the large enterprise, which is I own generative AI, not Ben. And then you go around and say, Hey, I own generative AI. And it's like, no, no, no, my team is building. So, so there's this, you know, food fight internally …”
Ali Ghodsi Sep 25, 2023 ▶ 2:28
Disclosure
Ghodsi: Enterprise CEOs now bypass CIOs to discuss AI strategy directly
“I get to talk these days to the CEOs of these big companies who previously were not interested in what I'm doing. I would be talking to the CIO, but now suddenly they want to talk like, Hey, I want this generative AI. I want to talk strategy, strategy for my c…”
Ali Ghodsi Sep 25, 2023 ▶ 3:24
Assertion Not checkable as stated
Ghodsi: Enterprise CEOs want proprietary AI models to protect their data
“One of the things that's really interesting that's happened in the sort of brains of the CEOs and the boards is that they realize Maybe I can beat my competition. Maybe this is the kryptonite that will help me kill my enemy. I have the data with generative AI.…”
Ali Ghodsi Sep 25, 2023 ▶ 3:46
Insight
Ghodsi: Smaller custom models can beat large LLMs on domain accuracy
“And there you're better off if you have a good data set to train, you can train a smaller model. The latency will be faster to use it later, and it will be cheaper to use it later, and yes, you can have absolutely accuracy that beats the really large model, bu…”
Ali Ghodsi Sep 25, 2023 ▶ 7:29
Disclosure
Ghodsi: Databricks restricted MosaicML sales because of severe GPU shortages
“So we bought Mosaic. I did not unleash our sales force and go to market of 3000 people. To sell the thing that we bought because we just can't satisfy the demand. Like there's not enough GPUs.”
Ali Ghodsi Sep 25, 2023 ▶ 10:12
Insight
Ghodsi: Hype around the largest LLMs mirrors 2000s focus on Cisco
“So I think it's a little bit like right now like that. Who has the largest LLM? Obviously whoever can build the largest one that can train it the most obviously will own all of AI and all the future of humanity. But just like the internet, someone will show up…”
Ali Ghodsi Sep 25, 2023 ▶ 12:40
What-if
Ghodsi: AI progress would be far behind without Meta releasing Llama
“If the original Llama was never released, What would the state of the world and our view of AI be right now? We would be way further behind, right? And A, it was a big model you know, by what existed in open source and it was open sourced. And both of those th…”
Ali Ghodsi Sep 25, 2023 ▶ 14:20
Opinion
Ghodsi: All current LLM evaluation benchmarks are completely bullshit
“I kind of think all the benchmarks are bullshit, and so all these, so all the LLM benchmarks, here's how it works.”
Ali Ghodsi Sep 25, 2023 ▶ 17:52
Assertion Supported
Ghodsi: MMLU benchmark scores are inflated due to web data contamination
“MMLU is just a multi-choice question that's on the web. Ask a question. Here's, is the answer A, B, C, D, and then it says what the right answer is. And it's on the web. You can deliberately train on it and create an LLM that crushes it on that. Okay. Or you c…”
Ali Ghodsi Sep 25, 2023 ▶ 18:15
Assertion Not checkable as stated
Horowitz: No large language model has ever decided to do anything
“No LLM has ever decided to do anything.”
Ben Horowitz Sep 25, 2023 ▶ 21:24
Prediction Not checkable as stated
Ghodsi: Humanity is very far away from autonomous AI self-reproduction
“Once you have reproduction and, you know, the building of new ones automatically once you crack the code on that loop, yes, then I think we're fucked, but we're very far away from that.”
Ali Ghodsi Sep 25, 2023 ▶ 23:44
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.