Jan 2, 2019 · 30m · a16z

a16z Podcast | Why the Datacenter Needs an Operating System

Benjamin Heinemann · 18m spoken Steven Sinofsky · 10m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of the a16z Podcast, host Steven Sinofsky and guest Benjamin Hindman explore the concept of Data Center Operating Systems (DCOS), explaining how treating an entire cluster of servers as a single computer solves severe hardware inefficiency and simplifies distributed application management.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The host as informed peer 5.3 Guest teaching 3.7 Guest disagreement 0.3 The host pushing back 2.0
05100:0010:0020:0030:002:37–5:51 · The host as informed peer 5/10 DCOS Architecture and the Mesos Microkernel Sinofsky demonstrates familiarity with OS architecture, making jokes about Linux vs Unix and asking targeted questions about scheduling and kernel abstractions. Heinemann explains Mesos as a microkernel and clarifies the distinction between tasks and processes. The exchange is highly collaborative and educational.5:51–8:36 · The host as informed peer 4/10 Resource Allocation, Agent Architecture, and Communication Sinofsky asks how machines connect across a cluster and uses system bus metaphors. Heinemann explains agent processes and the Mesos master architecture in response. The conversation remains smooth and conversational without friction.8:36–11:04 · The host as informed peer 5/10 Interacting with DCOS via Command Line Interface Sinofsky probes how CLI abstractions work in practice and correctly points out that system diagnostics still require visibility into underlying threads and processes. Heinemann validates this, explaining how the DCOS CLI allows drilling down. Sinofsky briefly catches Heinemann switching between process and task terminology.11:04–13:08 · The host as informed peer 3/10 Package Management and Real-World Microservices Sinofsky asks where tasks originate and prompts Heinemann to provide concrete real-world service examples. Heinemann uses Twitter's microservice architecture as a clear illustration. The discussion is purely explanatory.13:08–15:59 · The host as informed peer 6/10 Security Primitives and Standardized Distributed Abstractions Sinofsky pushes back on the security implications of running code anywhere across a cluster, providing real-world enterprise examples like fragmented web and analytics stacks. Heinemann frames standardized security primitives as essential operating system responsibilities. Both share technical insight on distributed system design.15:59–18:26 · The host as informed peer 5/10 Integrating Containers, Docker, and Oversubscription Sinofsky asks whether DCOS makes container solutions like Docker obsolete or if they complement each other. Heinemann educates the host on Mesos container history back to 2009 with Solaris zones, explaining how modern container formats plug directly into DCOS.18:26–22:17 · The host as informed peer 8/10 Historical Analogies: Virtual Memory and Hardware Independence Sinofsky leads with high expertise, drawing a detailed historical parallel between DCOS resource abstraction and the introduction of virtual memory, referencing 640K limits and manual swap tuning. He challenges whether stubborn engineers will resist letting software manage placement. Heinemann enthusiastically agrees with the analogy.22:17–27:41 · The host as informed peer 7/10 DCOS Compared to IaaS and PaaS Solutions Sinofsky pushes on why DCOS is necessary over traditional IaaS and PaaS, giving strategic advice on how Enterprise CIOs should bypass VM virtualization overhead. Heinemann clarifies the structural differences between simple VM wrapping and providing system call APIs for distributed apps.27:41–30:31 · The host as informed peer 5/10 Future Roadmap: Stateful Services and Planned Maintenance Sinofsky asks about the future roadmap, and Heinemann explains stateful service support and planned maintenance primitives. Sinofsky cleanly synthesizes Heinemann's explanation as managing planned failures versus unplanned ones before closing the episode.2:37–5:51 · Guest teaching 4/10 DCOS Architecture and the Mesos Microkernel Sinofsky demonstrates familiarity with OS architecture, making jokes about Linux vs Unix and asking targeted questions about scheduling and kernel abstractions. Heinemann explains Mesos as a microkernel and clarifies the distinction between tasks and processes. The exchange is highly collaborative and educational.5:51–8:36 · Guest teaching 4/10 Resource Allocation, Agent Architecture, and Communication Sinofsky asks how machines connect across a cluster and uses system bus metaphors. Heinemann explains agent processes and the Mesos master architecture in response. The conversation remains smooth and conversational without friction.8:36–11:04 · Guest teaching 3/10 Interacting with DCOS via Command Line Interface Sinofsky probes how CLI abstractions work in practice and correctly points out that system diagnostics still require visibility into underlying threads and processes. Heinemann validates this, explaining how the DCOS CLI allows drilling down. Sinofsky briefly catches Heinemann switching between process and task terminology.11:04–13:08 · Guest teaching 3/10 Package Management and Real-World Microservices Sinofsky asks where tasks originate and prompts Heinemann to provide concrete real-world service examples. Heinemann uses Twitter's microservice architecture as a clear illustration. The discussion is purely explanatory.13:08–15:59 · Guest teaching 4/10 Security Primitives and Standardized Distributed Abstractions Sinofsky pushes back on the security implications of running code anywhere across a cluster, providing real-world enterprise examples like fragmented web and analytics stacks. Heinemann frames standardized security primitives as essential operating system responsibilities. Both share technical insight on distributed system design.15:59–18:26 · Guest teaching 4/10 Integrating Containers, Docker, and Oversubscription Sinofsky asks whether DCOS makes container solutions like Docker obsolete or if they complement each other. Heinemann educates the host on Mesos container history back to 2009 with Solaris zones, explaining how modern container formats plug directly into DCOS.18:26–22:17 · Guest teaching 2/10 Historical Analogies: Virtual Memory and Hardware Independence Sinofsky leads with high expertise, drawing a detailed historical parallel between DCOS resource abstraction and the introduction of virtual memory, referencing 640K limits and manual swap tuning. He challenges whether stubborn engineers will resist letting software manage placement. Heinemann enthusiastically agrees with the analogy.22:17–27:41 · Guest teaching 4/10 DCOS Compared to IaaS and PaaS Solutions Sinofsky pushes on why DCOS is necessary over traditional IaaS and PaaS, giving strategic advice on how Enterprise CIOs should bypass VM virtualization overhead. Heinemann clarifies the structural differences between simple VM wrapping and providing system call APIs for distributed apps.27:41–30:31 · Guest teaching 5/10 Future Roadmap: Stateful Services and Planned Maintenance Sinofsky asks about the future roadmap, and Heinemann explains stateful service support and planned maintenance primitives. Sinofsky cleanly synthesizes Heinemann's explanation as managing planned failures versus unplanned ones before closing the episode.2:37–5:51 · Guest disagreement 1/10 DCOS Architecture and the Mesos Microkernel Sinofsky demonstrates familiarity with OS architecture, making jokes about Linux vs Unix and asking targeted questions about scheduling and kernel abstractions. Heinemann explains Mesos as a microkernel and clarifies the distinction between tasks and processes. The exchange is highly collaborative and educational.5:51–8:36 · Guest disagreement 0/10 Resource Allocation, Agent Architecture, and Communication Sinofsky asks how machines connect across a cluster and uses system bus metaphors. Heinemann explains agent processes and the Mesos master architecture in response. The conversation remains smooth and conversational without friction.8:36–11:04 · Guest disagreement 0/10 Interacting with DCOS via Command Line Interface Sinofsky probes how CLI abstractions work in practice and correctly points out that system diagnostics still require visibility into underlying threads and processes. Heinemann validates this, explaining how the DCOS CLI allows drilling down. Sinofsky briefly catches Heinemann switching between process and task terminology.11:04–13:08 · Guest disagreement 0/10 Package Management and Real-World Microservices Sinofsky asks where tasks originate and prompts Heinemann to provide concrete real-world service examples. Heinemann uses Twitter's microservice architecture as a clear illustration. The discussion is purely explanatory.13:08–15:59 · Guest disagreement 1/10 Security Primitives and Standardized Distributed Abstractions Sinofsky pushes back on the security implications of running code anywhere across a cluster, providing real-world enterprise examples like fragmented web and analytics stacks. Heinemann frames standardized security primitives as essential operating system responsibilities. Both share technical insight on distributed system design.15:59–18:26 · Guest disagreement 0/10 Integrating Containers, Docker, and Oversubscription Sinofsky asks whether DCOS makes container solutions like Docker obsolete or if they complement each other. Heinemann educates the host on Mesos container history back to 2009 with Solaris zones, explaining how modern container formats plug directly into DCOS.18:26–22:17 · Guest disagreement 0/10 Historical Analogies: Virtual Memory and Hardware Independence Sinofsky leads with high expertise, drawing a detailed historical parallel between DCOS resource abstraction and the introduction of virtual memory, referencing 640K limits and manual swap tuning. He challenges whether stubborn engineers will resist letting software manage placement. Heinemann enthusiastically agrees with the analogy.22:17–27:41 · Guest disagreement 1/10 DCOS Compared to IaaS and PaaS Solutions Sinofsky pushes on why DCOS is necessary over traditional IaaS and PaaS, giving strategic advice on how Enterprise CIOs should bypass VM virtualization overhead. Heinemann clarifies the structural differences between simple VM wrapping and providing system call APIs for distributed apps.27:41–30:31 · Guest disagreement 0/10 Future Roadmap: Stateful Services and Planned Maintenance Sinofsky asks about the future roadmap, and Heinemann explains stateful service support and planned maintenance primitives. Sinofsky cleanly synthesizes Heinemann's explanation as managing planned failures versus unplanned ones before closing the episode.2:37–5:51 · The host pushing back 2/10 DCOS Architecture and the Mesos Microkernel Sinofsky demonstrates familiarity with OS architecture, making jokes about Linux vs Unix and asking targeted questions about scheduling and kernel abstractions. Heinemann explains Mesos as a microkernel and clarifies the distinction between tasks and processes. The exchange is highly collaborative and educational.5:51–8:36 · The host pushing back 1/10 Resource Allocation, Agent Architecture, and Communication Sinofsky asks how machines connect across a cluster and uses system bus metaphors. Heinemann explains agent processes and the Mesos master architecture in response. The conversation remains smooth and conversational without friction.8:36–11:04 · The host pushing back 2/10 Interacting with DCOS via Command Line Interface Sinofsky probes how CLI abstractions work in practice and correctly points out that system diagnostics still require visibility into underlying threads and processes. Heinemann validates this, explaining how the DCOS CLI allows drilling down. Sinofsky briefly catches Heinemann switching between process and task terminology.11:04–13:08 · The host pushing back 1/10 Package Management and Real-World Microservices Sinofsky asks where tasks originate and prompts Heinemann to provide concrete real-world service examples. Heinemann uses Twitter's microservice architecture as a clear illustration. The discussion is purely explanatory.13:08–15:59 · The host pushing back 3/10 Security Primitives and Standardized Distributed Abstractions Sinofsky pushes back on the security implications of running code anywhere across a cluster, providing real-world enterprise examples like fragmented web and analytics stacks. Heinemann frames standardized security primitives as essential operating system responsibilities. Both share technical insight on distributed system design.15:59–18:26 · The host pushing back 2/10 Integrating Containers, Docker, and Oversubscription Sinofsky asks whether DCOS makes container solutions like Docker obsolete or if they complement each other. Heinemann educates the host on Mesos container history back to 2009 with Solaris zones, explaining how modern container formats plug directly into DCOS.18:26–22:17 · The host pushing back 3/10 Historical Analogies: Virtual Memory and Hardware Independence Sinofsky leads with high expertise, drawing a detailed historical parallel between DCOS resource abstraction and the introduction of virtual memory, referencing 640K limits and manual swap tuning. He challenges whether stubborn engineers will resist letting software manage placement. Heinemann enthusiastically agrees with the analogy.22:17–27:41 · The host pushing back 3/10 DCOS Compared to IaaS and PaaS Solutions Sinofsky pushes on why DCOS is necessary over traditional IaaS and PaaS, giving strategic advice on how Enterprise CIOs should bypass VM virtualization overhead. Heinemann clarifies the structural differences between simple VM wrapping and providing system call APIs for distributed apps.27:41–30:31 · The host pushing back 1/10 Future Roadmap: Stateful Services and Planned Maintenance Sinofsky asks about the future roadmap, and Heinemann explains stateful service support and planned maintenance primitives. Sinofsky cleanly synthesizes Heinemann's explanation as managing planned failures versus unplanned ones before closing the episode.

speaking balance: gold is the host, purple is the guest (3 minute bins)

0:00 · the host 0% · guest 100%0:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%18:00 · the host 0% · guest 100%18:00 · the host 0% · guest 100%21:00 · the host 0% · guest 100%21:00 · the host 0% · guest 100%24:00 · the host 0% · guest 100%24:00 · the host 0% · guest 100%27:00 · the host 0% · guest 100%27:00 · the host 0% · guest 100%30:00 · the host 0% · guest 100%30:00 · the host 0% · guest 100%
Sharpest disagreement ▶ 26:19 Rejecting PaaS as a Complete Solution

Heinemann firmly reframes the host's premise that platform-as-a-service solves distributed computing, pointing out that PaaS merely abstracts machine startup while lacking ongoing runtime system call APIs.

Hardest push from the host ▶ 13:08 Challenging Code Isolation and Security

Sinofsky directly challenges the guest's model by pointing out that letting code execute anywhere creates unpredicted enterprise security vulnerabilities.

Biggest teaching moment ▶ 16:24 History of Containerization in Mesos

Heinemann educates the host on how Mesos had built-in containerization as early as 2009 using Solaris zones, long before modern tools like Docker popularized container formats.

The host holds their own ▶ 18:25 Deep Historical Analogy to Early Computing

Sinofsky showcases deep technical authority by drawing an extended historical comparison between data center operating systems and early virtual memory abstractions, recalling personal experience swap-tuning code in 640K environments.

the scores for every segment, with the reasoning behind each
ChapterTopicThe host as informed peerGuest teachingGuest disagreementThe host pushing backWhy
DCOS Architecture and the Mesos Microkernel 5412 Sinofsky demonstrates familiarity with OS architecture, making jokes about Linux vs Unix and asking targeted questions about scheduling and kernel abstractions. Heinemann explains Mesos as a microkernel and clarifies the distinction between tasks and processes. The exchange is highly collaborative and educational.
Resource Allocation, Agent Architecture, and Communication 4401 Sinofsky asks how machines connect across a cluster and uses system bus metaphors. Heinemann explains agent processes and the Mesos master architecture in response. The conversation remains smooth and conversational without friction.
Interacting with DCOS via Command Line Interface 5302 Sinofsky probes how CLI abstractions work in practice and correctly points out that system diagnostics still require visibility into underlying threads and processes. Heinemann validates this, explaining how the DCOS CLI allows drilling down. Sinofsky briefly catches Heinemann switching between process and task terminology.
Package Management and Real-World Microservices 3301 Sinofsky asks where tasks originate and prompts Heinemann to provide concrete real-world service examples. Heinemann uses Twitter's microservice architecture as a clear illustration. The discussion is purely explanatory.
Security Primitives and Standardized Distributed Abstractions 6413 Sinofsky pushes back on the security implications of running code anywhere across a cluster, providing real-world enterprise examples like fragmented web and analytics stacks. Heinemann frames standardized security primitives as essential operating system responsibilities. Both share technical insight on distributed system design.
Integrating Containers, Docker, and Oversubscription 5402 Sinofsky asks whether DCOS makes container solutions like Docker obsolete or if they complement each other. Heinemann educates the host on Mesos container history back to 2009 with Solaris zones, explaining how modern container formats plug directly into DCOS.
Historical Analogies: Virtual Memory and Hardware Independence 8203 Sinofsky leads with high expertise, drawing a detailed historical parallel between DCOS resource abstraction and the introduction of virtual memory, referencing 640K limits and manual swap tuning. He challenges whether stubborn engineers will resist letting software manage placement. Heinemann enthusiastically agrees with the analogy.
DCOS Compared to IaaS and PaaS Solutions 7413 Sinofsky pushes on why DCOS is necessary over traditional IaaS and PaaS, giving strategic advice on how Enterprise CIOs should bypass VM virtualization overhead. Heinemann clarifies the structural differences between simple VM wrapping and providing system call APIs for distributed apps.
Future Roadmap: Stateful Services and Planned Maintenance 5501 Sinofsky asks about the future roadmap, and Heinemann explains stateful service support and planned maintenance primitives. Sinofsky cleanly synthesizes Heinemann's explanation as managing planned failures versus unplanned ones before closing the episode.

Statements from this episode (11)

Assertion Supported
Sinofsky: 85% of enterprise datacenter resources go unused
“It leads to this unbelievable waste, and waste in a data center is, is a big mess. 85% of the resources go unused”
Steven Sinofsky Jan 2, 2019 ▶ 1:47
Assertion Supported
Hindman: Apache Mesos functions as a bare kernel requiring higher-level orchestration
“If you download Mesos by itself today, there's not really much you can do with it. Just like if you're downloading the Linux kernel by itself today, right?”
Benjamin Heinemann Jan 2, 2019 ▶ 7:18
Insight
Hindman: Non-standardized infrastructure degrades enterprise security auditing capabilities
“Oftentimes you have worse security because, you know, rather than a security team being able to audit just the one way in which everything gets to run, they have to audit a whole bunch of different processes.”
Benjamin Heinemann Jan 2, 2019 ▶ 13:46
Insight
Hindman: Organization-tailored distributed systems prevent software portability across companies
“Because people are building distributed systems in such a personalized way and a personalized for their organization or their company, you can't easily build a distributed system in one organization and move it to another organization.”
Benjamin Heinemann Jan 2, 2019 ▶ 14:00
Prediction Not checkable as stated
Hindman: Programmatic datacenter scheduling will outperform manual human tuning
“There will probably be a lot of people who believe that they can do it better, but time's going to show that actually we can start to do far more sophisticated things. And we will be able to do far better scheduling for utilization, for meeting SLAs, for servi…”
Benjamin Heinemann Jan 2, 2019 ▶ 19:46
Assertion Supported
Hindman: Twitter, Airbnb, Netflix, and PayPal run on Mesos
“The open source components that, that make up a large part of the DCOS are used by a large number of companies today. Some of the biggest users out there are companies like Twitter Airbnb, HubSpot, eBay, and PayPal are using it for running things. Netflix is u…”
Benjamin Heinemann Jan 2, 2019 ▶ 20:57
Prediction Not checkable as stated
Sinofsky: DCOS will enable rapid hardware innovation like ARM server adoption
“One of the things that an operating system brings is it allows hardware to proceed at a different pace of innovation. And so I, when I look at DCOS, I think, wow, this is really gonna free a set of people to go, well, let's just go replace our servers with ARM…”
Steven Sinofsky Jan 2, 2019 ▶ 21:43
Opinion
Sinofsky: Virtualizing existing servers for cloud migration is a waste of time
“Everybody wants to move to cloud, they don't know what that means, and so they're very quickly virtualizing the servers that they have laying around, and I'm a big believer that that's just not a useful, a good use of time.”
Steven Sinofsky Jan 2, 2019 ▶ 23:37
Assertion Not checkable as stated
Hindman: Server virtualization carries a 30% performance overhead
“You don't have to start paying that 30% virtualization overhead for running your applications, which can start to save a lot of money.”
Benjamin Heinemann Jan 2, 2019 ▶ 25:02
Prediction Not checkable as stated
Sinofsky: Rewritten enterprise apps will be native distributed systems within 10 years
“Your thousand apps over the next 10 years that get rewritten are all just going to squeeze in and use the right amount of resources.”
Steven Sinofsky Jan 2, 2019 ▶ 25:46
Assertion Not checkable as stated
Hindman: Most organizations handle datacenter maintenance manually rather than via software
“Usually the way this works in most, most organizations is a human walks up to another human and says, hey, I'm going to be taking this rack down. What can we actually do about this? Well, we can turn this into software.”
Benjamin Heinemann Jan 2, 2019 ▶ 29:24
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.