Matei Zaharia

Co-Founder and CTO, Databricks · 1 appearance on the record.

computed by AI from the episodes · how this works → · full disclaimer →

founderexecutiveacademicscientistengineer@matei_zaharia ↗LinkedIn ↗profiles.stanford.edu/matei-zaharia ↗Wikipedia ↗

Matei Zaharia created Apache Spark during his Ph.D. at UC Berkeley and co-founded Databricks. He has also co-developed open-source technologies such as MLflow and Delta Lake, earning the ACM Prize in Computing.

13statements → 6claims → 4claims resolved → 75%fully supported → 4.15/5average certainty → 1.46/5average debate potential → 4.5/5argument clarity · the sources → 6said about them ↓

3 supported 0 partly supported 1 contradicted 1 not yet assessed 1 not checkable as stated how the 6 claims stand · each chip opens the sources

6 assertions · 1 opinion · 4 insights · 2 disclosures · every statement was checked. The predictions and assertions are the 6 claims: statements the public record can support or contradict. 4 are resolved, 1 is not yet assessed, and 1 names no date, number or outcome precise enough to check. Everything else (opinions, insights, what ifs, disclosures) can never be settled by the record, so it carries no assessment.

The record, in short

What the tape says about how Matei argues and how the claims held up. Everything they said, and everything said about them, is in the tabs below.

Their most notable supported claim

Assertion Supported
Zaharia: MapReduce was created by Google for nightly web indexing
“MapReduce initially came out of Google, where it was used for web indexing, and the whole point was, I will run this giant job every night, and in the morning, it's built a new index of the web.”
Matei Zaharia Jan 2, 2019 ▶ 3:49 a16z Podcast | A Conversation With the Inventor of Spark

Their most notable contradicted claim

Assertion Contradicted
Zaharia: Databricks includes all engine improvements in open-source Spark
“It's the same Spark that anyone else gets in the open source. All the libraries, all the improvements we put into the engine, you can just download them and run them yourselves. Or if you want you know, you can talk to a vendor that provides support. Support o…”
Matei Zaharia Jan 2, 2019 ▶ 18:31 a16z Podcast | A Conversation With the Inventor of Spark

Expressed certainty vs assessment result

none yet certainty 1
none yet certainty 2
none yet certainty 3
67% certainty 4
100% certainty 5

weighted support: a fully supported claim counts one, a partly supported claim counts half. Each filled bar is clickable and opens exactly those claims; "none yet" means nothing said at that certainty level has resolved yet

Argument clarity: do they answer the question? how? →

4.5 / 5 directness 4.8 · coherence 4.9 · precision 4.4 · compression 4.1

answered every one of 10 assessed questions directly

This is a score against a rubric. It is not a rank. Every host question → answer exchange is scored with names hidden on directness, coherence, precision and compression, 1–5 each, on meaning alone: disfluencies are ignored, and only raw unedited episodes count. This is the score that measures thought. Every scored exchange, scores shown → · The rubric and its checks →

How they sound: speaking style how? →

242 words/min while actually speaking · 55.3 um and uh per 1k words · 22.2 false starts per 1k · 36.2% of pauses land inside a clause

Measured by listening to the audio itself: 2,982 words across 1 episode of raw-level tape, transcribed verbatim with every um and uh kept, each one attributed only where the alignment onto our timed stream is unambiguous. These are measurements of speaking style. We do not rank them: across this corpus, fluency and argument quality are nearly uncorrelated (ρ≈0.2), and smooth talking does not signal clear thinking. How it's measured →

Everything Matei Zaharia said on the a16z Podcast that made the record, most notable first. Filter by type, assessment or year in the ledger →

Insight
Zaharia: Data scientists prefer advanced work over answering routine user queries
“The interesting thing is nobody wants only the insiders to work with data, basically. Everyone wants to be able to access it directly. Actually, there was a great keynote about this at the Spark Summit by Gloria Lau, where she said that also the insiders thems…”
Matei Zaharia Jan 2, 2019 ▶ 4:55 a16z Podcast | A Conversation With the Inventor of Spark
Assertion Not publicly verifiable
Zaharia: Apache Spark is the most active open-source data processing project
“It's actually the most active open source project in data processing in general as far as we can tell.”
Matei Zaharia Jan 2, 2019 ▶ 9:15 a16z Podcast | A Conversation With the Inventor of Spark
Insight
Zaharia: Mentoring open-source contributors is slower initially but scales community
“At the beginning, you know, if you're someone working on it every day and someone comes in and wants help to, you know, to get some idea in, it's always faster for you to do it yourself than to help this other person. But you have to do that. You have to Help …”
Matei Zaharia Jan 2, 2019 ▶ 12:09 a16z Podcast | A Conversation With the Inventor of Spark
Insight
Zaharia: Testing infrastructure is essential to maintain development speed in open source
“The third thing I need that's really important to keep a project moving quickly is just really great infrastructure for testing, checking the quality, making sure that it continues to be good. And by investing in this kind of infrastructure, much the same as y…”
Matei Zaharia Jan 2, 2019 ▶ 13:10 a16z Podcast | A Conversation With the Inventor of Spark
Assertion Contradicted
Zaharia: Databricks includes all engine improvements in open-source Spark
“It's the same Spark that anyone else gets in the open source. All the libraries, all the improvements we put into the engine, you can just download them and run them yourselves. Or if you want you know, you can talk to a vendor that provides support. Support o…”
Matei Zaharia Jan 2, 2019 ▶ 18:31 a16z Podcast | A Conversation With the Inventor of Spark
Assertion Not checkable as stated
Zaharia: Apache Spark is easier to use than prior big data systems
“So Spark is software for processing large volumes of data on a cluster, and the things that make it unique are, first of all, it has a very powerful programming model that lets you do many kinds of advanced analytics and processing, such as machine learning or…”
Matei Zaharia Jan 2, 2019 ▶ 0:29 a16z Podcast | A Conversation With the Inventor of Spark
Assertion Supported
Zaharia: MapReduce was created by Google for nightly web indexing
“MapReduce initially came out of Google, where it was used for web indexing, and the whole point was, I will run this giant job every night, and in the morning, it's built a new index of the web.”
Matei Zaharia Jan 2, 2019 ▶ 3:49 a16z Podcast | A Conversation With the Inventor of Spark
Insight
Zaharia: Ad hoc data work requires iterative processing over batch runs
“When you work with data, you want to ask multiple questions repeatedly when you're doing ad hoc exploration of the data, as opposed to, you know, when you have a certain application that, you know, okay, I'm just going to run this every night.”
Matei Zaharia Jan 2, 2019 ▶ 4:27 a16z Podcast | A Conversation With the Inventor of Spark
Assertion Supported
Zaharia: Toyota uses Spark to analyze social media feedback on cars
“One of the coolest ones I saw was a talk from Toyota about how they use Spark to improve, you know, to basically look at social media feedback, what people are writing about their cars, and figure out things like, oh, is there a problem with the brakes on the …”
Matei Zaharia Jan 2, 2019 ▶ 6:44 a16z Podcast | A Conversation With the Inventor of Spark
Assertion Supported
Legacy Hadoop tools like Hive, Pig, and Mahout now run on Spark
“So in particular you know, one of the things we saw is many of the projects that were built on top of Hadoop, such as Hive, which is a SQL processing at scale and Pig and Mahout for machine learning are starting to run on top of Spark as well, so that users of…”
Matei Zaharia Jan 2, 2019 ▶ 13:59 a16z Podcast | A Conversation With the Inventor of Spark
Opinion
Zaharia: Third-party integrations are Spark's most valuable asset for users
“So I think even beyond the activity happening in Spark itself, these projects on top and on the side are one of the most valuable things for the users.”
Matei Zaharia Jan 2, 2019 ▶ 14:44 a16z Podcast | A Conversation With the Inventor of Spark
Disclosure
Zaharia: Spark was originally designed to run Netflix Prize recommendation algorithms
“So it's actually one of the applications that I first tried to support in Spark was you know, the recommendation algorithm he was working on.”
Matei Zaharia Jan 2, 2019 ▶ 16:41 a16z Podcast | A Conversation With the Inventor of Spark
Disclosure
Matei Zaharia interned at Facebook in 2007 when it had 300 employees
“I was a PhD student at UC Berkeley, and we actually started working with Hadoop users back in 2007. And I did, for example, an internship at Facebook when Facebook was only about 300 people and they were just starting to set up Hadoop.”
Matei Zaharia Jan 2, 2019 ▶ 1:44 a16z Podcast | A Conversation With the Inventor of Spark

The other half of the tape: Matei Zaharia's own voice is left out of every number here. Other people bring the name up 6 times in 3 episodes on the a16z Podcast. every mention, with the transcript →

Who brings them up most Ben Horowitz 3Michael Franklin 1Ion Stoica 1Ali Ghodsi 1

Every mention by year

tap a year for its mentions
0031522019202020212022202320242025episodesmentions
0122019202020212022202320242025episodes it came up in
001.312.522019202020212022202320242025episodesmentions per episode
2025 5 mentions in 2 episodes 3 per episode
2019 1 mention in 1 episode

Appearances (1)

EpisodeDateSpeaking time
a16z Podcast | A Conversation With the Inventor of Spark Jan 2, 2019 13m
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.