Ion Stoica

Professor of Computer Science, UC Berkeley · 1 appearance on the record.

computed by AI from the episodes · how this works → · full disclaimer →

academicscientistfounderexecutiveengineerpeople.eecs.berkeley.edu/~istoica ↗Wikipedia ↗

Stoica is a computer science professor at UC Berkeley who co-created Apache Spark, Apache Mesos, and Ray. He also co-founded enterprise infrastructure companies Databricks and Anyscale, as well as AI evaluation platform LMArena.

14statements → 11claims → 8claims resolved → 100%fully supported → 4.14/5average certainty → 1.79/5average debate potential → 3said about them ↓

8 supported 0 partly supported 0 contradicted 3 not checkable as stated how the 11 claims stand · each chip opens the sources

1 prediction · 10 assertions · 1 opinion · 2 disclosures · every statement was checked. The prediction and assertions are the 11 claims: statements the public record can support or contradict. 8 are resolved, and 3 name no date, number or outcome precise enough to check. Everything else (opinions, insights, what ifs, disclosures) can never be settled by the record, so it carries no assessment.

The record, in short

What the tape says about how Ion argues and how the claims held up. Everything they said, and everything said about them, is in the tabs below.

Their most notable supported claim

Assertion Supported
Stoica: Apache Spark won the TerraSort benchmark processing data out of memory
“Just October last year, we had this we won this kind of TerraSort benchmark. And in those, in that benchmark, the data, it's not in memory. Right? It's SSDs and so forth.”
Ion Stoica Apr 2, 2015 ▶ 8:04 Ion Stoica, Databricks // Creating Apache Spark // Data Driven NYC (FirstMark Capital)

Expressed certainty vs assessment result

none yet certainty 1
none yet certainty 2
none yet certainty 3
100% certainty 4
100% certainty 5

weighted support: a fully supported claim counts one, a partly supported claim counts half. Each filled bar is clickable and opens exactly those claims; "none yet" means nothing said at that certainty level has resolved yet

How they sound: speaking style how? →

207 words/min while actually speaking · 58.3 um and uh per 1k words

No argument clarity score for Ion Stoica: only 1 usable question→answer exchange on raw tape (a fair score needs 8+). We do not score a sample that small. Roundtable and news formats yield far fewer direct exchanges than interviews.

Measured by listening to the audio itself: 3,329 words across 1 episode of raw-level tape, transcribed verbatim with every um and uh kept, each one attributed only where the alignment onto our timed stream is unambiguous. These are measurements of speaking style. We do not rank them: across this corpus, fluency and argument quality are nearly uncorrelated (ρ≈0.2), and smooth talking does not signal clear thinking. How it's measured →

Everything Ion Stoica said on the MAD Podcast that made the record, most notable first. Filter by type, assessment or year in the ledger →

Prediction Not checkable as stated
Stoica predicts big data infrastructure will eventually unify around one system
“Now, I do think that looking forward you are going to see more and more of this unification. This happens in many other industries and technologies. I do think this will happen in big data because it's so much easier if you have only one system than if you nee…”
Ion Stoica Apr 2, 2015 ▶ 21:53 Ion Stoica, Databricks // Creating Apache Spark // Data Driven NYC (FirstMark Capital)
Assertion Supported
Stoica: Apache Spark won the TerraSort benchmark processing data out of memory
“Just October last year, we had this we won this kind of TerraSort benchmark. And in those, in that benchmark, the data, it's not in memory. Right? It's SSDs and so forth.”
Ion Stoica Apr 2, 2015 ▶ 8:04 Ion Stoica, Databricks // Creating Apache Spark // Data Driven NYC (FirstMark Capital)
Assertion Supported
Stoica: Apache Spark supports all major data workloads with one engine
“While we spark, You can use only one engine and only one API to support all these workloads.”
Ion Stoica Apr 2, 2015 ▶ 12:00 Ion Stoica, Databricks // Creating Apache Spark // Data Driven NYC (FirstMark Capital)
Assertion Supported
Stoica: Hadoop's HDFS read/write cycle crippled early iterative machine learning
“If you look at the machine learning, it's, fundamentally, it's an iterative algorithm, and every iteration is turned into a Hadoop job. So between the iteration, you write the data and read the data from HDFS, so that's why it's very slow.”
Ion Stoica Apr 2, 2015 ▶ 4:08 Ion Stoica, Databricks // Creating Apache Spark // Data Driven NYC (FirstMark Capital)
Opinion
Stoica: Hadoop remains a very great batch processing engine
“Hadoop is still a very great, ah, batch engine.”
Ion Stoica Apr 2, 2015 ▶ 11:23 Ion Stoica, Databricks // Creating Apache Spark // Data Driven NYC (FirstMark Capital)
Assertion Not checkable as stated
Stoica: Spark Streaming and Storm have roughly similar throughput
“They are roughly, you know, in terms of the throughput, they are roughly similar.”
Ion Stoica Apr 2, 2015 ▶ 12:51 Ion Stoica, Databricks // Creating Apache Spark // Data Driven NYC (FirstMark Capital)
Assertion Not checkable as stated
Stoica: Spark Streaming micro-batching is unfit for high-frequency trading latency
“It's very hard to, you know, if you want millisecond latency from the time you, ah, the data entered in the system until you get the result, it's very hard to get. You can get latencies of several hundreds of milliseconds, but milliseconds, very hard. Ah, you,…”
Ion Stoica Apr 2, 2015 ▶ 13:40 Ion Stoica, Databricks // Creating Apache Spark // Data Driven NYC (FirstMark Capital)
Assertion Supported
Stoica: Apache Spark has exceeded 500 active open-source contributors
“We exceeded 500 contributors. It is the most active big data project right now, Spark.”
Ion Stoica Apr 2, 2015 ▶ 16:47 Ion Stoica, Databricks // Creating Apache Spark // Data Driven NYC (FirstMark Capital)
Assertion Supported
Stoica: Databricks still provides the majority of Apache Spark open-source contributions
“A lot of contributions, still the majority of contributions comes from Databricks.”
Ion Stoica Apr 2, 2015 ▶ 17:00 Ion Stoica, Databricks // Creating Apache Spark // Data Driven NYC (FirstMark Capital)
Assertion Supported
Stoica: Early Hadoop was limited to batch processing
“So at that point, in big data space we there was Hadoop just started, but of course that was, by, back then it was mostly, you know, batch, computation, so you could do historical analysis, but not much more than that.”
Ion Stoica Apr 2, 2015 ▶ 3:03 Ion Stoica, Databricks // Creating Apache Spark // Data Driven NYC (FirstMark Capital)
Disclosure
Stoica: Apache Spark was created for iterative machine learning and interactive queries
“And Spark was, ah, you know, we targeted first some workloads which are not covered by Hadoop, and from all this experience I mentioned earlier, we look at iterative, iterative computations to support machine learning, as well as interactive computation, right…”
Ion Stoica Apr 2, 2015 ▶ 5:45 Ion Stoica, Databricks // Creating Apache Spark // Data Driven NYC (FirstMark Capital)
Disclosure
Stoica: Databricks offers only cloud services, not a custom Spark distribution
“We don't have our own distribution. We provide only the service.”
Ion Stoica Apr 2, 2015 ▶ 17:30 Ion Stoica, Databricks // Creating Apache Spark // Data Driven NYC (FirstMark Capital)
Assertion Supported
Stoica: Berkeley's Succinct enables query processing on compressed data
“There is a related project that, ah, Berkeley is called, ah, succinct, ah, which actually go even more than, you know, beyond that. It's, ah it's a project that allows you to, you know, provide allow you to have, you know, query processing on the compressed da…”
Ion Stoica Apr 2, 2015 ▶ 19:31 Ion Stoica, Databricks // Creating Apache Spark // Data Driven NYC (FirstMark Capital)
Assertion Supported
Stoica: Apache Spark originally ran on Mesos before adding YARN and standalone support
“As originally was built to run on top of Mesos. Today is working on Yarn, working, you know, standalone, and is working also in addition to HDFS, you know, imports and exports data to many other data sources.”
Ion Stoica Apr 2, 2015 ▶ 7:07 Ion Stoica, Databricks // Creating Apache Spark // Data Driven NYC (FirstMark Capital)

The other half of the tape: Ion Stoica's own voice is left out of every number here. Other people bring the name up 3 times in 2 episodes on the MAD Podcast. every mention, with the transcript →

Who brings them up most Haoyuan Li 2Matt Turck 1

Every mention by year

tap a year for its mentions
001121201620172018201920202021episodesmentions
011201620172018201920202021episodes it came up in
0010.521201620172018201920202021episodesmentions per episode

Appearances (1)

EpisodeDateSpeaking time
Ion Stoica, Databricks // Creating Apache Spark // Data Driven NYC (FirstMark Capital) Apr 2, 2015 19m
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.