Ion Stoica

Professor of Computer Science, UC Berkeley · 2 appearances on the record.

computed by AI from the episodes · how this works → · full disclaimer →

academicscientistfounderexecutiveengineerpeople.eecs.berkeley.edu/~istoica ↗Wikipedia ↗

Stoica is a computer science professor at UC Berkeley who co-created Apache Spark, Apache Mesos, and Ray. He also co-founded enterprise infrastructure companies Databricks and Anyscale, as well as AI evaluation platform LMArena.

12statements → 9claims → 2claims resolved → 4/5average certainty → 1.67/5average debate potential → 15said about them ↓

2 supported 0 partly supported 0 contradicted 1 not yet assessed 6 not checkable as stated how the 9 claims stand · each chip opens the sources

1 prediction · 8 assertions · 2 insights · 1 disclosure · every statement was checked. The prediction and assertions are the 9 claims: statements the public record can support or contradict. 2 are resolved, 1 is not yet assessed, and 6 name no date, number or outcome precise enough to check. Everything else (opinions, insights, what ifs, disclosures) can never be settled by the record, so it carries no assessment.

The record, in short

What the tape says about how Ion argues and how the claims held up. Everything they said, and everything said about them, is in the tabs below.

Their most notable supported claim

Assertion Supported
Stoica: UC Berkeley's AMPLab generates zero patents and commits to open source
“The other thing about what AMLAB and all these labs have done is that they take a very strong stand about being open source and we generate no patents.”
Ion Stoica Jan 2, 2019 ▶ 7:32 a16z Podcast | A New Lab Rises

Expressed certainty vs assessment result

none yet certainty 1
none yet certainty 2
none yet certainty 3
100% certainty 4
100% certainty 5

weighted support: a fully supported claim counts one, a partly supported claim counts half. Each filled bar is clickable and opens exactly those claims; "none yet" means nothing said at that certainty level has resolved yet

How they sound: speaking style how? →

226 words/min while actually speaking · 18.4 um and uh per 1k words · 13.6 false starts per 1k · 23.2% of pauses land inside a clause

No argument clarity score for Ion Stoica: only 2 usable question→answer exchanges on raw tape (a fair score needs 8+). We do not score a sample that small. Roundtable and news formats yield far fewer direct exchanges than interviews.

Measured by listening to the audio itself: 7,767 words across 2 episodes of raw-level tape, transcribed verbatim with every um and uh kept, each one attributed only where the alignment onto our timed stream is unambiguous. These are measurements of speaking style. We do not rank them: across this corpus, fluency and argument quality are nearly uncorrelated (ρ≈0.2), and smooth talking does not signal clear thinking. How it's measured →

Everything Ion Stoica said on the a16z Podcast that made the record, most notable first. Filter by type, assessment or year in the ledger →

Assertion Not checkable as stated
Stoica: Top technical experts lack time to perform AI evaluation labeling
“You're getting some people from that area who are willing to do the labeling, but the best people are not willing. Fundamentally, they don't have time.”
Ion Stoica May 29, 2025 ▶ 9:51 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
Disclosure
Stoica: Apache Spark was created to prove Mesos's value
“So actually, Mesos was the first project we developed. This is what started the stack. And Mesos was by design to support multiple cluster computing framework. We started with Hadoop. And actually, one of the reasons we designed Spark, it's To show that it's m…”
Ion Stoica Jan 2, 2019 ▶ 4:03 a16z Podcast | A New Lab Rises
Assertion Not checkable as stated
Stoica: Many published AI research algorithms are hard to reproduce
“Many of the algorithms which are published are hard to reproduce.”
Ion Stoica Jan 2, 2019 ▶ 19:48 a16z Podcast | A New Lab Rises
Assertion Not checkable as stated
Stoica: Over 70% of daily LMArena prompts are completely unique
“And basically measures out how many more, you know, fresh prompts you have in one day compared to what you've seen in the past three months, right? And by a similarity score of something like 70, 75%, you have over 70 of these prompts are fresh.”
Ion Stoica May 29, 2025 ▶ 20:13 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
Assertion Supported
Stoica: UC Berkeley's AMPLab generates zero patents and commits to open source
“The other thing about what AMLAB and all these labs have done is that they take a very strong stand about being open source and we generate no patents.”
Ion Stoica Jan 2, 2019 ▶ 7:32 a16z Podcast | A New Lab Rises
Assertion Not checkable as stated
Stoica: Over 60% of UC Berkeley PhD applicants apply for AI
“When you look at the PhD students who applied to Berkeley, PhD applicants. It turns out that well over 60% are applying for AI.”
Ion Stoica Jan 2, 2019 ▶ 18:44 a16z Podcast | A New Lab Rises
Prediction Not checkable as stated
Stoica: Computing functionality will migrate bidirectionally between cloud and edge
“Things which are now done in the cloud is going to migrate some of the functionality for on the edge. Also, on, in, on the other side, things which are now done only at the edge, like self-driving cars will migrate, some of the functionality will migrate to th…”
Ion Stoica Jan 2, 2019 ▶ 23:09 a16z Podcast | A New Lab Rises
Assertion Open · timeframe May 2026
Stoica: Users grading answers to their own questions is the gold standard
“This is now from information retrieval field for decades, and it's called gold standard, when people evaluate the answer to their own questions. When an expert evaluates someone else questions, and the answer is called the silver, if I remember correctly.”
Ion Stoica May 29, 2025 ▶ 1:08:26 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
Assertion Not checkable as stated
Stoica: Facebook's early big data cluster had 80 nodes and three people
“When we started working with Facebook, Facebook, you know, has an entire cluster, big cluster for big data. It was 80 nodes. And their big data team was like three people.”
Ion Stoica Jan 2, 2019 ▶ 6:31 a16z Podcast | A New Lab Rises
Insight
Stoica: ML models degrade over time as real-world data evolves
“The models you developed on some data set, and for instance, the data or the queries are going to evolve over time. And because the environment or the world around you evolves, what you learned and which is embedded in the model may not be as relevant or as go…”
Ion Stoica Jan 2, 2019 ▶ 25:46 a16z Podcast | A New Lab Rises
Insight
Stoica: Silicon Valley proximity keeps UC Berkeley research anchored in industry problems
“Academia, it's allowing you to do more experimentation. It's set for that. And at Berkeley, we are in a privileged position. Of course, being close to the Silicon Valley, we have a lot of feedback. From the industry. So we are very anchored in what are the rea…”
Ion Stoica Jan 2, 2019 ▶ 1:15 a16z Podcast | A New Lab Rises
Assertion Supported
Stoica: UC Berkeley computer science research labs operate on five-year limits
“So the labs are like one rule is around five years.”
Ion Stoica Jan 2, 2019 ▶ 2:01 a16z Podcast | A New Lab Rises

The other half of the tape: Ion Stoica's own voice is left out of every number here. Other people bring the name up 15 times in 5 episodes on the a16z Podcast. every mention, with the transcript →

Who brings them up most Elena Burger 5Ali Ghodsi 4Michael Franklin 3Ben Horowitz 2Chris Granger 1

Every mention by year

tap a year for its mentions
0031512017201820192020202120222023202420252026episodesmentions
0112017201820192020202120222023202420252026episodes it came up in
002.50.5512017201820192020202120222023202420252026episodesmentions per episode
2026 5 mentions in 1 episode
2025 5 mentions in 1 episode
2024 1 mention in 1 episode
2019 3 mentions in 1 episode
2017 1 mention in 1 episode

Appearances (2)

EpisodeDateSpeaking time
Beyond Leaderboards: LMArena’s Mission to Make AI Reliable May 29, 2025 27m
a16z Podcast | A New Lab Rises Jan 2, 2019 16m
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.