Disclosure certainty 5/5 debate potential 1/5

Stoica: Databricks offers only cloud services, not a custom Spark distribution

Ion Stoica · Ion Stoica, Databricks // Creating Apache Spark // Data Driven NYC (FirstMark Capital) · Apr 2, 2015 · at 17:30

Ion Stoica, co-founder and CEO of Databricks, outlines Databricks' commercial product strategy relative to open-source software distributions.

0:00 / 0:03exact quote · 3.8s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“We don't have our own distribution. We provide only the service.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Ion Stoica

Prediction Not checkable as stated
Stoica predicts big data infrastructure will eventually unify around one system
“Now, I do think that looking forward you are going to see more and more of this unification. This happens in many other industries and technologies. I do think this will happen in big data because it's so much easier if you have only one system than if you nee…”
Ion Stoica Apr 2, 2015 ▶ 21:53 Ion Stoica, Databricks // Creating Apache Spark // Data Driven NYC (FirstMark Capital)
Assertion Supported
Stoica: Apache Spark won the TerraSort benchmark processing data out of memory
“Just October last year, we had this we won this kind of TerraSort benchmark. And in those, in that benchmark, the data, it's not in memory. Right? It's SSDs and so forth.”
Ion Stoica Apr 2, 2015 ▶ 8:04 Ion Stoica, Databricks // Creating Apache Spark // Data Driven NYC (FirstMark Capital)
Assertion Supported
Stoica: Apache Spark supports all major data workloads with one engine
“While we spark, You can use only one engine and only one API to support all these workloads.”
Ion Stoica Apr 2, 2015 ▶ 12:00 Ion Stoica, Databricks // Creating Apache Spark // Data Driven NYC (FirstMark Capital)
Assertion Supported
Stoica: Hadoop's HDFS read/write cycle crippled early iterative machine learning
“If you look at the machine learning, it's, fundamentally, it's an iterative algorithm, and every iteration is turned into a Hadoop job. So between the iteration, you write the data and read the data from HDFS, so that's why it's very slow.”
Ion Stoica Apr 2, 2015 ▶ 4:08 Ion Stoica, Databricks // Creating Apache Spark // Data Driven NYC (FirstMark Capital)
Opinion
Stoica: Hadoop remains a very great batch processing engine
“Hadoop is still a very great, ah, batch engine.”
Ion Stoica Apr 2, 2015 ▶ 11:23 Ion Stoica, Databricks // Creating Apache Spark // Data Driven NYC (FirstMark Capital)
Assertion Not checkable as stated
Stoica: Spark Streaming and Storm have roughly similar throughput
“They are roughly, you know, in terms of the throughput, they are roughly similar.”
Ion Stoica Apr 2, 2015 ▶ 12:51 Ion Stoica, Databricks // Creating Apache Spark // Data Driven NYC (FirstMark Capital)
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.