why aren't all 11 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Opinion
Borgman: Apache Spark is not built to support high concurrency workloads
“Anyone who's trying to do high concurrency will usually realize that Spark is not built for high concurrency.”
Opinion
Elprin: Apache Spark still generates more industry hype than actual business value
“I think there's still more hype around Spark than actual value extraction from it.”
Assertion Not checkable as stated
Stefan Groschupf: Apache Flink is already faster than Spark
“And what's really interesting is Flink is already faster than Spark.”
Assertion Not checkable as stated
Handy: 100x more people write SQL than Spark or Scala
“There are actually two orders of magnitude more human beings on the planet that can write SQL than can write Spark or Scala or whatever.”
Opinion
Bob Muglia: Apache Spark scenarios are complementary to Snowflake
“Spark, I think, is being used very, very broadly for advanced analytics, machine learning, in some cases for streaming data, and those scenarios are all very, very complimentary to Snowflake.”
Assertion Not checkable as stated
Kandel: Hand-coding tools remain the most common data preparation method
“So probably still today the most common is using kind of hand coding tools so programming languages, Python, Spark SAS and so on.”
Assertion Not checkable as stated
Barclays accelerated Spark jobs from hours to seconds using Alluxio
“They used Tachyon to accelerate Spark jobs from hours to seconds”
Disclosure
Ghodsi: Databricks re-wrote Apache Spark's execution engine in C++
“Two or three years ago, we set out to re-implement all Spark in C++ in what we call the really, really fast, what's called MPP engine, Massive Apparel Processing Engine.”
Assertion Not checkable as stated
Uber engineers frequently crashed Kafka clusters with unthrottled Spark executor writes
“Kafka was, in general, like, a nice way where people used to pipe the results of, like, their Spark jobs. But often cases, what they do is, like, they hit Kafka hard and bring Kafka down because they're trying to, like, actually send data from, like, hundred e…”
Disclosure
AXA US builds its analytics infrastructure on a Cloudera Hadoop stack
“So the one that we've got here in the U.S., it's primarily a Cloudera Hadoop stack that we've used a blueprint that was essentially blessed by our brethren over in French, in France and with that, we've got R and Python and Spark. We use some Dataiku along the…”
Disclosure
Essas: OpenTable routes event streams through Kafka, Cassandra, and Spark
“All of our events flowing through Kafka, they've been populated into Cassandra, which then we run Spark instances that kind of model on top of the data.”