The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

clear all ✕

why aren't all 11 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 0 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Opinion
Borgman: Apache Spark is not built to support high concurrency workloads
“Anyone who's trying to do high concurrency will usually realize that Spark is not built for high concurrency.”
Justin Borgman Mar 19, 2019 ▶ 18:41 Optionality in Data Architecture // Justin Borgman, Starburst Data (FirstMark's Data Driven NYC)
Opinion
Elprin: Apache Spark still generates more industry hype than actual business value
“I think there's still more hype around Spark than actual value extraction from it.”
Nick (Domino Data Lab) Nov 9, 2016 ▶ 15:56 Lessons Learned from Advanced Data Science Orgs // Domino Data Lab [FirstMark's Data Driven]
Assertion Not checkable as stated
Stefan Groschupf: Apache Flink is already faster than Spark
“And what's really interesting is Flink is already faster than Spark.”
Stefan Groschupf Dec 17, 2015 ▶ 10:35 The Acceleration of Innovation in Big Data w/ Stefan Groschupf, Datameer
Assertion Not checkable as stated
Handy: 100x more people write SQL than Spark or Scala
“There are actually two orders of magnitude more human beings on the planet that can write SQL than can write Spark or Scala or whatever.”
Tristan Handy Jun 28, 2022 ▶ 7:34 The Next Layer of the Modern Data Stack | dbt's Tristan Handy
Opinion
Bob Muglia: Apache Spark scenarios are complementary to Snowflake
“Spark, I think, is being used very, very broadly for advanced analytics, machine learning, in some cases for streaming data, and those scenarios are all very, very complimentary to Snowflake.”
Bob Muglia Apr 9, 2018 ▶ 18:28 Fireside Chat with Bob Muglia, CEO at Snowflake (FirstMark's Data Driven)
Assertion Not checkable as stated
Kandel: Hand-coding tools remain the most common data preparation method
“So probably still today the most common is using kind of hand coding tools so programming languages, Python, Spark SAS and so on.”
Sean Kandel Jul 13, 2017 ▶ 6:41 Three Loops of Analytics Efficiency // Sean Kandel, Trifacta (FirstMark's Data Driven)
Assertion Not checkable as stated
Barclays accelerated Spark jobs from hours to seconds using Alluxio
“They used Tachyon to accelerate Spark jobs from hours to seconds”
Haoyuan Li Apr 13, 2016 ▶ 12:03 A Virtual Distributed Storage System // Haoyuan Li, Alluxio (Hosted by FirstMark)
Disclosure
Ghodsi: Databricks re-wrote Apache Spark's execution engine in C++
“Two or three years ago, we set out to re-implement all Spark in C++ in what we call the really, really fast, what's called MPP engine, Massive Apparel Processing Engine.”
Ali Ghodsi May 24, 2021 ▶ 23:12 Fireside Chat: Ali Ghodsi (Founder & CEO, Databricks) with Matt Turck (Partner, FirstMark)
Assertion Not checkable as stated
Uber engineers frequently crashed Kafka clusters with unthrottled Spark executor writes
“Kafka was, in general, like, a nice way where people used to pipe the results of, like, their Spark jobs. But often cases, what they do is, like, they hit Kafka hard and bring Kafka down because they're trying to, like, actually send data from, like, hundred e…”
Praveen Murugesan Sep 30, 2016 ▶ 10:21 The Uber Big Data Story // Praveen Murugesan, Uber (Data Driven NYC / FirstMark)
Disclosure
AXA US builds its analytics infrastructure on a Cloudera Hadoop stack
“So the one that we've got here in the U.S., it's primarily a Cloudera Hadoop stack that we've used a blueprint that was essentially blessed by our brethren over in French, in France and with that, we've got R and Python and Spark. We use some Dataiku along the…”
Louis DiModugno May 23, 2016 ▶ 10:44 Big Data in Insurance // Louis DiModugno, Chief Data Officer at AXA US [FirstMark's Data Driven]
Disclosure
Essas: OpenTable routes event streams through Kafka, Cassandra, and Spark
“All of our events flowing through Kafka, they've been populated into Cassandra, which then we run Spark instances that kind of model on top of the data.”
Joseph Essas Jun 19, 2015 ▶ 2:32 Joseph Essas, OpenTable // Mining Diner Talk (Hosted by FirstMark Capital)
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.