Xin: Single HTAP database engines compromise ecosystem compatibility and performance
“This is sort of the holy grail of database engineering is, why not build a single system that can do both of this? But it ends up just being a lot of compromises. And one, I think one of the first issue is that, hey, each, they say Postgres has a massive ecosy…”
Needham: SPARC reactor targets 100MW thermal power with 10x energy gain
“That'll produce about a hundred megawatts of thermal power. So a lot like 500 to a thousand times what the national ignition facility did, like a substantial amount of power. And his design point is actually a gain of like 10, so 10 times, actually it's 11 mor…”
Needham: CFS ARC commercial plant will generate 400MW electric and 1GW thermal
“Spark is the, you know, the basis of our power plant, which we call ARC, and that'll be a 400 megawatt electric plant, a one gigawatt thermal plant”
Ury: Early Romantic Chemistry Is Often Anxiety Triggered by Ambiguity
“The second myth is that if you feel the spark, it's definitely a good thing, and actually, sometimes certain people give us the spark because we actually don't know how they feel about us, so sort of that hot, cold feeling makes us wonder, does he like me? Doe…”
The VC Industry Faces a Reckoning With Many Firms Shutting Down
“I think there will be a lot of folks that go away in this cycle. I watched it in the last go round. You know, Spark and USV and all these firms didn't exist before the last cycle. And they really made their names coming out of the O eight crisis. And on the ot…”
Mason: Missive's comparison pages explicitly state product shortcomings to deter poor-fit users
“It's a honest take. When there are some shortcomings in Missive compared to the other product, we do state it clearly, and we don't want to bring users in who won't like Missive and would actually rather use Spark or, like, a more personal-oriented email app.”
Handy: 100x more people write SQL than Spark or Scala
“There are actually two orders of magnitude more human beings on the planet that can write SQL than can write Spark or Scala or whatever.”
Ghodsi: Databricks re-wrote Apache Spark's execution engine in C++
“Two or three years ago, we set out to re-implement all Spark in C++ in what we call the really, really fast, what's called MPP engine, Massive Apparel Processing Engine.”
Borgman: Apache Spark is not built to support high concurrency workloads
“Anyone who's trying to do high concurrency will usually realize that Spark is not built for high concurrency.”
Hirasaki: Spark acquisition vaulted Roche forward in gene therapy platform technology
“And then with the Spark acquisition same thing BMS, I'm sorry Roche was a laggard in the space, and now suddenly leaped ahead with not only a marketed gene therapy, but this AAV platform that it could use for delivering other gene therapies.”
Hirasaki: Roche Allows Acquired Companies Like Spark to Operate Semi-Autonomously
“Some big pharma companies will just tend to fully integrate the company whereas others will like Roche and Spark will allow the acquired companies to operate semi-autonomously.”
Leibert: Writing Visual Basic is more joyful than writing Go applications
“And frankly, it's more joyful to write Visual Basic than it is to write Go, right? That you actually achieve Business results. I mean, seriously, right? Like, writing a Spark job, you see the output right away, whereas if you write a large Go application, I me…”
Stanek: Business users prefer spreadsheet interfaces over Spark and Hadoop
“Some of the most frequently used kind of data analytics tools extremely basic, because they actually look and feel like sheet of paper, like two-dimensional sheet of paper, and you know, so that's the problem with analytics, that on one hand, we have, you know…”
Moghe: Spark and Hadoop do not replace existing data warehouses
“Spark doesn't subsume data warehousing. Hadoop doesn't subsume, you know, streaming. So they're just like different technologies for different jobs.”
Zaharia: Apache Spark is easier to use than prior big data systems
“So Spark is software for processing large volumes of data on a cluster, and the things that make it unique are, first of all, it has a very powerful programming model that lets you do many kinds of advanced analytics and processing, such as machine learning or…”
Zaharia: Apache Spark is the most active open-source data processing project
“It's actually the most active open source project in data processing in general as far as we can tell.”
Zaharia: Spark was originally designed to run Netflix Prize recommendation algorithms
“So it's actually one of the applications that I first tried to support in Spark was you know, the recommendation algorithm he was working on.”
Nguyen: Apache Spark Would Have Failed Earlier Due to Memory Costs
“Now Spark, if it was created six, five, six years before its time would have completely failed because memory was so much more expensive.”
Bob Muglia: Apache Spark scenarios are complementary to Snowflake
“Spark, I think, is being used very, very broadly for advanced analytics, machine learning, in some cases for streaming data, and those scenarios are all very, very complimentary to Snowflake.”
Analytics frameworks like Spark and Hadoop assume exclusive resource access
“Most things, Hadoop, Spark, Storm, they think they're running by themselves. And so they compete for resources in really interesting ways.”
Kandel: Hand-coding tools remain the most common data preparation method
“So probably still today the most common is using kind of hand coding tools so programming languages, Python, Spark SAS and so on.”
Particle's first Kickstarter failed after raising $125K on a $250K goal
“We had a 250,000 dollar goal. We raised 125,000, but Kickstarter is all or nothing. So that means we raised zero and that product never came to be.”
Elprin: Apache Spark still generates more industry hype than actual business value
“I think there's still more hype around Spark than actual value extraction from it.”
Uber engineers frequently crashed Kafka clusters with unthrottled Spark executor writes
“Kafka was, in general, like, a nice way where people used to pipe the results of, like, their Spark jobs. But often cases, what they do is, like, they hit Kafka hard and bring Kafka down because they're trying to, like, actually send data from, like, hundred e…”
Turck: Spark Made Real-Time Massive Data Processing Possible
“What those guys are doing would simply not have been possible to do up until, you know, two or three years ago when the emergence of frameworks like Spark made possible, what was an alibi, it became possible to be processing You know, there's massive amounts o…”
AXA US builds its analytics infrastructure on a Cloudera Hadoop stack
“So the one that we've got here in the U.S., it's primarily a Cloudera Hadoop stack that we've used a blueprint that was essentially blessed by our brethren over in French, in France and with that, we've got R and Python and Spark. We use some Dataiku along the…”
Barclays accelerated Spark jobs from hours to seconds using Alluxio
“They used Tachyon to accelerate Spark jobs from hours to seconds”
Stefan Groschupf: Apache Flink is already faster than Spark
“And what's really interesting is Flink is already faster than Spark.”
Essas: OpenTable routes event streams through Kafka, Cassandra, and Spark
“All of our events flowing through Kafka, they've been populated into Cassandra, which then we run Spark instances that kind of model on top of the data.”