Apache Hive, every mention
33 scenes · ← back to Apache Hive
tap a year for its mentions
every year anyone Matt Turck 7Praveen Murugesan 5Dave Burgess 4Justin Borgman 3Todd Papaioannou 2Stefan Groschupf 2Mike Driscoll 2M.C. Srivas 2Ion Stoica 2Ben Rogojan (Seattle Data Guy) 2
Verbatim, from the transcripts: the passages where Apache Hive comes up
Understanding Data Engineering in 2025 | Ben Rogojan, Seattle Data Guy
- ▶ 40:39 Ben Rogojan (Seattle Data Guy) Um, we still occasionally had hive jobs. 2 times in the scene
AI at ZoomInfo: Superpowering GTM teams | Ali Dasdan, CTO, ZoomInfo
- ▶ 22:42 Ali Dasdan We are lucky that, yeah, we created an HBase, sort of HHive-like system ourselves, uh, with hundred petabytes of size of data, right?
A Conversation with Chris Wiggins - Author of "How Data Happened"
- ▶ 20:10 Chris Wiggins Um, it was, so when I showed up at the New York Times in 2013, if you wanted to get your hands on data, you needed to write your own MapReduce jobs in Hive and hit buckets of unstructured JSON sitting in S three.
Data Visibility & Control | BigID Co-Founder & CEO Dimitri Sirota
- ▶ 8:57 Dimitri Sirota Now, they may differ, so for instance, for Hadoop, we could do, like, MapReduce, uh, we could do, um, direct, uh, connectivity to, um, uh, uh, Hive or Spark we could use, but they all use native protocols to scan the underlying system.
Behavioral Data Creation for AI | Snowplow Co-Founder & CEO Alex Dean
Fireside Chat: Ali Ghodsi (Founder & CEO, Databricks) with Matt Turck (Partner, FirstMark)
- ▶ 17:39 Ali Ghodsi Yeah, actually, the four technological breakthroughs that kind of happened at the same time, 2016, 17, at the same time, the one we contributed was Delta Lake, there was Hootie, there was uh, Hive acid, and there was icebergs.
Fireside Chat: Dave Burgess (Head of Data Engineering, Pinterest) w/ Matt Turck (Partner, FirstMark)
- ▶ 5:58 Dave Burgess And so we query Intelli, uh, to Presto, to Hive, uh, to Spark SQL, uh, to MySQL, but you can do it with other engines too.
- ▶ 8:21 Dave Burgess We used to use Hive, but we're migrating of Hive to Smart SQL, so we just have the, the two main engines. 2 times in the scene
- ▶ 31:08 Dave Burgess They usually either Spark or Hive or Presto jobs or Spark SQL and, uh, just process the data in every step and, and persist the data back to S three along the way.
Data Observability and Pipelines: OpenLineage and Marquez
- ▶ 10:13 Julien Le Dem Or the SQL and Hadoop types like Hive and Presto.
- ▶ 21:00 Matt Turck Julien, what are your thoughts on where the Hive Metastore, uh, fits in the metadata and lineage story? 7 times in the scene
Optionality in Data Architecture // Justin Borgman, Starburst Data (FirstMark's Data Driven NYC)
- ▶ 1:54 Justin Borgman So, uh, I'm sure most of you have heard of Hive. 3 times in the scene
Making Big Data Accessible Using the Cloud // Ashish Thusoo, Qubole [FirstMark's Data Driven]
- ▶ 3:07 Ashish Thusoo So, um, you know, big data, essentially the emergence of these new systems, we hear about systems like Hadoop, Spark, Hive, and so on and so forth.
The Uber Big Data Story // Praveen Murugesan, Uber (Data Driven NYC / FirstMark)
- ▶ 4:41 Praveen Murugesan And, ah, on top of HDFS, we basically have, like, Spark, and, ah, Presto, and Hive.
- ▶ 7:18 Praveen Murugesan So, a few things, like I talked about, strict schema management, so we actually built, like, a central schema repository which is used for schema management, and then, ah, we unlocked, like, a whole bunch of new tools with, like, data on…
- ▶ 15:21 Praveen Murugesan and then it's provided as, like, a Hive UDF. 3 times in the scene
A Virtual Distributed Storage System // Haoyuan Li, Alluxio (Hosted by FirstMark)
- ▶ 12:41 Haoyuan Li They run a streaming workload on top of it, and they also use Aluxio to share the data efficiently between streaming processing and their batch processing, and use Hive in this particular use case.
A Fireside Chat with MapR CTO M.C. Srivas (Data Driven NYC / FirstMark)
- ▶ 2:26 M.C. Srivas So I said, Haru, I can't even spell this thing, and then I looked more and had even more strange names like Hive, and Hive actually didn't exist, it was Pig, and I said, what is Pig, and then it has something else, and I was just, alright,…
- ▶ 20:58 M.C. Srivas We introduced JSON to Hadoop and Spark and in MapReduce and in Hive and everywhere.
The Acceleration of Innovation in Big Data w/ Stefan Groschupf, Datameer
- ▶ 6:19 Stefan Groschupf Anybody's using Hyphia?
- ▶ 22:08 Stefan Groschupf Therefore, I do not believe in kind of the schema on write, hive approach, but very much on the schema on read approach.
Ion Stoica, Databricks // Creating Apache Spark // Data Driven NYC (FirstMark Capital)
- ▶ 4:59 Ion Stoica Um, and two of newly hired, if I remember correctly, from Yahoo, they were just starting the work on Hive.
- ▶ 22:49 Ion Stoica Um, you know, it's, ah, ah, in terms of the parser, it's a high parser, and recently we also released this, ah, DataFrames API, ah, ah, which is very similar with, you know, the Pandas, ah, which is very popular.
Michael Rubenstein and Catherine Williams, App Nexus // Data Driven #31 // Nov 2014
- ▶ 10:18 Catherine Williams We started using, you know, Hadoop, MapReduce, Hive.
John Rauser, Pinterest // Big Data at Pinterest // Data Driven NYC (Hosted by FirstMark Capital)
- ▶ 17:42 John Rauser Um, that's been a pretty substantial game changer for us, uh, moving away from Hive.
Tobi Knaup, Mesosphere // Data Driven #29 // Sep 2014 (Hosted by FirstMark Capital)
- ▶ 10:31 Tobi Knaup And a couple different pieces, you know, Flume, HDFS, Hive, Pig, pretty typical for, for a big data stack.
Ashish Thusoo, Qubole // Data Driven #26 // April 2014 (Hosted by FirstMark Capital)
- ▶ 1:21 Ashish Thusoo I was also one of the creators of Bachi Hive.
Vaclav Petricek, eHarmony // Data Driven NYC 19 // October 2013
- ▶ 12:37 Vaclav Petricek And then, uh, what we use is Hive for the data massaging, um, Impala also for some quick analytics, and, uh, RStudio,
Panel: Big data and VCs (Accel, IA Ventures, Data Collective) // Data Driven NYC #8// Oct 2012
Panel: Continuuity, Sailthru and Visual Revenue // Data Driven NYC #7 // June 2012
- ▶ 15:25 Todd Papaioannou But I really think of it as kind of like the core of a kernel of a distributed operating system, and if you look at the menagerie of things that exist in the Hadoop ecosystem, whether it's Hadoop or Hive or HBase or, you know, all of the…
- ▶ 20:08 Todd Papaioannou Facebook built Hive, right, because they needed a tool to sit on top of Hadoop, you know, to allow their business analysts to kind of sequel interface to, to this big data platform.
Panel: Metamarkets, Kaggle and Quid // Data Driven NYC #4 // Mar 2012
- ▶ 29:52 Mike Driscoll Uh, then if you've got Hive installed on your Hadoop cluster, they can run their queries and get their data, but it's critical that they have, you know, a certain level of data engineering expertise. 2 times in the scene
Kirill Sheynkman, RTP Ventures // Data Driven #3 // Feb 2012 (Interviewed by Matt Turck)
- ▶ 8:24 Kirill Sheynkman So, with no SQL databases, with, ah, Hadoop, Pig, and, and, and Hive and things like that,