MapReduce, every mention
64 scenes across 8 shows · ← back to MapReduce
the MAD Podcast 62
the a16z Podcast 24
the Y Combinator Startup Podcast 10
Acquired 4
Latent Space 2
TBPN 2
the Official SaaStr Podcast 1
All-In 1
every year every show
the MAD Podcast 62
the a16z Podcast 24
the Y Combinator Startup Podcast 10
Acquired 4
Latent Space 2
TBPN 2
the Official SaaStr Podcast 1
All-In 1
Verbatim, from the transcripts: passages where MapReduce comes up on the MAD Podcast, the a16z Podcast, the Y Combinator Startup Podcast, Acquired, Latent Space
Demis Is Out as DeepMind CEO, Revolut Founder Yacht Drama, AI’s Great Reverse Bank Run | Diet TBPN
- ▶ 2:20 John Coogan Uh, he has been, you know, very key to so many different, uh, Google projects, Integral and MapReduce, and, uh, actually being able to scale Google systems.
Jeff Dean: The 1% Rule for Building in AI · Y Combinator
- ▶ 0:24 Diana Hu So, um, you built MapReduce, Bigtable, TensorFlow, the TPU, Gemini, we could spend a whole hour on all the things you've done.
- ▶ 40:31 Jeff Dean The origin of MapReduce is another good example. 3 times in the scene
- ▶ 55:07 Diana Hu Now, one last thing, I'm pretty sure someone in this room or multiple people will eventually build something as consequential as you've done with MapReduce, TPU, distillation, et cetera, et cetera.
The Agent Cloud: Databricks’ Bet on the Future of AI — Matei Zaharia and Reynold Xin
- ▶ 55:43 Matei Zaharia Cause that's what Spark was, was for the large scale map reduce like stuff.
Weekly Recap: Gwyneth Paltrow Saves Astronomer, Figma’s IPO, Top AI Researchers Ranked
- ▶ 22:17 John Coogan Built, uh, MapReduce, Bigtable, TensorFlow.
Google Part I: Origins of Search. How the Best Business in Human History Happened (Audio) · Acquired
- ▶ 1:00:46 David Rosenthal He also implemented the first version of AdWords, built AdSense, rewrote the core search pipeline five times, co-invented and implemented Bigtable MapReduce, TensorFlow, and Gemini.
- ▶ 1:11:11 David Rosenthal And they built stuff like GFS, the Google file system, MapReduce.
- ▶ 1:11:16 David Rosenthal Yahoo would eventually feel like they needed to copy MapReduce to be competitive, and they would open source that as Hadoop, so people, Apache Hadoop, that is a Yahoo copy of Google's MapReduce.
- ▶ 2:58:33 David Rosenthal Big table, MapReduce, et cetera, et cetera.
Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
- ▶ 43:21 Ion Stoica MapReduce, Google file systems, all of that happening at Google.
How To Build The Future: Aravind Srinivas · Y Combinator
- ▶ 33:29 Aravind Srinivas They're not just like, oh, a page rank or like, uh, MapReduce or, uh, you know, all these advances that they made in like visual, like deep learning and, and, and like BERT, Transformers.
Why is everyone cloning Deep Research?
- ▶ 47:59 Shawn Wang You know, thing or whatever in the old Google days might be like MapReduce or, you know, whatever, but like it's a, it's a different scale and nature of work.
Dataiku's Secret to Scaling AI in Global Enterprises | Florian Douetteau, CEO, Dataiku
- ▶ 6:05 Florian Douetteau You had, uh, Google pushing new technologies, you had, like, those, uh, those days of MapReduce, those days of deep learning starting to work,
The Death of Big Data and Why It’s Time To Think Small | Jordan Tigani, CEO, MotherDuck
- ▶ 3:28 Jordan Tigani You know, after Google came out with, you know, MapReduce and, and GFS and Bigtable, kind of everybody's... 2 times in the scene
E156: Ivy League antisemitism, macro, SaaS recovery, Gemini, Figma deal delay + big Friedberg update
- ▶ 1:09:05 Chamath Palihapitiya So, TensorFlow, MapReduce, BigTable, Spanner, the guy is just an absolute animal.
A Conversation with Chris Wiggins - Author of "How Data Happened"
- ▶ 20:10 Chris Wiggins Um, it was, so when I showed up at the New York Times in 2013, if you wanted to get your hands on data, you needed to write your own MapReduce jobs in Hive and hit buckets of unstructured JSON sitting in S three.
Not Just Another Cloud Database | SurrealDB Co-Founders Jaime & Tobie Morgan Hitchcock
- ▶ 8:27 Jaime & Tobie Morgan Hitchcock We wanted to have real-time data that was constantly changing, but at the same time, without having to run the MapReduce or large-scale analytics processes and workloads on that data, we wanted to pull out that data in real-time for…
Data Visibility & Control | BigID Co-Founder & CEO Dimitri Sirota
- ▶ 8:57 Dimitri Sirota Now, they may differ, so for instance, for Hadoop, we could do, like, MapReduce, uh, we could do, um, direct, uh, connectivity to, um, uh, uh, Hive or Spark we could use, but they all use native protocols to scan the underlying system.
Things That Don't Scale, The Software Edition – Dalton Caldwell and Michael Seibel · Y Combinator
- ▶ 22:30 Dalton Caldwell Um, and this was the genesis for them to create MapReduce, which they wrote a paper about. 2 times in the scene
- ▶ 23:21 Dalton Caldwell Again, in terms of do things that don't scale, did they build MapReduce before they had any users? 2 times in the scene
$100 Million ARR Pivot: From Platform Product to Vertical Apps With Treasure Data CEO Kazuki Ohta
- ▶ 8:28 Kazuki Ohta And in terms of data warehouse, we built a really scalable columnar storage with the Hadoop MapReduce data processing job, really storing three hundred and sixty billion records, executing almost two million jobs in total.
Fireside Chat: Ali Ghodsi (Founder & CEO, Databricks) with Matt Turck (Partner, FirstMark)
- ▶ 2:50 Ali Ghodsi Every iteration of the data has to run a MapReduce job.
Dynamic Range Sharding with Spanner // Daniel Chia, Google Spanner (FirstMark's Data Driven NYC)
- ▶ 13:53 Daniel Chia And so when you spin up a MapReduce job with 50,000 workers and all try and write to the same little data range that previously had no data, that's not going to work very well because we haven't had time to split it yet.
a16z Podcast | A New Lab Rises
- ▶ 11:09 Ion Stoica If you think about the MapReduce and Google File System, which really started as a big data movement. 2 times in the scene
a16z Podcast | The Strategies and Tactics of Big
- ▶ 17:09 Steven Sinofsky If Google wants to replace MapReduce with some new thing, or if they are going to roll out, like, a new way of doing machine learning internally, they'll work on that for the same amount of time that Apple works on a phone.
a16z Podcast | The Storage Renaissance
- ▶ 13:28 Mike Matchett When you had Hadoop and MapReduce, they could partition certain categories of problems and run them in parallel, but not really a lot of machine learning algorithms.
a16z Podcast | A Conversation With the Inventor of Spark
- ▶ 0:57 unnamed speaker Before Spark, the most widely used system was probably MapReduce, which was, uh, uh, invented at Google and popularized through the, the open source Hadoop project. 7 times in the scene
a16z Podcast | The Cool Stuff Only Happens at Scale
- ▶ 1:41 Vijay Pande But you can't do MapReduce with everything. 2 times in the scene
- ▶ 6:21 Vijay Pande I think I've gotten people thinking about this, and MapReduce and things like that, those abstractions have played a huge role.
a16z Podcast | Making Sense of Big Data, Machine Learning, and Deep Learning
- ▶ 11:41 Christopher Nguyen Uh, so big compute is the first example of the compute you can think of is MapReduce. 8 times in the scene
a16z Podcast | Why the Datacenter Needs an Operating System
- ▶ 0:23 Steven Sinofsky There are things like Kafka and Spark and MapReduce and Cassandra, and it's super, super
Fireside Chat: Mike Tuchen, CEO of Talend (TLND) (FirstMark's Data Driven NYC)
- ▶ 15:38 Mike Tuchen So we built a, initially in our, in our initial technology, we built a, um, SQL, or excuse me, we did build a SQL one, but we started out with a Java compiler, um, and then we, um, did a SQL one, and then when Hadoop came along, um, we did…
3 Heretical Ideas on the Future of Data // Ajay Kulkarni, TimescaleDB (FirstMark's Data Driven NYC)
- ▶ 2:50 Ajay Kulkarni The relational databases they had in the past just couldn't scale for these workloads, so the new companies started developing the first big data non-relational systems, and this includes Google, who published the MapReduce paper and the…
Fireside Chat with Bob Muglia, CEO at Snowflake (FirstMark's Data Driven)
- ▶ 17:46 Bob Muglia And then the, the, the, the typical approaches that people have traditionally used with Hadoop, variations of MapReduce, um, are now being seen, I think, as, as, now that there are alternatives, such as Snowflake that are available, that…
Surveillance Platform for Banks // Mayur Thakur, Goldman Sachs (FirstMark's Data Driven)
- ▶ 16:53 Mayur Thakur We use MapReduce.
A Kafka-Powered Real-Time Streaming Platform // Neha Narkhede, Confluent [FirstMark's Data Driven]
- ▶ 23:52 Neha Narkhede Uh, a lot of these other stream processing systems, I didn't get a, uh, chance to get into the details. 3 times in the scene
Why Marketing is All About Data // Nitay Joffe, ActionIQ [FirstMark's Data Driven]
- ▶ 16:43 Nitay Joffe No SQL, no MapReduce, none of that kind of stuff.
A Fireside Chat with MapR CTO M.C. Srivas (Data Driven NYC / FirstMark)
- ▶ 20:58 M.C. Srivas We introduced JSON to Hadoop and Spark and in MapReduce and in Hive and everywhere.
The Acceleration of Innovation in Big Data w/ Stefan Groschupf, Datameer
- ▶ 1:16 Stefan Groschupf That's why we looked into the Google papers and implemented MapReduce and the HDFS, et cetera, et cetera.
- ▶ 4:38 Stefan Groschupf and then we built a MapReduce thing, and then, um, Peter Thiel, not sure, who knows Peter Thiel? 2 times in the scene
- ▶ 12:07 Stefan Groschupf Replacing MapReduce with TES or with Spark is really straightforward, um, especially if you use BI on top of Hadoop, um, because it basically automatically does it for you, but obviously you can't just rip out hardware and replace it,…
Oren Falkowitz, Area 1 Security // The Future of Cybersecurity (Hosted by FirstMark Capital)
- ▶ 4:19 Oren Falkowitz Maybe, you know, eight or nine years since Google published the Bigtable paper and the MapReduce paper and some of those early big data papers which have spawned, you know, this community and all the great technological revolution, now…
Ion Stoica, Databricks // Creating Apache Spark // Data Driven NYC (FirstMark Capital)
- ▶ 6:25 Matt Turck Uh, can you, uh, maybe help us understand, um, so this HDFS, which is the storage, uh, and MapReduce, which is, uh, the processing, 4 times in the scene
- ▶ 7:24 Matt Turck Yes, so perhaps more specifically, can you go through the, the, the key advantages of Spark over MapReduce? 4 times in the scene
Paul Dix, InfluxDB // Open-Source Time Series Database // Data Driven NYC (FirstMark Capital)
Chris Wiggins, NY Times // Data Science at The New York Times (Hosted by FirstMark Capital)
- ▶ 18:24 Chris Wiggins Elastic MapReduce, proper MapReduce in Java. 2 times in the scene
Mike Olson, Cloudera // The Cloudera Story (Hosted by FirstMark Capital)
- ▶ 4:55 Mike Olson Google published the Google File System and MapReduce papers, and like the entire database industry, I thought that this thing was a joke, man.
- ▶ 6:52 Matt Turck This Spark, this data flow, and, you know, maybe MapReduce is not what people need, and all of this is evolving.
- ▶ 9:33 Mike Olson These days when people talk about Hadoop, what they mean is HDFS and MapReduce, yeah, yarn for resource management. 5 times in the scene
- ▶ 14:39 Mike Olson The Hadoop that was then available, which was HDFS and MapReduce, was powerful, transformative.
- ▶ 18:18 Mike Olson The inventor of Hadoop, Google Ventures, the inventor of MapReduce and GFS took a stake.
Michael Rubenstein and Catherine Williams, App Nexus // Data Driven #31 // Nov 2014
- ▶ 10:18 Catherine Williams We started using, you know, Hadoop, MapReduce, Hive.
Mike Abbott, KPCB // Data Driven #30 // Oct 2014 (Hosted by FirstMark Capital)
- ▶ 12:56 Mike Abbott A third of the people that I recruited from Google was how could we, you know, let non-engineers run MapReduce jobs, ask more questions of the data, because if we let more non-engineers ask questions, there is a higher probability that we…
John Rauser, Pinterest // Big Data at Pinterest // Data Driven NYC (Hosted by FirstMark Capital)
- ▶ 1:16 John Rauser I'd walk them through an example of MapReduce at a level of simplicity that this audience would probably find comical. 4 times in the scene
Tobi Knaup, Mesosphere // Data Driven #29 // Sep 2014 (Hosted by FirstMark Capital)
- ▶ 3:33 Tobi Knaup And, um, so MapReduce, for example, was, uh, started as a, um, you know, Google wrote a paper about it.
Tasso Argyros, Aster Data // Data Driven NYC 24 // February 2014 (Hosted by FirstMark Capital)
- ▶ 2:47 Tasso Argyros A declarative language of SQL, right, and the procedural language of MapReduce.
- ▶ 14:39 Tasso Argyros In 2007, we had the brilliant idea of coming out with a SQL MapReduce Azure positioning, and of course, nobody knew what MapReduce back then, right? 6 times in the scene
Scott Sorensen, CTO of Ancestry.com // Data Driven 19 // October 2013 (Hosted by FirstMark Capital)
- ▶ 7:38 Scott Sorensen And MapReduce and R were introduced into our tool set at that time.
Mike Dauber, Battery Ventures // Data Driven NYC 19 // October 2013 (interviewed by Matt Turck)
- ▶ 2:52 Mike Dauber And then came 2003, 2004, and Google came out with their, you know, their big table MapReduce papers.
Panel: Continuuity, Sailthru and Visual Revenue // Data Driven NYC #7 // June 2012
- ▶ 15:17 Todd Papaioannou And two, MapReduce allows you to write a simple program that can scan all of that data.
Panel: Metamarkets, Kaggle and Quid // Data Driven NYC #4 // Mar 2012
- ▶ 49:39 unnamed speaker That it will scale, that it will map reduce, that whatever these words are.
Kirill Sheynkman, RTP Ventures // Data Driven #3 // Feb 2012 (Interviewed by Matt Turck)
- ▶ 4:56 Kirill Sheynkman It's, it's, it's pretty, ah, hairy stuff, but it's, it's a system that has APIs in Scala and in Java, and they're working on JavaScript and, and Python, where you can take a function and basically the MapReduce becomes three lines of code.