MapReduce, every mention

64 scenes across 8 shows · ← back to MapReduce

tap a year for its mentions
0013425820132014201520162017201820192020202120222023202420252026episodesmentions
04820132014201520162017201820192020202120222023202420252026episodes it came up in
00244820132014201520162017201820192020202120222023202420252026episodesmentions per episode

the MAD Podcast 62the a16z Podcast 24the Y Combinator Startup Podcast 10Acquired 4Latent Space 2TBPN 2the Official SaaStr Podcast 1All-In 1

every year every show the MAD Podcast 62 the a16z Podcast 24 the Y Combinator Startup Podcast 10 Acquired 4 Latent Space 2 TBPN 2 the Official SaaStr Podcast 1 All-In 1

Verbatim, from the transcripts: passages where MapReduce comes up on the MAD Podcast, the a16z Podcast, the Y Combinator Startup Podcast, Acquired, Latent Space

loading…

Demis Is Out as DeepMind CEO, Revolut Founder Yacht Drama, AI’s Great Reverse Bank Run | Diet TBPN Aug 5, 2026 · 1 mention

  • ▶ 2:20 John Coogan Uh, he has been, you know, very key to so many different, uh, Google projects, Integral and MapReduce, and, uh, actually being able to scale Google systems.

Jeff Dean: The 1% Rule for Building in AI · Y Combinator Jul 30, 2026 · 5 mentions

  • ▶ 0:24 Diana Hu So, um, you built MapReduce, Bigtable, TensorFlow, the TPU, Gemini, we could spend a whole hour on all the things you've done.
  • ▶ 40:31 Jeff Dean The origin of MapReduce is another good example. 3 times in the scene
  • ▶ 55:07 Diana Hu Now, one last thing, I'm pretty sure someone in this room or multiple people will eventually build something as consequential as you've done with MapReduce, TPU, distillation, et cetera, et cetera.

The Agent Cloud: Databricks’ Bet on the Future of AI — Matei Zaharia and Reynold Xin Jun 24, 2026 · 1 mention

Weekly Recap: Gwyneth Paltrow Saves Astronomer, Figma’s IPO, Top AI Researchers Ranked Aug 2, 2025 · 1 mention

Google Part I: Origins of Search. How the Best Business in Human History Happened (Audio) · Acquired Jun 30, 2025 · 4 mentions

  • ▶ 1:00:46 David Rosenthal He also implemented the first version of AdWords, built AdSense, rewrote the core search pipeline five times, co-invented and implemented Bigtable MapReduce, TensorFlow, and Gemini.
  • ▶ 1:11:11 David Rosenthal And they built stuff like GFS, the Google file system, MapReduce.
  • ▶ 1:11:16 David Rosenthal Yahoo would eventually feel like they needed to copy MapReduce to be competitive, and they would open source that as Hadoop, so people, Apache Hadoop, that is a Yahoo copy of Google's MapReduce.
  • ▶ 2:58:33 David Rosenthal Big table, MapReduce, et cetera, et cetera.

Beyond Leaderboards: LMArena’s Mission to Make AI Reliable May 29, 2025 · 1 mention

How To Build The Future: Aravind Srinivas · Y Combinator Feb 21, 2025 · 1 mention

  • ▶ 33:29 Aravind Srinivas They're not just like, oh, a page rank or like, uh, MapReduce or, uh, you know, all these advances that they made in like visual, like deep learning and, and, and like BERT, Transformers.

Why is everyone cloning Deep Research? Feb 18, 2025 · 1 mention

  • ▶ 47:59 Shawn Wang You know, thing or whatever in the old Google days might be like MapReduce or, you know, whatever, but like it's a, it's a different scale and nature of work.

Dataiku's Secret to Scaling AI in Global Enterprises | Florian Douetteau, CEO, Dataiku Dec 12, 2024 · 1 mention

  • ▶ 6:05 Florian Douetteau You had, uh, Google pushing new technologies, you had, like, those, uh, those days of MapReduce, those days of deep learning starting to work,

The Death of Big Data and Why It’s Time To Think Small | Jordan Tigani, CEO, MotherDuck Oct 24, 2024 · 2 mentions

  • ▶ 3:28 Jordan Tigani You know, after Google came out with, you know, MapReduce and, and GFS and Bigtable, kind of everybody's... 2 times in the scene

E156: Ivy League antisemitism, macro, SaaS recovery, Gemini, Figma deal delay + big Friedberg update Dec 8, 2023 · 1 mention

A Conversation with Chris Wiggins - Author of "How Data Happened" May 31, 2023 · 1 mention

  • ▶ 20:10 Chris Wiggins Um, it was, so when I showed up at the New York Times in 2013, if you wanted to get your hands on data, you needed to write your own MapReduce jobs in Hive and hit buckets of unstructured JSON sitting in S three.

Not Just Another Cloud Database | SurrealDB Co-Founders Jaime & Tobie Morgan Hitchcock Mar 28, 2023 · 1 mention

  • ▶ 8:27 Jaime & Tobie Morgan Hitchcock We wanted to have real-time data that was constantly changing, but at the same time, without having to run the MapReduce or large-scale analytics processes and workloads on that data, we wanted to pull out that data in real-time for…

Data Visibility & Control | BigID Co-Founder & CEO Dimitri Sirota Jan 30, 2023 · 1 mention

  • ▶ 8:57 Dimitri Sirota Now, they may differ, so for instance, for Hadoop, we could do, like, MapReduce, uh, we could do, um, direct, uh, connectivity to, um, uh, uh, Hive or Spark we could use, but they all use native protocols to scan the underlying system.

Things That Don't Scale, The Software Edition – Dalton Caldwell and Michael Seibel · Y Combinator Feb 16, 2022 · 4 mentions

  • ▶ 22:30 Dalton Caldwell Um, and this was the genesis for them to create MapReduce, which they wrote a paper about. 2 times in the scene
  • ▶ 23:21 Dalton Caldwell Again, in terms of do things that don't scale, did they build MapReduce before they had any users? 2 times in the scene

$100 Million ARR Pivot: From Platform Product to Vertical Apps With Treasure Data CEO Kazuki Ohta Dec 14, 2021 · 1 mention

  • ▶ 8:28 Kazuki Ohta And in terms of data warehouse, we built a really scalable columnar storage with the Hadoop MapReduce data processing job, really storing three hundred and sixty billion records, executing almost two million jobs in total.

Fireside Chat: Ali Ghodsi (Founder & CEO, Databricks) with Matt Turck (Partner, FirstMark) May 24, 2021 · 1 mention

Dynamic Range Sharding with Spanner // Daniel Chia, Google Spanner (FirstMark's Data Driven NYC) Feb 25, 2019 · 1 mention

  • ▶ 13:53 Daniel Chia And so when you spin up a MapReduce job with 50,000 workers and all try and write to the same little data range that previously had no data, that's not going to work very well because we haven't had time to split it yet.

a16z Podcast | A New Lab Rises Jan 2, 2019 · 2 mentions

  • ▶ 11:09 Ion Stoica If you think about the MapReduce and Google File System, which really started as a big data movement. 2 times in the scene

a16z Podcast | The Strategies and Tactics of Big Jan 2, 2019 · 1 mention

  • ▶ 17:09 Steven Sinofsky If Google wants to replace MapReduce with some new thing, or if they are going to roll out, like, a new way of doing machine learning internally, they'll work on that for the same amount of time that Apple works on a phone.

a16z Podcast | The Storage Renaissance Jan 2, 2019 · 1 mention

  • ▶ 13:28 Mike Matchett When you had Hadoop and MapReduce, they could partition certain categories of problems and run them in parallel, but not really a lot of machine learning algorithms.

a16z Podcast | A Conversation With the Inventor of Spark Jan 2, 2019 · 7 mentions

  • ▶ 0:57 unnamed speaker Before Spark, the most widely used system was probably MapReduce, which was, uh, uh, invented at Google and popularized through the, the open source Hadoop project. 7 times in the scene

a16z Podcast | The Cool Stuff Only Happens at Scale Jan 2, 2019 · 3 mentions

  • ▶ 1:41 Vijay Pande But you can't do MapReduce with everything. 2 times in the scene
  • ▶ 6:21 Vijay Pande I think I've gotten people thinking about this, and MapReduce and things like that, those abstractions have played a huge role.

a16z Podcast | Making Sense of Big Data, Machine Learning, and Deep Learning Jan 2, 2019 · 8 mentions

  • ▶ 11:41 Christopher Nguyen Uh, so big compute is the first example of the compute you can think of is MapReduce. 8 times in the scene

a16z Podcast | Why the Datacenter Needs an Operating System Jan 2, 2019 · 1 mention

  • ▶ 0:23 Steven Sinofsky There are things like Kafka and Spark and MapReduce and Cassandra, and it's super, super

Fireside Chat: Mike Tuchen, CEO of Talend (TLND) (FirstMark's Data Driven NYC) Oct 17, 2018 · 1 mention

  • ▶ 15:38 Mike Tuchen So we built a, initially in our, in our initial technology, we built a, um, SQL, or excuse me, we did build a SQL one, but we started out with a Java compiler, um, and then we, um, did a SQL one, and then when Hadoop came along, um, we did…

3 Heretical Ideas on the Future of Data // Ajay Kulkarni, TimescaleDB (FirstMark's Data Driven NYC) Sep 17, 2018 · 1 mention

  • ▶ 2:50 Ajay Kulkarni The relational databases they had in the past just couldn't scale for these workloads, so the new companies started developing the first big data non-relational systems, and this includes Google, who published the MapReduce paper and the…

Fireside Chat with Bob Muglia, CEO at Snowflake (FirstMark's Data Driven) Apr 9, 2018 · 1 mention

  • ▶ 17:46 Bob Muglia And then the, the, the, the typical approaches that people have traditionally used with Hadoop, variations of MapReduce, um, are now being seen, I think, as, as, now that there are alternatives, such as Snowflake that are available, that…

Surveillance Platform for Banks // Mayur Thakur, Goldman Sachs (FirstMark's Data Driven) Dec 19, 2017 · 1 mention

A Kafka-Powered Real-Time Streaming Platform // Neha Narkhede, Confluent [FirstMark's Data Driven] Jun 16, 2016 · 3 mentions

  • ▶ 23:52 Neha Narkhede Uh, a lot of these other stream processing systems, I didn't get a, uh, chance to get into the details. 3 times in the scene

Why Marketing is All About Data // Nitay Joffe, ActionIQ [FirstMark's Data Driven] Jun 16, 2016 · 1 mention

A Fireside Chat with MapR CTO M.C. Srivas (Data Driven NYC / FirstMark) Dec 17, 2015 · 1 mention

  • ▶ 20:58 M.C. Srivas We introduced JSON to Hadoop and Spark and in MapReduce and in Hive and everywhere.

The Acceleration of Innovation in Big Data w/ Stefan Groschupf, Datameer Dec 17, 2015 · 4 mentions

  • ▶ 1:16 Stefan Groschupf That's why we looked into the Google papers and implemented MapReduce and the HDFS, et cetera, et cetera.
  • ▶ 4:38 Stefan Groschupf and then we built a MapReduce thing, and then, um, Peter Thiel, not sure, who knows Peter Thiel? 2 times in the scene
  • ▶ 12:07 Stefan Groschupf Replacing MapReduce with TES or with Spark is really straightforward, um, especially if you use BI on top of Hadoop, um, because it basically automatically does it for you, but obviously you can't just rip out hardware and replace it,…

Oren Falkowitz, Area 1 Security // The Future of Cybersecurity (Hosted by FirstMark Capital) Oct 21, 2015 · 1 mention

  • ▶ 4:19 Oren Falkowitz Maybe, you know, eight or nine years since Google published the Bigtable paper and the MapReduce paper and some of those early big data papers which have spawned, you know, this community and all the great technological revolution, now…

Ion Stoica, Databricks // Creating Apache Spark // Data Driven NYC (FirstMark Capital) Apr 2, 2015 · 8 mentions

  • ▶ 6:25 Matt Turck Uh, can you, uh, maybe help us understand, um, so this HDFS, which is the storage, uh, and MapReduce, which is, uh, the processing, 4 times in the scene
  • ▶ 7:24 Matt Turck Yes, so perhaps more specifically, can you go through the, the, the key advantages of Spark over MapReduce? 4 times in the scene

Paul Dix, InfluxDB // Open-Source Time Series Database // Data Driven NYC (FirstMark Capital) Apr 2, 2015 · 2 mentions

  • ▶ 5:16 Paul Dix This is the key thing that we learned from Hadoop and Google's MapReduce framework, which is
  • ▶ 19:19 Paul Dix So under the covers, we basically have our own MapReduce implementation that's,

Chris Wiggins, NY Times // Data Science at The New York Times (Hosted by FirstMark Capital) Jan 16, 2015 · 2 mentions

Mike Olson, Cloudera // The Cloudera Story (Hosted by FirstMark Capital) Dec 18, 2014 · 9 mentions

  • ▶ 4:55 Mike Olson Google published the Google File System and MapReduce papers, and like the entire database industry, I thought that this thing was a joke, man.
  • ▶ 6:52 Matt Turck This Spark, this data flow, and, you know, maybe MapReduce is not what people need, and all of this is evolving.
  • ▶ 9:33 Mike Olson These days when people talk about Hadoop, what they mean is HDFS and MapReduce, yeah, yarn for resource management. 5 times in the scene
  • ▶ 14:39 Mike Olson The Hadoop that was then available, which was HDFS and MapReduce, was powerful, transformative.
  • ▶ 18:18 Mike Olson The inventor of Hadoop, Google Ventures, the inventor of MapReduce and GFS took a stake.

Michael Rubenstein and Catherine Williams, App Nexus // Data Driven #31 // Nov 2014 Nov 20, 2014 · 1 mention

Mike Abbott, KPCB // Data Driven #30 // Oct 2014 (Hosted by FirstMark Capital) Oct 16, 2014 · 1 mention

  • ▶ 12:56 Mike Abbott A third of the people that I recruited from Google was how could we, you know, let non-engineers run MapReduce jobs, ask more questions of the data, because if we let more non-engineers ask questions, there is a higher probability that we…

John Rauser, Pinterest // Big Data at Pinterest // Data Driven NYC (Hosted by FirstMark Capital) Oct 16, 2014 · 4 mentions

  • ▶ 1:16 John Rauser I'd walk them through an example of MapReduce at a level of simplicity that this audience would probably find comical. 4 times in the scene

Tobi Knaup, Mesosphere // Data Driven #29 // Sep 2014 (Hosted by FirstMark Capital) Sep 22, 2014 · 1 mention

  • ▶ 3:33 Tobi Knaup And, um, so MapReduce, for example, was, uh, started as a, um, you know, Google wrote a paper about it.

Tasso Argyros, Aster Data // Data Driven NYC 24 // February 2014 (Hosted by FirstMark Capital) Mar 3, 2014 · 7 mentions

  • ▶ 2:47 Tasso Argyros A declarative language of SQL, right, and the procedural language of MapReduce.
  • ▶ 14:39 Tasso Argyros In 2007, we had the brilliant idea of coming out with a SQL MapReduce Azure positioning, and of course, nobody knew what MapReduce back then, right? 6 times in the scene

Scott Sorensen, CTO of Ancestry.com // Data Driven 19 // October 2013 (Hosted by FirstMark Capital) Jan 16, 2014 · 1 mention

Mike Dauber, Battery Ventures // Data Driven NYC 19 // October 2013 (interviewed by Matt Turck) Dec 5, 2013 · 1 mention

  • ▶ 2:52 Mike Dauber And then came 2003, 2004, and Google came out with their, you know, their big table MapReduce papers.

Panel: Continuuity, Sailthru and Visual Revenue // Data Driven NYC #7 // June 2012 Dec 5, 2013 · 1 mention

Panel: Metamarkets, Kaggle and Quid // Data Driven NYC #4 // Mar 2012 Dec 5, 2013 · 1 mention

  • ▶ 49:39 unnamed speaker That it will scale, that it will map reduce, that whatever these words are.

Kirill Sheynkman, RTP Ventures // Data Driven #3 // Feb 2012 (Interviewed by Matt Turck) Dec 5, 2013 · 1 mention

  • ▶ 4:56 Kirill Sheynkman It's, it's, it's pretty, ah, hairy stuff, but it's, it's a system that has APIs in Scala and in Java, and they're working on JavaScript and, and Python, where you can take a function and basically the MapReduce becomes three lines of code.
Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.