Amazon S3, every mention
65 scenes, the whole family · ← back to Amazon S3
tap a year for its mentions
every year anyone Matt Turck 14Justin Borgman 8Dave Burgess 7Zach Sherman 5Savin Goyal 5Haoyuan Li 5Chris Wiggins 4Bob Muglia 4Spencer Kimball 3Prat Moghe 3
Verbatim, from the transcripts: the passages where Amazon S3 comes up
Everything Gets Rebuilt: The New AI Agent Stack | Harrison Chase, LangChain
- ▶ 18:09 Harrison Chase So yeah, it could be anything under the hood database S three, real file system.
Rewriting Success: What InfluxDB 3.0 Teaches About Scaling—and Scrapping—Your Core Tech
- ▶ 0:27 Matt Turck On the deep tech side, we unpack how InfluxDB obliterates high cardinality bottlenecks, streams straight to S-Tree and Parquet, and why F-Dap might become the next LAMP stack.
- ▶ 16:08 Matt Turck So the idea is to live on top of S-III, is that, is that the right? 4 times in the scene
Trino, Iceberg and the Battle for the Lakehouse | Justin Borgman, CEO, Starburst
- ▶ 29:25 Justin Borgman It is connecting to your storage, uh, so it's your own S three buckets, your own, you know, RDS, your own MySQL database, uh, but the compute and the control plane is managed by us, and so we're able to offer a very seamless, easy to use,…
- ▶ 32:53 Justin Borgman So even if you're just accessing S three and you're going to be querying iceberg tables, 2 times in the scene
Understanding Data Engineering in 2025 | Ben Rogojan, Seattle Data Guy
- ▶ 39:46 Ben Rogojan (Seattle Data Guy) I mean, I think the big thing is that it sets a standard for, for how you're going to end up storing data and interacting with it, which just, you know, instead of, uh, you know, Databricks stores their data in Delta, um, Snowflake stores…
The Death of Big Data and Why It’s Time To Think Small | Jordan Tigani, CEO, MotherDuck
- ▶ 13:45 Jordan Tigani Because of separation of storage and compute, that other data just sort of sits there and is culled on, you know, on AWS S three,
Building The Database That Can Do It All | Tobie Morgan Hitchcock, CEO of SurrealDB
- ▶ 41:14 Tobie Morgan Hitchcock We then connect our compute loads to that, but the data itself is actually residing in object storage, so S-three in AWS situation, and that gives us nine nines of durability on the data that resides there. 2 times in the scene
Turbocharging Postgres for Time Series & Vectors — Timescale CTO Mike Freedman | Data Driven NYC
- ▶ 5:01 Mike Freedman Tiered storage transparently tiers huge tables on both high performance disk and bottomless S three object storage.
- ▶ 25:07 Mike Freedman There's a lot of stuff that's really, we're taking Postgres in a cloud native approach. 2 times in the scene
The $9B Startup Going After Snowflake and Databricks | Renen Hallak, CEO of VAST Data
- ▶ 20:25 Matt Turck Uh, so cloud that's going to be the AWS's, the S-III's of the world possibly.
From Xbox to Databricks: Carly Taylor’s Rebel Path in Data Science & Gaming AI
- ▶ 35:02 Matt Turck And I guess what, what are the, you know, the way you recommend people get started is that, uh, you know, you need to have like all, uh, modern data stack in place or go buy Snowflake or just put a bunch of, um, data on, on an S three.
Building LlamaIndex: Jerry Liu on Scaling Retrieval-Augmented AI
An AI Assistant to Work Faster with Notion's Head of Data, Daniel Sternberg
- ▶ 4:40 Daniel Sternberg We also have our own data lake infrastructure that we've built out ourselves on top of S three.
A Conversation with Chris Wiggins - Author of "How Data Happened"
- ▶ 20:10 Chris Wiggins Um, it was, so when I showed up at the New York Times in 2013, if you wanted to get your hands on data, you needed to write your own MapReduce jobs in Hive and hit buckets of unstructured JSON sitting in S three.
The 2023 MAD (Machine Learning, Artificial Intelligence & Data) Landscape
- ▶ 8:00 Matt Turck Why don't you just dump everything in S three, which is very cheap.
A Cloud SQL Database Built for Survival | Cockroach CEO & Co-Founder, Spencer Kimball
- ▶ 14:57 Spencer Kimball So if everyone in here is aware of Snowflake, I mean, they're, they're building on the, the cloud data storage primitives, like S three or Google cloud storage, and that's, that's a huge benefit by having that primitive, that's able,…
Separating Data Hype From Substance | Fireside Chat with Mode Co-Founder Benn Stancil
- ▶ 33:25 Benn Stancil Um, all that stuff can be rebuilt on top of, like, a data infrastructure instead of on top of just sort of AWS and S-three and EC-two and all that kind of stuff.
Fireside Chat: DeVaris Brown (Founder & CEO, Meroxa) with Matt Turck (Partner, FirstMark)
- ▶ 14:54 DeVaris Brown So, you know, if you want to have a, because you can multiplex or, you know, kind of multi, you know, branch out that same stream to multiple destinations, uh, it's really easy for us to have like an immutable S three bucket
Fireside Chat: Dev Ittycheria (President & CEO, MongoDB) with Matt Turck (Partner, FirstMark)
- ▶ 14:19 Matt Turck Search your, like, your S three folders. 2 times in the scene
Fireside Chat: Dave Burgess (Head of Data Engineering, Pinterest) w/ Matt Turck (Partner, FirstMark)
- ▶ 8:54 Dave Burgess So we put all of this data into S three 3 times in the scene
- ▶ 12:30 Dave Burgess And that's, you know, mostly from this data that comes into, that comes into S three and AWS.
- ▶ 30:56 Dave Burgess And so, and so Kafka, yeah, Kafka basically, uh, brings all the data into S three 3 times in the scene
Fireside Chat: Arjun Narayan (Founder & CEO, Materialize) with Matt Turck (Partner, FirstMark)
- ▶ 15:34 Arjun Narayan The integrations to tools like Kafka, uh, the integrations to, to, to, to pull batch data from S three.
- ▶ 22:25 Arjun Narayan Basically the ability to store large historical data sets on very cheap storage like Amazon S three, um, such that, you know, you don't have to really think about separating out your historical data set.
Fireside Chat: Savin Goyal (ML Infra team (Metaflow), Netflix) with Matt Turck (Partner, FirstMark)
- ▶ 3:23 Savin Goyal Uh, so we use SG as
- ▶ 5:51 Savin Goyal So if say you were a data scientist at Netflix, then you could in theory use a Jupyter notebook, get access to, uh, your data from ST via a bunch of different query engines, and then, uh, 2 times in the scene
- ▶ 15:46 Savin Goyal Uh, you know, let's say if you're using S three as your data store, but now if you're using Kubernetes, uh, as your compute cluster and say, if you want to use step functions or Argo,
- ▶ 20:27 Savin Goyal They care about what sort of modeling libraries, uh, they should be using, but not so much about, you know, how the data is actually being stored in S three, as long as they make efficient access, uh, to S three.
Introducing Kedro
- ▶ 11:39 Kedro Product Manager In most cases, your data catalog could look something like this, where your data is an S three or is your blog storage or Google cloud storage or do file system somewhere.
Data Observability and Pipelines: OpenLineage and Marquez
- ▶ 2:35 Julien Le Dem You may be depending on an S three bucket or a table in snowflake, but you don't know how it's produced, right?
- ▶ 18:12 Julien Le Dem Your, um, data infrastructure, and you have an injection, and then you add a storage layer for streaming and for batch processing using things like Kafka or Sree or HDFS, and then you would have stream and batch processing, and usually you…
Fireside Chat: Alok Gupta (Head of Data Science & ML, DoorDash) with Matt Turck (Partner, FirstMark)
- ▶ 21:12 Matt Turck Am I understanding correctly that instead of running Databricks on something like S-III, you are running on Snowflake. 4 times in the scene
Fireside Chat: Jeremiah Lowin (Prefect), Tristan Handy (dbt) with Matt Turck (Partner, FirstMark)
- ▶ 11:08 Tristan Handy Um, that's my, you know, I know that that's a technology centric answer to that question, but like, it's the same way that like EC two or S three has changed the way that we think about not just like building applications, but
Fireside Chat: George Fraser (Founder & CEO, Fivetran) with Matt Turck (Partner, FirstMark)
- ▶ 13:31 George Fraser Uh, so things like Dropbox or S three and, uh, we support events.
Optionality in Data Architecture // Justin Borgman, Starburst Data (FirstMark's Data Driven NYC)
- ▶ 2:55 Justin Borgman Um, this is, uh, maybe a little hard for some of you guys to see, uh, but it shows at the top, uh, a variety of popular BI tools that you can connect to using ODBC or JDBC drivers, and on the bottom, a variety of data sources, and you'll…
- ▶ 8:10 Justin Borgman You have EC two, which is your compute, and you have S three, which is your storage. 2 times in the scene
- ▶ 10:17 Justin Borgman So even though that data lives in S-III,
- ▶ 10:34 Justin Borgman So, you know, a few years ago, a lot of people were moving data from, let's say, a large database or data warehousing appliance into HDFS, the Hadoop file system, but now many people are moving that into S three.
A New Kind of Logging System // Zach Sherman & Ben Johnson, Timber (FirstMark's Data Driven NYC)
- ▶ 9:11 Zach Sherman Um, the data plane requests information from the control plane about where data should be going, and then it flows through from sources through this transformation step, uh, into the actual syncs, and so that could be S three, or Splunk,… 2 times in the scene
- ▶ 10:54 Zach Sherman And so this lets you actually sample the data and determine like, hey, maybe I only want to send, uh, 20% of the data to Splunk and a hundred percent of it to S three because I'm in financial services and I need to store that data for…
- ▶ 12:51 Zach Sherman Basically, people are sampling the request that they sent to Splunk, still using it because it's a great tool, but they're also using a different time series database, and then they're sending all of that data to a cheaper long-term store,…
- ▶ 12:51 Zach Sherman Basically, people are sampling the request that they sent to Splunk, still using it because it's a great tool, but they're also using a different time series database, and then they're sending all of that data to a cheaper long-term store,…
Make AI Less Mysterious // Serkan Piantino, Spell (FirstMark's Data Driven)
- ▶ 23:03 Serkan Piantino Um, in a normal workflow, you have a huge number of logs, and let's say they sit in S three.
Data Pipelines at Braze // Jon Hyman, Braze (FirstMark's Data Driven)
Building an Operating System for AI // Diego Oppenheimer, Algorithmia (FirstMark's Data Driven)
- ▶ 15:50 Diego Oppenheimer So if I'm actually going to be calling data in a specific cloud client, it doesn't matter if that data is in S-III.
Big Data as a Service // Prat Moghe, Cazena (FirstMark's Data Driven)
- ▶ 7:26 Prat Moghe And then, uh, what about S three versus blob store? 3 times in the scene
Rethinking Predictive Analytics // Yaniv Altshuler, Endor (FirstMark's Data Driven)
- ▶ 13:37 Yaniv Altshuler This is the raw data, also partially hashed, and the, the integration is just taking the data, the large CSV file, putting it on an AWS S-tree bucket that we created for this purpose, and telling the system which column we want to ask…
The Uber Big Data Story // Praveen Murugesan, Uber (Data Driven NYC / FirstMark)
- ▶ 3:46 Praveen Murugesan And, ah, for all the analytic log data, which basically tends to be, like, more in volume, ah, we left that data a copy in S three, ah, which was pretty smart at that point, because we really were not consuming it, but we still wanted to,…
A Virtual Distributed Storage System // Haoyuan Li, Alluxio (Hosted by FirstMark)
- ▶ 3:32 Haoyuan Li Say for example, you may have cloud storage like Amazon S three, Microsoft storage, have Google storage like Google cloud storage,
- ▶ 6:00 Haoyuan Li Like you have Amazon, SIII, Swift, OSS from Alibaba, you have HDFS, Gloucester, and this list goes on and on. 2 times in the scene
- ▶ 13:31 Haoyuan Li And there's another one, you run Impala on top of Aluxio, Aluxio on top of Amazon S III. 2 times in the scene
Combining Machine Learning With Expert Human Judgement // Eric Colson, Stitch Fix
- ▶ 15:07 Eric Colson We use the usual tools, um, S-Tree and Spark.
Ramana Rao, Livefyre // Real-Time Social Engagement (Hosted by FirstMark Capital)
- ▶ 13:23 Ramana Rao We push static things and many pre-computed things out to CDNs, ah, or stored right there on S-III.
Spencer Kimball, Cockroach Labs // A New Kind of Database (Hosted by FirstMark Capital)
- ▶ 21:30 Spencer Kimball You know, they, they built their whole own, uh, alternative to S 2 times in the scene
Joseph Essas, OpenTable // Mining Diner Talk (Hosted by FirstMark Capital)
- ▶ 2:43 Joseph Essas We'll also store it in S-III for, um, kind of more long-term processing and trying to get some insights out of it.
Bob Muglia, Snowflake // Navigating The Cloud (Hosted by FirstMark Capital)
- ▶ 6:35 Bob Muglia Our primary data store with Snowflake is S-III.
- ▶ 14:35 Bob Muglia The data is actually replicated, the core data is actually replicated using S-III,
- ▶ 22:49 Bob Muglia So what we do is the data is permanently stored in blob storage in S-three. 2 times in the scene
Chris Wiggins, NY Times // Data Science at The New York Times (Hosted by FirstMark Capital)
- ▶ 18:27 Chris Wiggins There's a lot of use of, um, both S-III and HDFS, so we do a lot of MapReduce jobs to chomp up huge JSON buckets of our own design in S-III via EC-II in order to render it down to a bite-sized data table so we can beat it down with… 3 times in the scene
Tobi Knaup, Mesosphere // Data Driven #29 // Sep 2014 (Hosted by FirstMark Capital)
- ▶ 19:20 Aaron Franco Am I writing to bare metal, or am I, like, is it automating it to S three?
Ashish Thusoo, Qubole // Data Driven #26 // April 2014 (Hosted by FirstMark Capital)
- ▶ 9:14 Ashish Thusoo Should I store my data in, um, S-three, uh, or, you know, should I, you know, use Glacier as a storage, which is much more, you know, offline storage and things like that.
- ▶ 19:15 Ashish Thusoo So just recently, for example, last, uh, whatever it was, last week or two weeks back, the costs, the storage costs in the cloud, for example, S-three and so on and so forth, were cut to one third, like a 70% drop. 2 times in the scene
Fireside chat with Dwight Merriman // Data Driven #8 // Sep 2012 (interviewed by Matt Turck)
- ▶ 48:25 Dwight Merriman Like, this is about five years ago, so like, Amazon S III, right?
Panel: Metamarkets, Kaggle and Quid // Data Driven NYC #4 // Mar 2012
- ▶ 39:01 Mike Driscoll Already sort of drinking the Kool-Aid, in a sense, ah, and by that I would define them as, ah, companies that are already on the cloud, because if you have a company that's storing its data in S-III, it's, like, gonna be a lot easier to,…