Amazon S3, every mention
15 scenes (2021), the whole family · ← back to Amazon S3
tap a year for its mentions
every year 2021 anyone Matt Turck 14Justin Borgman 8Dave Burgess 7Zach Sherman 5Savin Goyal 5Haoyuan Li 5Chris Wiggins 4Bob Muglia 4Spencer Kimball 3Prat Moghe 3
Verbatim, from the transcripts: the passages where Amazon S3 comes up
Fireside Chat: DeVaris Brown (Founder & CEO, Meroxa) with Matt Turck (Partner, FirstMark)
- ▶ 14:54 DeVaris Brown So, you know, if you want to have a, because you can multiplex or, you know, kind of multi, you know, branch out that same stream to multiple destinations, uh, it's really easy for us to have like an immutable S three bucket
Fireside Chat: Dev Ittycheria (President & CEO, MongoDB) with Matt Turck (Partner, FirstMark)
- ▶ 14:19 Matt Turck Search your, like, your S three folders. 2 times in the scene
Fireside Chat: Dave Burgess (Head of Data Engineering, Pinterest) w/ Matt Turck (Partner, FirstMark)
- ▶ 8:54 Dave Burgess So we put all of this data into S three 3 times in the scene
- ▶ 12:30 Dave Burgess And that's, you know, mostly from this data that comes into, that comes into S three and AWS.
- ▶ 30:56 Dave Burgess And so, and so Kafka, yeah, Kafka basically, uh, brings all the data into S three 3 times in the scene
Fireside Chat: Arjun Narayan (Founder & CEO, Materialize) with Matt Turck (Partner, FirstMark)
- ▶ 15:34 Arjun Narayan The integrations to tools like Kafka, uh, the integrations to, to, to, to pull batch data from S three.
- ▶ 22:25 Arjun Narayan Basically the ability to store large historical data sets on very cheap storage like Amazon S three, um, such that, you know, you don't have to really think about separating out your historical data set.
Fireside Chat: Savin Goyal (ML Infra team (Metaflow), Netflix) with Matt Turck (Partner, FirstMark)
- ▶ 3:23 Savin Goyal Uh, so we use SG as
- ▶ 5:51 Savin Goyal So if say you were a data scientist at Netflix, then you could in theory use a Jupyter notebook, get access to, uh, your data from ST via a bunch of different query engines, and then, uh, 2 times in the scene
- ▶ 15:46 Savin Goyal Uh, you know, let's say if you're using S three as your data store, but now if you're using Kubernetes, uh, as your compute cluster and say, if you want to use step functions or Argo,
- ▶ 20:27 Savin Goyal They care about what sort of modeling libraries, uh, they should be using, but not so much about, you know, how the data is actually being stored in S three, as long as they make efficient access, uh, to S three.
Introducing Kedro
- ▶ 11:39 Kedro Product Manager In most cases, your data catalog could look something like this, where your data is an S three or is your blog storage or Google cloud storage or do file system somewhere.
Data Observability and Pipelines: OpenLineage and Marquez
- ▶ 2:35 Julien Le Dem You may be depending on an S three bucket or a table in snowflake, but you don't know how it's produced, right?
- ▶ 18:12 Julien Le Dem Your, um, data infrastructure, and you have an injection, and then you add a storage layer for streaming and for batch processing using things like Kafka or Sree or HDFS, and then you would have stream and batch processing, and usually you…
Fireside Chat: Alok Gupta (Head of Data Science & ML, DoorDash) with Matt Turck (Partner, FirstMark)
- ▶ 21:12 Matt Turck Am I understanding correctly that instead of running Databricks on something like S-III, you are running on Snowflake. 4 times in the scene