why aren't all 14 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Insight
Physical AI moats require real-world edge data you cannot crawl online
“These are not the tokens you're going to find online. Like you can't crawl Reddit and find out about what happened on a construction site, right? Nor can you do you know, test time reasoning about it. You can't just, you can simulate all kinds of environments,…”
Insight
Data startups push horizontal platforms, but enterprise buyers want single features
“The companies I see, it feels like a number of the startups I see in the space get pressure to expand to be more horizontal solutions. And then frankly, like we want them for one thing, but not for the three other things that they want to do.”
Opinion
Hanlon: AWS Redshift's core architecture has fallen behind competitors
“Redshift feels way behind, whereas a number of years ago, it was fantastic. The core architecture there has not kept up.”
Assertion Partly supported
Google, Facebook, Microsoft, and OpenAI used Reddit data for AI models
“Google, Facebook, Microsoft, and OpenAI all used Reddit's data to train their conversational AI models.”
Assertion Not checkable as stated
Hanlon: Reddit ingested 55-60B daily events managed by one person in 2019
“I came in, we were doing probably about 55 or sixty billion events a day into a data warehouse that one person was managing.”
Disclosure
Reddit's 2021 home feed will include recommendations beyond user subscriptions
“You'll see this year that the home feed and people's home feed will be breaking the subscription wall and having recommendations in there for people.”
Insight
Pesenti: Do not train production AI chatbots on Reddit data
“So I would not advise to train a production bot on Reddit.”
Disclosure
Reddit isolates malformed data events in an automated penalty box
“You can do stuff for event validation to say, ah, events that don't match certain things get dropped and sort of get put in the penalty box, which we've done and said, okay, developers, you can't actually join these to anything. You can't actually use them if …”
Prediction Not checkable as stated
Hanlon: Reddit's data team will probably reach ~120 people by end of 2021
“By the end of the year, the data organization I just described will be probably the third largest organization at Reddit with a 120 or so people.”
Assertion Not checkable as stated
Jack Hanlon: 2020 was the first year every Reddit surface implemented personalization
“Last year was the first year that every Reddit surface saw personalization.”
Disclosure
Reddit is leaning heavily into becoming a Kafka shop for streaming
“So, you know, I think we're ending up leaning heavily into being a Kafka shop, building more of the stuff into base plate, into more of these core services”
Assertion Not checkable as stated
Hanlon: Reddit grew from around 600 to 800 employees in 14 months
“When I got to Reddit overall, I'd say Reddit was maybe 600 people 14 months ago. Right now, oh, maybe 800.”
Disclosure
Reddit runs production on AWS and analytics/ML on GCP
“From the top view, I'd say we're AWS prod with originally a lot of Postgres For a variety of prod systems, we actually transit almost all that stuff over to GCP, and we use TensorFlow and BigQuery for our analytics and our model concerns”
Disclosure
Hanlon: Reddit uses a hub-and-spoke model for data science
“Data science is hub and spoke. So so effectively that is all of those folks report into the director of data science, who's, who's one of my team members and, but they actually, when we're in a, when we're in an office setting, they sit with the downstream tea…”