Simon Eskildsen, founder of Turbopuffer, reflects on his decade managing databases and infrastructure at Shopify.
Insight
Eskildsen: Model weights compress reasoning, not all world knowledge
“We can take all of the world's knowledge, all of the exabytes and exabytes of data that there is, and we can use those tokens to train a model, but we can't compress all of that into a few terabytes of weights, right? We can compress into a few terabytes of we…”
Insight
Eskildsen: Object-storage-first databases trade 200ms write latency for pure upside
“The only real downside to that is that if you go all in on object storage, every write will take a couple hundred milliseconds of latency, but from there, it's really all upside, right? You do the first query, it takes half a second”
Assertion Supported
Eskildsen: Neon retrofitted Postgres for S3, while Turbopuffer built pure object-storage consensus
“I think neon neon was first to, and they're trying to retrofit it onto Postgres. And then they built this whole architecture where you have it in memory, and then you sort of like, you know, mmap back to S-III, and I think that was very novel at the time to do…”
Disclosure
Eskildsen: Offered to return capital if Turbopuffer lacked PMF by year-end
“I don't think I've said this publicly before, but I just called Locky and was like, well, Locky, like, if this doesn't have PMF by the end of the year, like, we'll just like return all the money to you. But it's just like, I don't really, Justine and I don't w…”
Assertion Open · timeframe Mar 2027
Eskildsen: Turbopuffer outperforms Lucene on long LLM search queries
“Turbo Puffer today has a fairly start of the state of the art full text search engine.
We beat Lucene on some queries, in particular, very long queries that we've optimized for, because those are the text search queries we see today.”
Assertion Not checkable as stated
Eskildsen: Vector Search for Readwise Would Have Cost Six Times Its Total Infra Bill
“But this was a company that was spending maybe five grand a month in total on all of their infrastructure. And when I did the napkin math on running the embeddings of all the articles, putting them into a vector index, putting it in prod, it's going to be like…”