Ben Firshman, CEO of Replicate, addresses whether industry GPU hardware crunches are constraining inference workloads at Replicate.
Insight
Firshman: Academic ML researchers make poor software users due to six-month cycles
“And they were like, what was funny about that is they were like, not very good users. Like, they were doing great work, obviously, but the way that research worked is that they just made, like, one thing every six months, and they just fired and forgot it. Lik…”
Assertion Contradicted
Firshman: Early 2021 Discord AI bots originated Midjourney's collaborative interface
“It was the start of, it was the start of mid-journey, and, you know, it's where that kind of user interface came from. Like, what's beautiful about the user interface is, like, You could see what other people are doing, and that you could riff off other people…”
Assertion Partly supported
Firshman: Llama 2 costs $25M to train but $50 to fine-tune
“Lama II as a base model is that, like, yeah, it costs twenty five million dollars to train to start with, but then you can fine tune it for, like, 50 bucks.”
Opinion
Firshman: Building Fig and Docker Compose in Python was a mistake
“We used Python, which was a big mistake, where Python's really hard to get booting up fast, because you have to load up the whole Python runtime before it can run anything.”
Insight
Firshman: CLIs are a more natural fit for LLMs than GUIs
“It's almost more natural to a CLI than it is in a graphical user interface, because it feels like there's back and forth with the computer. Yeah. Almost funnily like a language model. So I think there's some interesting intersection of like CLIs and language m…”
Disclosure
Firshman: Llama 2 release was Replicate's biggest week of growth ever
“Llama II was, like, our biggest week of growth ever, because, like, tons of people wanted to tinker with it and run it.”