why aren't all 24 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Opinion
Goyal: Open-source inference providers are far less reliable than OpenAI
“They are nowhere near as reliable as, I mean, every single time I use any of those products and run a benchmark, I find a bug, text the CEO, and they fix something. It's nowhere near where OpenAI is.”
Opinion
Chintala: Inference Moats from Fast CUDA Kernels Last Only Months
“I think, like, Together and Fireworks and all these people are trying to build some faster CUDA kernels and faster, like, you know, hardware kernels in general. But those modes only last for a month or two. Like, these ideas quickly propagate.”
Opinion
Morcos: AI training is commoditized while data curation remains hard
“Mosaic was the first one to really recognize that there was a huge opportunity in making this easy. And now this has largely been commoditized by things like SageMaker and Together and lots of different folks that help you on the training side. But on the data…”
Opinion
Prakash: Transformer architectures will not reach 5,000 tokens per second inference
“We are running into you know, the limits of how fast you can make transformers and you know, we want inference at 5000 tokens per second. And I don't think we will get there with transformers”
Opinion
Prakash: Foundation models lack data moats and are constrained by capital
“There aren't really big data moats around foundation models. They are built from a subset of the web. What is difficult is the cost of capital to build these, and what, one of the ways in which you can reduce this cost is by making more efficient systems.”
Assertion Supported
Prakash: SemiAnalysis inference report erred by assuming unpriced input tokens
“I think there were some errors in that analysis. In particular we were trying to decode it, and one of the things we noticed is that it assumed that input tokens weren't being priced. So I think that may have been an error in the model.”
Assertion Not checkable as stated
Prakash: AnyScale LLM benchmarks conflicted with internal provider testing
“Everyone sort of had a reaction to it because it just didn't match their benchmarks that we've all run internally against different services.”
Opinion
Zhang: The world is not running out of AI training data
“I don't think we are running out of data on earth. Right, so think about it globally... But I do think there are many organizations in the world have enough data to actually train, like, very, very good models, right? So, I mean, they are not public available,…”
Assertion Not checkable as stated
Swyx: Most AI researchers are scaling Transformers, not alternative architectures
“I think most people that I talk to are not seriously pursuing alternative architectures. There are some notable exceptions, primarily together with the Mombard architecture recursal with RWKV. There's like the XLSTM that was created by a separate Hogwriter. An…”
Disclosure
Prakash: Top 5 models on Together AI inference are fine-tuned open models
“I would say right now the top five models on our inference stack are probably all fine-tuned versions of open models.”
Disclosure
Prakash: Together AI keeps its inference stack proprietary for competitive advantage
“I think on the inference stack, there are open source inference stacks which are pretty good, and it gives us, you know, definitely today it gives us a competitive advantage to have the best one, and So we're not sort of rushing out to release everything about…”
Insight
Zhang: Good benchmarks must anticipate how developers over-optimize toward metrics
“A good benchmark should think about how it's going to incentivize the field to actually move forward, right? So the benchmark will become kind of standard. How are people going to over optimize to the benchmark because people are going to do that? And when peo…”
Insight
Zhang: Combining fine-tuning and RAG provides superior performance boosts
“Combining all those techniques all together, right? So we'll give you essentially another boost, right? So that kind of one thing that we learn on the technical side.”
Disclosure
Prakash: Together AI operates a fleet of 7,000 to 8,000 GPUs
“We have close to seven to 8000 GPUs today. It's growing monthly.”
Assertion Not checkable as stated
Prakash: No provider satisfies machine-to-machine AI demand for extreme token speeds
“There are applications that are, you know, consuming the tokens that are produced from unmodel, so they're not necessarily being read or heard by humans. So that's a place where we see that level of requirement today that really nobody can quite satisfy.”
Insight
Zhang: Pre-training data evolved from a model byproduct into a standalone asset
“So, so I think one fundamental thing that changed in the last year, essentially, in the beginning when people think about data, is, is always like a byproduct of the model, right? You release the model, you also release the data, right? The data side is there …”
Assertion Supported
Prakash: Together AI builds cloud footprint avoiding major hyperscalers
“The way we structure our cloud is by combining data centers around the world instead of you know we are today not located in hyperscalers. We have built a footprint of you know, AI supercomputers in this sort of a disaggregated, decentralized manner.”
Prediction Not checkable as stated
Prakash predicts data marketplaces for AI training will increasingly emerge
“I think you need to have some kind of marketplace for figuring out how to get this you know, data into models and have, I think you'll increasingly see more of that”
Assertion Not checkable as stated
Prakash: 40% to 45% of Together AI's team is dedicated to research
“It's like 40, 45% I was counting this morning.”
Insight
Zhang: Next 10x AI inference gain requires multi-layer co-optimization
“If you only push on one direction, you are going to reach diminution return really, really quickly. Yeah, there's only that much you can do on the system side, only that much you can do on the algorithm side. And since the only big thing that's going to happen…”
Prediction Not checkable as stated
Prakash: Near-instant model inference will spawn new application paradigms
“Once you can get sort of an immediate answer from a model it starts working in a different way and you know, new types of applications will be created.”
Assertion Supported
Zhang: RedPajama-V2 features 40 pre-computed quality signals for custom filtering
“So that's why in REST-PYRON V-II, we kind of overlay the data set, it's like, 40 different pre-computed quality signal, right? If you want to reproduce your best effort, like, C-Four filter, it's kind of like, 20 lines of code.”
Disclosure
Tri Dao joins Together AI as Chief Scientist
“Yeah, yeah, so I just joined this week actually, and it's been really exciting.”
Assertion Not checkable as stated
Prakash: Together AI has 38 employees as of February 2024
“We are 38 people on the team and we are hiding across all the areas you know.”