Feb 8, 2024 · 1h 15m · latent-space

Building an open AI company - with Ce and Vipul of Together AI

Vipul Ved Prakash · 26m spoken Ce Zhang · 24m spoken Shawn Wang · 8m spoken Alessio Fanelli · 6m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Together AI co-founders Vipul Ved Prakash and Ce Zhang join the Latent Space podcast to discuss their open-source platform, disaggregated cloud infrastructure, and multi-dimensional inference optimizations. They detail dataset initiatives like RedPajama, hybrid architectures such as StripedHyena, and the systems engineering required to deliver high-throughput, serverless AI for developers.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 23.5% of the talking time here. How this is scored →

The hosts as informed peer 4.6 Guest teaching 5.2 Guest disagreement 1.3 The hosts pushing back 1.7
05100:0020:0040:001:00:002:28–5:37 · The hosts as informed peer 4/10 Vipul's Apple Background and Lessons on AI Systems Alessio asks an insightful question contrasting Apple's closed, polished ecosystem with Together's open philosophy. Vipul shares his deep background in spam filtering, early open domain Q&A at Apple, and his perspective on scaling laws.5:38–10:17 · The hosts as informed peer 4/10 Founding Together AI and Research-Led Systems Architecture Hosts ask about the intersection of academic research and startup founding. Ce and Vipul explain why data movement across stacks and gradient compression for decentralized computing drove the company's genesis.10:17–15:27 · The hosts as informed peer 5/10 RedPajama Dataset Evolution: V1 to V2 Swyx and Alessio dig into RedPajama V1 and V2, observing the shift toward modular quality filtering signals. Ce explains the philosophy of turning dataset artifacts into tunable, multi-signal platforms rather than static snapshots.15:27–20:30 · The hosts as informed peer 6/10 Custom Model Training and Importance Resampling (DSIR) Swyx pushes back by arguing that targeted importance resampling (DSIR) violates the pursuit of general intelligence. Ce counters by reframing the trade-off space around deployment costs, meta-learning, and domain-targeted models.20:30–24:57 · The hosts as informed peer 5/10 Exploring Global Data Limits and Marketplace Dynamics Swyx brings up debates around YouTube token quality from Whisper and proprietary data walled gardens. Vipul and Ce acknowledge data fragmentation and propose fair marketplace dynamics for creators.24:58–32:10 · The hosts as informed peer 5/10 GPU Fleet Economics and Disaggregated Supercomputing Hosts press for specific numbers on GPU cluster size and allocation across pre-training versus inference. Vipul discloses their 7,000 to 8,000 GPU fleet scale and calculates overall industry Capex economics.32:11–39:27 · The hosts as informed peer 6/10 Inference Stack Optimization and Cloud Architecture Alessio references SemiAnalysis's critique of Together's pricing and speculative decoding setup. Vipul systematically clarifies where Dylan Patel's model made incorrect assumptions regarding input token pricing and hardware configurations.39:30–43:04 · The hosts as informed peer 4/10 Multi-Dimensional Inference Co-Optimization Ce and Vipul explain the compounding benefits of co-optimizing across algorithms, model architectures, and custom kernels rather than isolating single improvements.43:04–49:17 · The hosts as informed peer 5/10 Industry Benchmarking Challenges and Standardization Alessio asks about the AnyScale benchmark controversy. Ce and Vipul emphasize the systemic dangers of benchmarks creating bad optimization incentives and advocate for neutral third-party measurement.49:18–52:10 · The hosts as informed peer 4/10 Fine-Tuning Spectrum and Enterprise Customer Engagements Swyx asks about customer engagement tiers for fine-tuning. Vipul outlines the spectrum from full consultative pre-training to serverless fine-tuning workflows.52:10–54:51 · The hosts as informed peer 5/10 Advancements and Future Potential in Embeddings Alessio asks whether embedding models have reached a plateau. Ce explains why embeddings remain in their infancy, particularly regarding fine-grained semantics, negation, and data flywheel loops.54:51–1:03:41 · The hosts as informed peer 5/10 State Space Models, Hybrids, and StripedHyena Swyx questions why researchers should care about subquadratic state space models beyond simple context length. Ce educates the hosts on memory footprint, execution patterns, and hybrid layer grafting in StripedHyena.1:03:41–1:06:17 · The hosts as informed peer 4/10 The Case for 5,000 Tokens per Second Inference Swyx asks why anyone needs 5,000 tokens per second when humans read much slower. Vipul clarifies that machine-to-machine consumption, hardware card throughput, and interactive UX fundamentally require extreme generation speeds.1:06:19–1:11:40 · The hosts as informed peer 4/10 Delivering a Pure Serverless AI Developer Experience Swyx shares his past experience running out of credits on Together. Vipul and Ce explain their pivot to a fully serverless, friction-free developer experience and outline hiring priorities.1:11:40–1:14:49 · The hosts as informed peer 3/10 Lightning Round: Unsolved Questions and Positive AI Frameworks Ce discusses edge satellite communications as an alternative research passion, while Vipul calls for replacing science-fiction doomerism with constructive frameworks for advanced intelligence.2:28–5:37 · Guest teaching 5/10 Vipul's Apple Background and Lessons on AI Systems Alessio asks an insightful question contrasting Apple's closed, polished ecosystem with Together's open philosophy. Vipul shares his deep background in spam filtering, early open domain Q&A at Apple, and his perspective on scaling laws.5:38–10:17 · Guest teaching 5/10 Founding Together AI and Research-Led Systems Architecture Hosts ask about the intersection of academic research and startup founding. Ce and Vipul explain why data movement across stacks and gradient compression for decentralized computing drove the company's genesis.10:17–15:27 · Guest teaching 6/10 RedPajama Dataset Evolution: V1 to V2 Swyx and Alessio dig into RedPajama V1 and V2, observing the shift toward modular quality filtering signals. Ce explains the philosophy of turning dataset artifacts into tunable, multi-signal platforms rather than static snapshots.15:27–20:30 · Guest teaching 6/10 Custom Model Training and Importance Resampling (DSIR) Swyx pushes back by arguing that targeted importance resampling (DSIR) violates the pursuit of general intelligence. Ce counters by reframing the trade-off space around deployment costs, meta-learning, and domain-targeted models.20:30–24:57 · Guest teaching 4/10 Exploring Global Data Limits and Marketplace Dynamics Swyx brings up debates around YouTube token quality from Whisper and proprietary data walled gardens. Vipul and Ce acknowledge data fragmentation and propose fair marketplace dynamics for creators.24:58–32:10 · Guest teaching 6/10 GPU Fleet Economics and Disaggregated Supercomputing Hosts press for specific numbers on GPU cluster size and allocation across pre-training versus inference. Vipul discloses their 7,000 to 8,000 GPU fleet scale and calculates overall industry Capex economics.32:11–39:27 · Guest teaching 6/10 Inference Stack Optimization and Cloud Architecture Alessio references SemiAnalysis's critique of Together's pricing and speculative decoding setup. Vipul systematically clarifies where Dylan Patel's model made incorrect assumptions regarding input token pricing and hardware configurations.39:30–43:04 · Guest teaching 5/10 Multi-Dimensional Inference Co-Optimization Ce and Vipul explain the compounding benefits of co-optimizing across algorithms, model architectures, and custom kernels rather than isolating single improvements.43:04–49:17 · Guest teaching 5/10 Industry Benchmarking Challenges and Standardization Alessio asks about the AnyScale benchmark controversy. Ce and Vipul emphasize the systemic dangers of benchmarks creating bad optimization incentives and advocate for neutral third-party measurement.49:18–52:10 · Guest teaching 5/10 Fine-Tuning Spectrum and Enterprise Customer Engagements Swyx asks about customer engagement tiers for fine-tuning. Vipul outlines the spectrum from full consultative pre-training to serverless fine-tuning workflows.52:10–54:51 · Guest teaching 5/10 Advancements and Future Potential in Embeddings Alessio asks whether embedding models have reached a plateau. Ce explains why embeddings remain in their infancy, particularly regarding fine-grained semantics, negation, and data flywheel loops.54:51–1:03:41 · Guest teaching 7/10 State Space Models, Hybrids, and StripedHyena Swyx questions why researchers should care about subquadratic state space models beyond simple context length. Ce educates the hosts on memory footprint, execution patterns, and hybrid layer grafting in StripedHyena.1:03:41–1:06:17 · Guest teaching 5/10 The Case for 5,000 Tokens per Second Inference Swyx asks why anyone needs 5,000 tokens per second when humans read much slower. Vipul clarifies that machine-to-machine consumption, hardware card throughput, and interactive UX fundamentally require extreme generation speeds.1:06:19–1:11:40 · Guest teaching 4/10 Delivering a Pure Serverless AI Developer Experience Swyx shares his past experience running out of credits on Together. Vipul and Ce explain their pivot to a fully serverless, friction-free developer experience and outline hiring priorities.1:11:40–1:14:49 · Guest teaching 4/10 Lightning Round: Unsolved Questions and Positive AI Frameworks Ce discusses edge satellite communications as an alternative research passion, while Vipul calls for replacing science-fiction doomerism with constructive frameworks for advanced intelligence.2:28–5:37 · Guest disagreement 1/10 Vipul's Apple Background and Lessons on AI Systems Alessio asks an insightful question contrasting Apple's closed, polished ecosystem with Together's open philosophy. Vipul shares his deep background in spam filtering, early open domain Q&A at Apple, and his perspective on scaling laws.5:38–10:17 · Guest disagreement 1/10 Founding Together AI and Research-Led Systems Architecture Hosts ask about the intersection of academic research and startup founding. Ce and Vipul explain why data movement across stacks and gradient compression for decentralized computing drove the company's genesis.10:17–15:27 · Guest disagreement 1/10 RedPajama Dataset Evolution: V1 to V2 Swyx and Alessio dig into RedPajama V1 and V2, observing the shift toward modular quality filtering signals. Ce explains the philosophy of turning dataset artifacts into tunable, multi-signal platforms rather than static snapshots.15:27–20:30 · Guest disagreement 3/10 Custom Model Training and Importance Resampling (DSIR) Swyx pushes back by arguing that targeted importance resampling (DSIR) violates the pursuit of general intelligence. Ce counters by reframing the trade-off space around deployment costs, meta-learning, and domain-targeted models.20:30–24:57 · Guest disagreement 1/10 Exploring Global Data Limits and Marketplace Dynamics Swyx brings up debates around YouTube token quality from Whisper and proprietary data walled gardens. Vipul and Ce acknowledge data fragmentation and propose fair marketplace dynamics for creators.24:58–32:10 · Guest disagreement 1/10 GPU Fleet Economics and Disaggregated Supercomputing Hosts press for specific numbers on GPU cluster size and allocation across pre-training versus inference. Vipul discloses their 7,000 to 8,000 GPU fleet scale and calculates overall industry Capex economics.32:11–39:27 · Guest disagreement 2/10 Inference Stack Optimization and Cloud Architecture Alessio references SemiAnalysis's critique of Together's pricing and speculative decoding setup. Vipul systematically clarifies where Dylan Patel's model made incorrect assumptions regarding input token pricing and hardware configurations.39:30–43:04 · Guest disagreement 1/10 Multi-Dimensional Inference Co-Optimization Ce and Vipul explain the compounding benefits of co-optimizing across algorithms, model architectures, and custom kernels rather than isolating single improvements.43:04–49:17 · Guest disagreement 2/10 Industry Benchmarking Challenges and Standardization Alessio asks about the AnyScale benchmark controversy. Ce and Vipul emphasize the systemic dangers of benchmarks creating bad optimization incentives and advocate for neutral third-party measurement.49:18–52:10 · Guest disagreement 1/10 Fine-Tuning Spectrum and Enterprise Customer Engagements Swyx asks about customer engagement tiers for fine-tuning. Vipul outlines the spectrum from full consultative pre-training to serverless fine-tuning workflows.52:10–54:51 · Guest disagreement 1/10 Advancements and Future Potential in Embeddings Alessio asks whether embedding models have reached a plateau. Ce explains why embeddings remain in their infancy, particularly regarding fine-grained semantics, negation, and data flywheel loops.54:51–1:03:41 · Guest disagreement 1/10 State Space Models, Hybrids, and StripedHyena Swyx questions why researchers should care about subquadratic state space models beyond simple context length. Ce educates the hosts on memory footprint, execution patterns, and hybrid layer grafting in StripedHyena.1:03:41–1:06:17 · Guest disagreement 2/10 The Case for 5,000 Tokens per Second Inference Swyx asks why anyone needs 5,000 tokens per second when humans read much slower. Vipul clarifies that machine-to-machine consumption, hardware card throughput, and interactive UX fundamentally require extreme generation speeds.1:06:19–1:11:40 · Guest disagreement 1/10 Delivering a Pure Serverless AI Developer Experience Swyx shares his past experience running out of credits on Together. Vipul and Ce explain their pivot to a fully serverless, friction-free developer experience and outline hiring priorities.1:11:40–1:14:49 · Guest disagreement 1/10 Lightning Round: Unsolved Questions and Positive AI Frameworks Ce discusses edge satellite communications as an alternative research passion, while Vipul calls for replacing science-fiction doomerism with constructive frameworks for advanced intelligence.2:28–5:37 · The hosts pushing back 1/10 Vipul's Apple Background and Lessons on AI Systems Alessio asks an insightful question contrasting Apple's closed, polished ecosystem with Together's open philosophy. Vipul shares his deep background in spam filtering, early open domain Q&A at Apple, and his perspective on scaling laws.5:38–10:17 · The hosts pushing back 1/10 Founding Together AI and Research-Led Systems Architecture Hosts ask about the intersection of academic research and startup founding. Ce and Vipul explain why data movement across stacks and gradient compression for decentralized computing drove the company's genesis.10:17–15:27 · The hosts pushing back 2/10 RedPajama Dataset Evolution: V1 to V2 Swyx and Alessio dig into RedPajama V1 and V2, observing the shift toward modular quality filtering signals. Ce explains the philosophy of turning dataset artifacts into tunable, multi-signal platforms rather than static snapshots.15:27–20:30 · The hosts pushing back 4/10 Custom Model Training and Importance Resampling (DSIR) Swyx pushes back by arguing that targeted importance resampling (DSIR) violates the pursuit of general intelligence. Ce counters by reframing the trade-off space around deployment costs, meta-learning, and domain-targeted models.20:30–24:57 · The hosts pushing back 2/10 Exploring Global Data Limits and Marketplace Dynamics Swyx brings up debates around YouTube token quality from Whisper and proprietary data walled gardens. Vipul and Ce acknowledge data fragmentation and propose fair marketplace dynamics for creators.24:58–32:10 · The hosts pushing back 2/10 GPU Fleet Economics and Disaggregated Supercomputing Hosts press for specific numbers on GPU cluster size and allocation across pre-training versus inference. Vipul discloses their 7,000 to 8,000 GPU fleet scale and calculates overall industry Capex economics.32:11–39:27 · The hosts pushing back 3/10 Inference Stack Optimization and Cloud Architecture Alessio references SemiAnalysis's critique of Together's pricing and speculative decoding setup. Vipul systematically clarifies where Dylan Patel's model made incorrect assumptions regarding input token pricing and hardware configurations.39:30–43:04 · The hosts pushing back 1/10 Multi-Dimensional Inference Co-Optimization Ce and Vipul explain the compounding benefits of co-optimizing across algorithms, model architectures, and custom kernels rather than isolating single improvements.43:04–49:17 · The hosts pushing back 2/10 Industry Benchmarking Challenges and Standardization Alessio asks about the AnyScale benchmark controversy. Ce and Vipul emphasize the systemic dangers of benchmarks creating bad optimization incentives and advocate for neutral third-party measurement.49:18–52:10 · The hosts pushing back 1/10 Fine-Tuning Spectrum and Enterprise Customer Engagements Swyx asks about customer engagement tiers for fine-tuning. Vipul outlines the spectrum from full consultative pre-training to serverless fine-tuning workflows.52:10–54:51 · The hosts pushing back 1/10 Advancements and Future Potential in Embeddings Alessio asks whether embedding models have reached a plateau. Ce explains why embeddings remain in their infancy, particularly regarding fine-grained semantics, negation, and data flywheel loops.54:51–1:03:41 · The hosts pushing back 2/10 State Space Models, Hybrids, and StripedHyena Swyx questions why researchers should care about subquadratic state space models beyond simple context length. Ce educates the hosts on memory footprint, execution patterns, and hybrid layer grafting in StripedHyena.1:03:41–1:06:17 · The hosts pushing back 2/10 The Case for 5,000 Tokens per Second Inference Swyx asks why anyone needs 5,000 tokens per second when humans read much slower. Vipul clarifies that machine-to-machine consumption, hardware card throughput, and interactive UX fundamentally require extreme generation speeds.1:06:19–1:11:40 · The hosts pushing back 1/10 Delivering a Pure Serverless AI Developer Experience Swyx shares his past experience running out of credits on Together. Vipul and Ce explain their pivot to a fully serverless, friction-free developer experience and outline hiring priorities.1:11:40–1:14:49 · The hosts pushing back 1/10 Lightning Round: Unsolved Questions and Positive AI Frameworks Ce discusses edge satellite communications as an alternative research passion, while Vipul calls for replacing science-fiction doomerism with constructive frameworks for advanced intelligence.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 38% · guest 62%0:00 · the hosts 38% · guest 62%3:00 · the hosts 12.8% · guest 87.2%3:00 · the hosts 12.8% · guest 87.2%6:00 · the hosts 11.7% · guest 88.3%6:00 · the hosts 11.7% · guest 88.3%9:00 · the hosts 24.9% · guest 75.1%9:00 · the hosts 24.9% · guest 75.1%12:00 · the hosts 20% · guest 80%12:00 · the hosts 20% · guest 80%15:00 · the hosts 23.2% · guest 76.8%15:00 · the hosts 23.2% · guest 76.8%18:00 · the hosts 23.7% · guest 76.3%18:00 · the hosts 23.7% · guest 76.3%21:00 · the hosts 30.2% · guest 69.8%21:00 · the hosts 30.2% · guest 69.8%24:00 · the hosts 24.9% · guest 75.1%24:00 · the hosts 24.9% · guest 75.1%27:00 · the hosts 35.4% · guest 64.6%27:00 · the hosts 35.4% · guest 64.6%30:00 · the hosts 31.8% · guest 68.2%30:00 · the hosts 31.8% · guest 68.2%33:00 · the hosts 9.6% · guest 90.4%33:00 · the hosts 9.6% · guest 90.4%36:00 · the hosts 6.7% · guest 93.3%36:00 · the hosts 6.7% · guest 93.3%39:00 · the hosts 29.4% · guest 70.6%39:00 · the hosts 29.4% · guest 70.6%42:00 · the hosts 26.4% · guest 73.6%42:00 · the hosts 26.4% · guest 73.6%45:00 · the hosts 29.9% · guest 70.1%45:00 · the hosts 29.9% · guest 70.1%48:00 · the hosts 26.9% · guest 73.1%48:00 · the hosts 26.9% · guest 73.1%51:00 · the hosts 23.5% · guest 76.5%51:00 · the hosts 23.5% · guest 76.5%54:00 · the hosts 46.3% · guest 53.7%54:00 · the hosts 46.3% · guest 53.7%57:00 · the hosts 28.7% · guest 71.3%57:00 · the hosts 28.7% · guest 71.3%1:00:00 · the hosts 13.4% · guest 86.6%1:00:00 · the hosts 13.4% · guest 86.6%1:03:00 · the hosts 25.2% · guest 74.8%1:03:00 · the hosts 25.2% · guest 74.8%1:06:00 · the hosts 24.8% · guest 75.2%1:06:00 · the hosts 24.8% · guest 75.2%1:09:00 · the hosts 18.8% · guest 81.2%1:09:00 · the hosts 18.8% · guest 81.2%1:12:00 · the hosts 2.1% · guest 97.9%1:12:00 · the hosts 2.1% · guest 97.9%1:15:00 · the hosts 0% · guest 0%1:15:00 · the hosts 0% · guest 0%
Sharpest disagreement ▶ 19:29 Ce directly pushes back on Swyx's AGI critique

Ce rejects the premise that targeted importance resampling contradicts general intelligence principles, reframing the problem as meta-learning and real-world deployment efficiency.

Hardest push from the hosts ▶ 19:07 Swyx challenges DSIR against general intelligence goals

Swyx pushes back on domain-specific data filtering by asserting that predetermining task distributions runs counter to the foundational premise of training AGI.

Biggest teaching moment ▶ 56:00 Ce explains why SSM advantages extend beyond long context

Ce educates Swyx on the broader systems benefits of state space models, demonstrating how smaller state footprints and decoupling quadratic dependencies enable massive batch sizes and cheaper execution patterns.

The host holds their own ▶ 32:10 Alessio confronts guests with SemiAnalysis inference tear-downs

Alessio quotes detailed technical findings from Dylan Patel on Together's memory bandwidth and speculative decoding architectures, forcing the guests to provide a granular technical response.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Vipul's Apple Background and Lessons on AI Systems 4511 Alessio asks an insightful question contrasting Apple's closed, polished ecosystem with Together's open philosophy. Vipul shares his deep background in spam filtering, early open domain Q&A at Apple, and his perspective on scaling laws.
Founding Together AI and Research-Led Systems Architecture 4511 Hosts ask about the intersection of academic research and startup founding. Ce and Vipul explain why data movement across stacks and gradient compression for decentralized computing drove the company's genesis.
RedPajama Dataset Evolution: V1 to V2 5612 Swyx and Alessio dig into RedPajama V1 and V2, observing the shift toward modular quality filtering signals. Ce explains the philosophy of turning dataset artifacts into tunable, multi-signal platforms rather than static snapshots.
Custom Model Training and Importance Resampling (DSIR) 6634 Swyx pushes back by arguing that targeted importance resampling (DSIR) violates the pursuit of general intelligence. Ce counters by reframing the trade-off space around deployment costs, meta-learning, and domain-targeted models.
Exploring Global Data Limits and Marketplace Dynamics 5412 Swyx brings up debates around YouTube token quality from Whisper and proprietary data walled gardens. Vipul and Ce acknowledge data fragmentation and propose fair marketplace dynamics for creators.
GPU Fleet Economics and Disaggregated Supercomputing 5612 Hosts press for specific numbers on GPU cluster size and allocation across pre-training versus inference. Vipul discloses their 7,000 to 8,000 GPU fleet scale and calculates overall industry Capex economics.
Inference Stack Optimization and Cloud Architecture 6623 Alessio references SemiAnalysis's critique of Together's pricing and speculative decoding setup. Vipul systematically clarifies where Dylan Patel's model made incorrect assumptions regarding input token pricing and hardware configurations.
Multi-Dimensional Inference Co-Optimization 4511 Ce and Vipul explain the compounding benefits of co-optimizing across algorithms, model architectures, and custom kernels rather than isolating single improvements.
Industry Benchmarking Challenges and Standardization 5522 Alessio asks about the AnyScale benchmark controversy. Ce and Vipul emphasize the systemic dangers of benchmarks creating bad optimization incentives and advocate for neutral third-party measurement.
Fine-Tuning Spectrum and Enterprise Customer Engagements 4511 Swyx asks about customer engagement tiers for fine-tuning. Vipul outlines the spectrum from full consultative pre-training to serverless fine-tuning workflows.
Advancements and Future Potential in Embeddings 5511 Alessio asks whether embedding models have reached a plateau. Ce explains why embeddings remain in their infancy, particularly regarding fine-grained semantics, negation, and data flywheel loops.
State Space Models, Hybrids, and StripedHyena 5712 Swyx questions why researchers should care about subquadratic state space models beyond simple context length. Ce educates the hosts on memory footprint, execution patterns, and hybrid layer grafting in StripedHyena.
The Case for 5,000 Tokens per Second Inference 4522 Swyx asks why anyone needs 5,000 tokens per second when humans read much slower. Vipul clarifies that machine-to-machine consumption, hardware card throughput, and interactive UX fundamentally require extreme generation speeds.
Delivering a Pure Serverless AI Developer Experience 4411 Swyx shares his past experience running out of credits on Together. Vipul and Ce explain their pivot to a fully serverless, friction-free developer experience and outline hiring priorities.
Lightning Round: Unsolved Questions and Positive AI Frameworks 3411 Ce discusses edge satellite communications as an alternative research passion, while Vipul calls for replacing science-fiction doomerism with constructive frameworks for advanced intelligence.

Statements from this episode (26)

Assertion Supported
Prakash: Together AI builds cloud footprint avoiding major hyperscalers
“The way we structure our cloud is by combining data centers around the world instead of you know we are today not located in hyperscalers. We have built a footprint of you know, AI supercomputers in this sort of a disaggregated, decentralized manner.”
Vipul Ved Prakash Feb 8, 2024 ▶ 2:06
Insight
Prakash: Algorithms that improve predictably with scale mark a new computing era
“We've never had algorithms That improving capabilities with scale out. It's a, this is almost, ah, You know, new era of computing.”
Vipul Ved Prakash Feb 8, 2024 ▶ 5:10
Insight
Zhang: ML systems optimization is fundamentally always about data movement
“Fundamentally, the thing we are actually optimizing is actually not that different. It's always about data movement across essentially all the stacks, right? So when you do distributed, like computing, it's about communication across different machines. When y…”
Ce Zhang Feb 8, 2024 ▶ 6:19
Opinion
Prakash: Foundation models lack data moats and are constrained by capital
“There aren't really big data moats around foundation models. They are built from a subset of the web. What is difficult is the cost of capital to build these, and what, one of the ways in which you can reduce this cost is by making more efficient systems.”
Vipul Ved Prakash Feb 8, 2024 ▶ 7:48
Assertion Supported
Zhang: RedPajama-V2 features 40 pre-computed quality signals for custom filtering
“So that's why in REST-PYRON V-II, we kind of overlay the data set, it's like, 40 different pre-computed quality signal, right? If you want to reproduce your best effort, like, C-Four filter, it's kind of like, 20 lines of code.”
Ce Zhang Feb 8, 2024 ▶ 13:29
Insight
Zhang: Pre-training data evolved from a model byproduct into a standalone asset
“So, so I think one fundamental thing that changed in the last year, essentially, in the beginning when people think about data, is, is always like a byproduct of the model, right? You release the model, you also release the data, right? The data side is there …”
Ce Zhang Feb 8, 2024 ▶ 14:43
Opinion
Prakash: Transformer architectures will not reach 5,000 tokens per second inference
“We are running into you know, the limits of how fast you can make transformers and you know, we want inference at 5000 tokens per second. And I don't think we will get there with transformers”
Vipul Ved Prakash Feb 8, 2024 ▶ 16:34
Insight
Zhang: Targeted task training yields smaller, cheaper, and more accurate models
“The benefit you can get out of that is you could build a, Better open model, often smaller, often easier to do inference if you know what you want, right? So I think the whole trade-off would be, and the x-axis would be how generic the hosting will be. The y-a…”
Ce Zhang Feb 8, 2024 ▶ 19:54
Opinion
Zhang: The world is not running out of AI training data
“I don't think we are running out of data on earth. Right, so think about it globally... But I do think there are many organizations in the world have enough data to actually train, like, very, very good models, right? So, I mean, they are not public available,…”
Ce Zhang Feb 8, 2024 ▶ 20:56
Prediction Not checkable as stated
Prakash predicts data marketplaces for AI training will increasingly emerge
“I think you need to have some kind of marketplace for figuring out how to get this you know, data into models and have, I think you'll increasingly see more of that”
Vipul Ved Prakash Feb 8, 2024 ▶ 24:10
Disclosure
Prakash: Top 5 models on Together AI inference are fine-tuned open models
“I would say right now the top five models on our inference stack are probably all fine-tuned versions of open models.”
Vipul Ved Prakash Feb 8, 2024 ▶ 25:22
Disclosure
Prakash: Together AI operates a fleet of 7,000 to 8,000 GPUs
“We have close to seven to 8000 GPUs today. It's growing monthly.”
Vipul Ved Prakash Feb 8, 2024 ▶ 27:33
Prediction Held up
Prakash predicts up to 5 million AI GPUs will sell in 2024
“There is four to five million GPUs that will be sold this year. NVIDIA and others.”
Vipul Ved Prakash Feb 8, 2024 ▶ 30:43
Assertion Supported
Prakash: SemiAnalysis inference report erred by assuming unpriced input tokens
“I think there were some errors in that analysis. In particular we were trying to decode it, and one of the things we noticed is that it assumed that input tokens weren't being priced. So I think that may have been an error in the model.”
Vipul Ved Prakash Feb 8, 2024 ▶ 33:09
Disclosure
Prakash: Together AI keeps its inference stack proprietary for competitive advantage
“I think on the inference stack, there are open source inference stacks which are pretty good, and it gives us, you know, definitely today it gives us a competitive advantage to have the best one, and So we're not sort of rushing out to release everything about…”
Vipul Ved Prakash Feb 8, 2024 ▶ 35:40
Insight
Zhang: Next 10x AI inference gain requires multi-layer co-optimization
“If you only push on one direction, you are going to reach diminution return really, really quickly. Yeah, there's only that much you can do on the system side, only that much you can do on the algorithm side. And since the only big thing that's going to happen…”
Ce Zhang Feb 8, 2024 ▶ 40:16
Insight
Zhang: Good benchmarks must anticipate how developers over-optimize toward metrics
“A good benchmark should think about how it's going to incentivize the field to actually move forward, right? So the benchmark will become kind of standard. How are people going to over optimize to the benchmark because people are going to do that? And when peo…”
Ce Zhang Feb 8, 2024 ▶ 44:23
Assertion Not checkable as stated
Prakash: AnyScale LLM benchmarks conflicted with internal provider testing
“Everyone sort of had a reaction to it because it just didn't match their benchmarks that we've all run internally against different services.”
Vipul Ved Prakash Feb 8, 2024 ▶ 45:48
Assertion Supported
Prakash: Small engineer collectives are driving top open-source leaderboard models
“You have small collectives of, you know engineers who have created, who are now creating the top models on open source leaderboards, and I have tried out all sorts of different sort of, you know, data recipes, creating synthetic data”
Vipul Ved Prakash Feb 8, 2024 ▶ 51:15
Prediction Not checkable as stated
Zhang predicts much faster, diverse new embedding models within couple years
“So I think for the next couple years, yeah, we will see a whole bunch of new embeddings maybe of different sites, and much, much faster than today.”
Ce Zhang Feb 8, 2024 ▶ 53:29
Assertion Not checkable as stated
Prakash: 40% to 45% of Together AI's team is dedicated to research
“It's like 40, 45% I was counting this morning.”
Vipul Ved Prakash Feb 8, 2024 ▶ 55:04
Assertion Not checkable as stated
Prakash: No provider satisfies machine-to-machine AI demand for extreme token speeds
“There are applications that are, you know, consuming the tokens that are produced from unmodel, so they're not necessarily being read or heard by humans. So that's a place where we see that level of requirement today that really nobody can quite satisfy.”
Vipul Ved Prakash Feb 8, 2024 ▶ 1:04:25
Prediction Not checkable as stated
Prakash: Near-instant model inference will spawn new application paradigms
“Once you can get sort of an immediate answer from a model it starts working in a different way and you know, new types of applications will be created.”
Vipul Ved Prakash Feb 8, 2024 ▶ 1:05:21
Insight
Zhang: Combining fine-tuning and RAG provides superior performance boosts
“Combining all those techniques all together, right? So we'll give you essentially another boost, right? So that kind of one thing that we learn on the technical side.”
Ce Zhang Feb 8, 2024 ▶ 1:08:27
Assertion Not checkable as stated
Prakash: Together AI has 38 employees as of February 2024
“We are 38 people on the team and we are hiding across all the areas you know.”
Vipul Ved Prakash Feb 8, 2024 ▶ 1:09:10
Opinion
Prakash: AI discourse is overly dominated by sci-fi-informed doomerism
“I think we have had this Very you know, sort of a doomerism view of it really kind of informed by science fiction, you know, dystopian science fiction and Terminator, and I don't think we have a kind of a positive or a realistic, really framework coming from, …”
Vipul Ved Prakash Feb 8, 2024 ▶ 1:13:19
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.