Feb 8, 2024 · 1h 15m · latent-space
Building an open AI company - with Ce and Vipul of Together AI
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
Together AI co-founders Vipul Ved Prakash and Ce Zhang join the Latent Space podcast to discuss their open-source platform, disaggregated cloud infrastructure, and multi-dimensional inference optimizations. They detail dataset initiatives like RedPajama, hybrid architectures such as StripedHyena, and the systems engineering required to deliver high-throughput, serverless AI for developers.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 23.5% of the talking time here. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Ce rejects the premise that targeted importance resampling contradicts general intelligence principles, reframing the problem as meta-learning and real-world deployment efficiency.
Hardest push from the hosts ▶ 19:07 Swyx challenges DSIR against general intelligence goalsSwyx pushes back on domain-specific data filtering by asserting that predetermining task distributions runs counter to the foundational premise of training AGI.
Biggest teaching moment ▶ 56:00 Ce explains why SSM advantages extend beyond long contextCe educates Swyx on the broader systems benefits of state space models, demonstrating how smaller state footprints and decoupling quadratic dependencies enable massive batch sizes and cheaper execution patterns.
The host holds their own ▶ 32:10 Alessio confronts guests with SemiAnalysis inference tear-downsAlessio quotes detailed technical findings from Dylan Patel on Together's memory bandwidth and speculative decoding architectures, forcing the guests to provide a granular technical response.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Vipul's Apple Background and Lessons on AI Systems | 4 | 5 | 1 | 1 | Alessio asks an insightful question contrasting Apple's closed, polished ecosystem with Together's open philosophy. Vipul shares his deep background in spam filtering, early open domain Q&A at Apple, and his perspective on scaling laws. | |
| Founding Together AI and Research-Led Systems Architecture | 4 | 5 | 1 | 1 | Hosts ask about the intersection of academic research and startup founding. Ce and Vipul explain why data movement across stacks and gradient compression for decentralized computing drove the company's genesis. | |
| RedPajama Dataset Evolution: V1 to V2 | 5 | 6 | 1 | 2 | Swyx and Alessio dig into RedPajama V1 and V2, observing the shift toward modular quality filtering signals. Ce explains the philosophy of turning dataset artifacts into tunable, multi-signal platforms rather than static snapshots. | |
| Custom Model Training and Importance Resampling (DSIR) | 6 | 6 | 3 | 4 | Swyx pushes back by arguing that targeted importance resampling (DSIR) violates the pursuit of general intelligence. Ce counters by reframing the trade-off space around deployment costs, meta-learning, and domain-targeted models. | |
| Exploring Global Data Limits and Marketplace Dynamics | 5 | 4 | 1 | 2 | Swyx brings up debates around YouTube token quality from Whisper and proprietary data walled gardens. Vipul and Ce acknowledge data fragmentation and propose fair marketplace dynamics for creators. | |
| GPU Fleet Economics and Disaggregated Supercomputing | 5 | 6 | 1 | 2 | Hosts press for specific numbers on GPU cluster size and allocation across pre-training versus inference. Vipul discloses their 7,000 to 8,000 GPU fleet scale and calculates overall industry Capex economics. | |
| Inference Stack Optimization and Cloud Architecture | 6 | 6 | 2 | 3 | Alessio references SemiAnalysis's critique of Together's pricing and speculative decoding setup. Vipul systematically clarifies where Dylan Patel's model made incorrect assumptions regarding input token pricing and hardware configurations. | |
| Multi-Dimensional Inference Co-Optimization | 4 | 5 | 1 | 1 | Ce and Vipul explain the compounding benefits of co-optimizing across algorithms, model architectures, and custom kernels rather than isolating single improvements. | |
| Industry Benchmarking Challenges and Standardization | 5 | 5 | 2 | 2 | Alessio asks about the AnyScale benchmark controversy. Ce and Vipul emphasize the systemic dangers of benchmarks creating bad optimization incentives and advocate for neutral third-party measurement. | |
| Fine-Tuning Spectrum and Enterprise Customer Engagements | 4 | 5 | 1 | 1 | Swyx asks about customer engagement tiers for fine-tuning. Vipul outlines the spectrum from full consultative pre-training to serverless fine-tuning workflows. | |
| Advancements and Future Potential in Embeddings | 5 | 5 | 1 | 1 | Alessio asks whether embedding models have reached a plateau. Ce explains why embeddings remain in their infancy, particularly regarding fine-grained semantics, negation, and data flywheel loops. | |
| State Space Models, Hybrids, and StripedHyena | 5 | 7 | 1 | 2 | Swyx questions why researchers should care about subquadratic state space models beyond simple context length. Ce educates the hosts on memory footprint, execution patterns, and hybrid layer grafting in StripedHyena. | |
| The Case for 5,000 Tokens per Second Inference | 4 | 5 | 2 | 2 | Swyx asks why anyone needs 5,000 tokens per second when humans read much slower. Vipul clarifies that machine-to-machine consumption, hardware card throughput, and interactive UX fundamentally require extreme generation speeds. | |
| Delivering a Pure Serverless AI Developer Experience | 4 | 4 | 1 | 1 | Swyx shares his past experience running out of credits on Together. Vipul and Ce explain their pivot to a fully serverless, friction-free developer experience and outline hiring priorities. | |
| Lightning Round: Unsolved Questions and Positive AI Frameworks | 3 | 4 | 1 | 1 | Ce discusses edge satellite communications as an alternative research passion, while Vipul calls for replacing science-fiction doomerism with constructive frameworks for advanced intelligence. |