Together AI, every mention

50 scenes · ← back to Together AI

tap a year for its mentions
0040880152023202420252026episodesmentions
08152023202420252026episodes it came up in
0047.58152023202420252026episodesmentions per episode

every year anyone Alessio Fanelli 41Shawn Wang 18Vipul Ved Prakash 4Michael Swix (Swyx) 4Stephanie Palazzolo 2Dylan Patel 2Vibhu (Veebu) 1Steve Ruiz 1Sherwin Wu 1Eugene Cheah 1

Verbatim, from the transcripts: the passages where Together AI comes up

loading…

Podcast Crossover: AIE, AGI, frontier lab strategy with ​ ⁨@matthew_berman⁩ and @swyxtv Jul 10, 2026 · 1 mention

The Future of AI Infra: from Kubernetes to Agent Sandboxes — Akshat Bubna, Modal CTO Jul 8, 2026 · 1 mention

  • ▶ 14:39 unnamed speaker There's a bunch of inference providers, which, you know, provide this fireworks, does this as a service together or whatnot.

⚡️ Reverse Engineering OpenAI's Training Data — Pratyush Maini, Datology Feb 10, 2026 · 1 mention

  • ▶ 26:17 unnamed speaker It's interesting to me that every single generation is a different company, you know, like, oh, it was Together AI,

Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith Jan 9, 2026 · 1 mention

  • ▶ 0:54 Shawn Wang But you had together, you had perplexity and, uh, I think we just started chatting there.

The Future of Email: Superhuman CTO on Your Inbox As the Real AI Agent (Not ChatGPT) — Loïc Houssier Dec 11, 2025 · 2 mentions

  • ▶ 32:55 Shawn Wang The inference provider for open models compared to, let's say the fireworks and the together AIs. 2 times in the scene

Why RL Won — Kyle Corbitt, OpenPipe (acq. CoreWeave) Oct 16, 2025 · 1 mention

DevDay 2025: Apps SDK, Agent Kit, MCP, Codex and why Prompting is More Important than Ever Oct 7, 2025 · 1 mention

  • ▶ 17:18 Sherwin Wu You know, take your pick from wherever, even like open source ones on together, uh, and see the, see the results, uh, in our product.

⚡️ Beyond Transformers with Power Retention Sep 23, 2025 · 4 mentions

  • ▶ 24:54 unnamed speaker So if you are together or fireworks or any of these like inference providers, should you just be doing this for every single model? 4 times in the scene

A Technical History of Generative Media Sep 8, 2025 · 1 mention

  • ▶ 53:20 unnamed speaker So it's really interesting because I think this is what Together AI did with Red Pajama is they, they actually built a dataset for language models to help people create more open language models that they can serve.

Better Data is All You Need — Ari Morcos, Datology Aug 29, 2025 · 1 mention

  • ▶ 59:52 Ari Morcos And now this has largely been commoditized by things like SageMaker and Together and lots of different folks that help you on the training side.

The AI Agenda: GPT5 leaks and the business of AI News — Steph Palazzolo, The Information Aug 6, 2025 · 3 mentions

  • ▶ 8:34 Stephanie Palazzolo You mentioned some of them, Fireworks, Modal, there's Together, there's Base 10. 2 times in the scene
  • ▶ 1:10:54 Shawn Wang There are some notable exceptions, primarily together with the Mombard architecture, um, recursal with RWKV.

The Shape of Compute (Chris Lattner of Modular) Jun 13, 2025 · 1 mention

  • ▶ 36:34 Shawn Wang And I think a lot of the other inference team, because effectively every team is, is a startup, like the fireworks together, you know, all those guys, your business model is very different from them.

The AI Coding Factory May 29, 2025 · 1 mention

  • ▶ 45:30 unnamed speaker We did an episode with Together AI maybe a year ago or so, and we were talking about what inference speed actually we needed, and they always argued we need to get to, like, 5000 tokens a second, and we were chatting whether or not that…

SF Compute: Commoditizing Compute Apr 11, 2025 · 4 mentions

  • ▶ 9:11 Michael Swix (Swyx) And then you had some analysis of where every player sits in this, including CoreWeave, but also Together, and Modal, and all these other guys.
  • ▶ 16:24 Michael Swix (Swyx) and then I promise we'll get to, uh, SF Compute, is why will, uh, DigitalOcean and Together lose money on their clusters? 3 times in the scene

DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing) Jan 19, 2025 · 2 mentions

  • ▶ 8:28 unnamed speaker So I think a lot of companies as well, like together, they'll also release like quantized versions of the Lama models, right? 2 times in the scene

2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents Jan 1, 2025 · 2 mentions

  • ▶ 41:03 Shawn Wang We have now interviewed both together and fireworks and replicates.
  • ▶ 56:30 Alessio Fanelli The code sandbox got acquired by Together AI last week, um, which they're now also gonna offer as an API.

2024 in Post-Transformer Architectures: State Space Models, RWKV [Latent Space LIVE! @ NeurIPS 2024] Dec 24, 2024 · 2 mentions

  • ▶ 0:13 Dan Fu Um, I met together AI
  • ▶ 24:50 Eugene Cheah And, and I believe together AI worked on the law cats for, for the Mamba side of things, and, and we took some ideas from there as well, and we essentially did that for RWKV.

[Paper Club] DocETL: Agentic Query Rewriting + Eval for Complex Document Processing w Shreya Shankar Nov 29, 2024 · 1 mention

  • ▶ 43:03 Vibhu (Veebu) So is the benefit there trying to save on performance of a large model or like optimization that you can't run a large one, like mixture of agents kind of shows how from, um, together shows how you can use mixture of agents of small models…

Why Compound AI + Open Source will beat Closed AI — with Lin Qiao, CEO of Fireworks AI Nov 25, 2024 · 1 mention

  • ▶ 38:55 Shawn Wang There's Replicate, there's Together, there's, like, Elepton, there's, like, a whole bunch of other players.

Production AI Engineering starts with Evals Oct 11, 2024 · 2 mentions

  • ▶ 1:37:21 unnamed speaker You can use Together, Fireworks, all these guys. 2 times in the scene

The Winds of AI Winter (Q2 Four Wars of the AI Stack Recap) Aug 2, 2024 · 2 mentions

  • ▶ 11:23 Alessio Fanelli I also don't know what the hosting up hosting options are as far as like scaling, you know, I don't know if the fireworks and togethers of the world, how much capacity they actually have to serve this model, because at the end of the day,…
  • ▶ 25:10 unnamed speaker No one thinks that he can do inference at 13 X cheaper than the fireworks together, right?

[LLM Paper Club] Llama 3.1 Paper: The Llama Family of Models Jul 29, 2024 · 5 mentions

  • ▶ 55:55 unnamed speaker Um, and also, you know, using the together API, very biased. 4 times in the scene
  • ▶ 1:05:43 unnamed speaker I'm writing an email now between cloud providers for 3.1 70 B just to see if like, you know, together's, uh, the eight versus, I don't know, fireworks or something like this has a difference or versus Glock.

Open Source AI is AI we can Trust — with Soumith Chintala of Meta AI Mar 6, 2024 · 1 mention

  • ▶ 28:39 unnamed speaker Thing here, like, uh, I think, like, Together and Fireworks and all these people are trying to build some faster CUDA kernels and faster, like, you know, hardware kernels in general.

Building an open AI company - with Ce and Vipul of Together AI Feb 8, 2024 · 52 mentions

  • ▶ 0:10 Shawn Wang Hey, and today we have, we're together with together. 6 times in the scene
  • ▶ 2:28 Alessio Fanelli I know before together you were Apple, so you go from like the most walled garden, private, we don't say anything company to, we want everything to be open and everybody to know somebody. 3 times in the scene
  • ▶ 5:45 Alessio Fanelli You know, we already had three DAO from together, and we talked about Hazy. 6 times in the scene
  • ▶ 10:18 Alessio Fanelli I want to try and fill in some of the blanks in the history of together. 6 times in the scene
  • ▶ 15:28 Alessio Fanelli Um, and you guys have a custom model training platform on, on Together Tube. 2 times in the scene
  • ▶ 25:22 Vipul Ved Prakash I would say right now the top five models on our inference stack are probably all fine-tuned versions of open models. 4 times in the scene
  • ▶ 32:20 Alessio Fanelli And this post he said, our model indicates that together it's better off using two a 180 gig system rather than a each 100 based system. 6 times in the scene
  • ▶ 41:27 Shawn Wang Um, I think I finally get the name of the company, like, bring it together.
  • ▶ 43:10 Alessio Fanelli Uh, so we're, we're actually having to meet on the podcast tomorrow who also talked about, uh, kind of came to your guys, uh, support about how, yeah, how important it's not just like, oh, together saying this benchmark is not good because…
  • ▶ 52:11 Alessio Fanelli Um, and then just to, I guess, wrap up the together, it's almost becoming like a platform as a service, you know, because now you release together embeddings. 2 times in the scene
  • ▶ 54:57 Shawn Wang Well, first of all, uh, how much of the company is dedicated to research? 3 times in the scene
  • ▶ 1:06:20 Alessio Fanelli Anything we missed about together as a product? 3 times in the scene
  • ▶ 1:08:57 Alessio Fanelli What should people that want to work together know? 6 times in the scene
  • ▶ 1:11:53 Alessio Fanelli Um, so maybe another way to think about it is if you weren't building together, what would you be working on? 3 times in the scene

The Four Wars of the AI Stack - Dec 2023 Recap Jan 26, 2024 · 5 mentions

  • ▶ 21:54 unnamed speaker Yeah, we can figure out a name, but we have Modal, Together, Replicate, um, there's a, there's a lot coming up. 3 times in the scene
  • ▶ 34:18 unnamed speaker And I mean, together it's doing so much for like three dial and like fresh attention to and whatnot. 2 times in the scene

The Accidental AI Canvas - with Steve Ruiz of tldraw Jan 5, 2024 · 1 mention

  • ▶ 1:10:00 Steve Ruiz so I guess the way that, the way that, yes, we, we probably could on for something like Together or that DrawFast thing, uh, make a, a tiny little SaaS app, you know, give me 10 dollars a month, play with this thing, and, uh,

The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis Dec 5, 2023 · 2 mentions

  • ▶ 26:25 Dylan Patel So there's, like, there's stuff like speculative decoding, and then, uh, you know, Together did something really cool, and they put in open source, of course, Medusa, right? 2 times in the scene

FlashAttention-2: Making Transformers 800% faster AND exact Aug 3, 2023 · 12 mentions

  • ▶ 0:43 unnamed speaker And in the meantime, just to get, you know, a low pressure thing, you're a chief scientist at Together as well, which, uh, is the company behind the red pajama. 2 times in the scene
  • ▶ 51:00 unnamed speaker Obviously, together, you know, when Red Pajama came out, which was a, you know, an open clone of, like, the Llama one, um, pre-training data set, it was a big thing in the industry. 3 times in the scene
  • ▶ 1:00:52 unnamed speaker What made you decide to join together? 7 times in the scene
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.