CLIP, every mention

21 scenes, the whole family · ← back to CLIP

tap a year for its mentions
002044082023202420252026episodesmentions
0482023202420252026episodes it came up in
002.54582023202420252026episodesmentions per episode

every year anyone Vibhu Sapra 13Peter Robicheaux 11comfyanonymous (Comfy) 6Ben Firshman 2Alessio Fanelli 2Varun Mohan 1Shawn Wang 1Ludwig Schmidt 1Karina Nguyen 1Jerry Liu 1

Verbatim, from the transcripts: the passages where CLIP comes up

loading…

Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten Aug 3, 2026 · 1 mention

Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He Jun 1, 2026 · 2 mentions

  • ▶ 14:04 unnamed speaker This was pretty common when we went from like, uh, Clip and Dali, right? 2 times in the scene

Terminal-Bench 2.0: the most impt coding agent benchmark of 2025 gets a v2! Launch + Q&A w/ founders Nov 8, 2025 · 1 mention

  • ▶ 22:15 Ludwig Schmidt So we started this back in the good old days, building datasets for club, then for language model pre-training, then for reasoning, fine tuning, and now for agent training data.

🕰️ The Oral History of Windsurf (ft. Varun Mohan, Scott Wu, Jeff Wang, Kevin Hou, Anshul R) Jul 28, 2025 · 1 mention

  • ▶ 8:26 Varun Mohan If it's a vision, they should be fine tuning resonant or using something like clip, like the less work they can do the better.

The Agent Reasoning Interface: Claude, ChatGPT Canvas, Tasks, Operator — with Karina Nguyen, OpenAI Feb 1, 2025 · 1 mention

AI Engineering for Art - with comfyanonymous Jan 4, 2025 · 8 mentions

  • ▶ 19:31 comfyanonymous (Comfy) They also kind of work on SDXL because SDXL has the, has two text encoders and one of them is the same as the, as the SD 1.5 clip L.
  • ▶ 20:06 Shawn Wang Do people experiment a lot on, just on the clip side, uh, there's like Siglip, there's Blip, like do people experiment a lot on those?
  • ▶ 20:23 comfyanonymous (Comfy) So basically someone, uh, fine tune the clip model to accept longer prompts. 5 times in the scene
  • ▶ 35:15 Alessio Fanelli So unlike all the other text to image, you have a very like deep, so you have like a separate node for like clip and code.

Best of 2024 in Vision [LS Live @ NeurIPS] Dec 22, 2024 · 12 mentions

  • ▶ 22:21 Peter Robicheaux So, it hypothesizes that models that have been initialized with, with Clip as their vision encoder, they don't have fine-grained details and the features extracted using Clip because, um, 8 times in the scene
  • ▶ 36:30 Peter Robicheaux It's image or internet scale data, very high quality data created by the, um, data filtering networks paper essentially, um, which is maybe the best clip data that exists. 3 times in the scene
  • ▶ 40:46 Isaac Robinson You know, we see all these super, super complicated things, and it's not super easy to build something that just transfers naturally like that, whereas image classification, you know, clip pre-training transfers super, super easily.

[Paper Club] Embeddings in 2024: OpenAI, Nomic Embed, Jina Embed, cde-small-v1 - with swyx Dec 1, 2024 · 3 mentions

  • ▶ 27:42 unnamed speaker So they took embeddings v three and jammed it into clip. 3 times in the scene

[Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz Oct 13, 2024 · 13 mentions

  • ▶ 2:00 Vibhu Sapra And some of the other papers we covered in paper club or stuff like clip open clip, which are captioners, the whole history of these and how captioning is pretty important for this. 6 times in the scene
  • ▶ 33:57 Vibhu Sapra The interesting thing there was the only closed code for the, the closed open data was just, they used Clip, large Clip, right? 7 times in the scene

[Paper Club] Berkeley Function Calling Paper Club! — Sam Julien, Writer Oct 5, 2024 · 1 mention

  • ▶ 39:44 unnamed speaker I think they pre-train it from clip.

High Agency Pydantic over VC Backed Frameworks — with Jason Liu of Instructor Apr 24, 2024 · 1 mention

  • ▶ 5:05 unnamed speaker Uh, one more system that I'm interested in finding out more about is your similarity search system using Clip, uh, and GPT-T-E-Embedding in FICE, um, where you, you said 50, over fifty million dollars in annual revenue.

A Brief History of the Open Source AI Hacker - with Ben Firshman of Replicate Feb 28, 2024 · 2 mentions

  • ▶ 28:39 Ben Firshman People started, OpenAI created this, created Clip, and people started smushing Clip and Gans together to produce image generation models, and this started with, um, you know, it was just a bunch of, like, tinkerers on Discord, basically. 2 times in the scene

The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert Jan 11, 2024 · 1 mention

  • ▶ 22:52 unnamed speaker That's for clip guidance.

RAG is a hack - with Jerry Liu of LlamaIndex Oct 12, 2023 · 1 mention

  • ▶ 7:14 Jerry Liu I actually think once the multimodal models come out, I think there's just, like, mathematically nicer properties if you can just get, like, joint multimodal embeddings, like clip, clip style.

RWKV: Reinventing RNNs for the Transformer Era Aug 31, 2023 · 1 mention

  • ▶ 51:32 unnamed speaker There's clip-guided autoencoder, but I don't think that's, yeah,
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.