CLIP, every mention
21 scenes, the whole family · ← back to CLIP
tap a year for its mentions
every year anyone Vibhu Sapra 13Peter Robicheaux 11comfyanonymous (Comfy) 6Ben Firshman 2Alessio Fanelli 2Varun Mohan 1Shawn Wang 1Ludwig Schmidt 1Karina Nguyen 1Jerry Liu 1
Verbatim, from the transcripts: the passages where CLIP comes up
Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
- ▶ 19:06 Alessio Fanelli Same with, uh, clip, a meta clip, where you go from just captioning to
Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
- ▶ 14:04 unnamed speaker This was pretty common when we went from like, uh, Clip and Dali, right? 2 times in the scene
Terminal-Bench 2.0: the most impt coding agent benchmark of 2025 gets a v2! Launch + Q&A w/ founders
- ▶ 22:15 Ludwig Schmidt So we started this back in the good old days, building datasets for club, then for language model pre-training, then for reasoning, fine tuning, and now for agent training data.
🕰️ The Oral History of Windsurf (ft. Varun Mohan, Scott Wu, Jeff Wang, Kevin Hou, Anshul R)
- ▶ 8:26 Varun Mohan If it's a vision, they should be fine tuning resonant or using something like clip, like the less work they can do the better.
The Agent Reasoning Interface: Claude, ChatGPT Canvas, Tasks, Operator — with Karina Nguyen, OpenAI
- ▶ 4:49 Karina Nguyen I used, like, Clip for, like, fashion recommendation search.
AI Engineering for Art - with comfyanonymous
- ▶ 19:31 comfyanonymous (Comfy) They also kind of work on SDXL because SDXL has the, has two text encoders and one of them is the same as the, as the SD 1.5 clip L.
- ▶ 20:06 Shawn Wang Do people experiment a lot on, just on the clip side, uh, there's like Siglip, there's Blip, like do people experiment a lot on those?
- ▶ 20:23 comfyanonymous (Comfy) So basically someone, uh, fine tune the clip model to accept longer prompts. 5 times in the scene
- ▶ 35:15 Alessio Fanelli So unlike all the other text to image, you have a very like deep, so you have like a separate node for like clip and code.
Best of 2024 in Vision [LS Live @ NeurIPS]
- ▶ 22:21 Peter Robicheaux So, it hypothesizes that models that have been initialized with, with Clip as their vision encoder, they don't have fine-grained details and the features extracted using Clip because, um, 8 times in the scene
- ▶ 36:30 Peter Robicheaux It's image or internet scale data, very high quality data created by the, um, data filtering networks paper essentially, um, which is maybe the best clip data that exists. 3 times in the scene
- ▶ 40:46 Isaac Robinson You know, we see all these super, super complicated things, and it's not super easy to build something that just transfers naturally like that, whereas image classification, you know, clip pre-training transfers super, super easily.
[Paper Club] Embeddings in 2024: OpenAI, Nomic Embed, Jina Embed, cde-small-v1 - with swyx
- ▶ 27:42 unnamed speaker So they took embeddings v three and jammed it into clip. 3 times in the scene
[Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz
- ▶ 2:00 Vibhu Sapra And some of the other papers we covered in paper club or stuff like clip open clip, which are captioners, the whole history of these and how captioning is pretty important for this. 6 times in the scene
- ▶ 33:57 Vibhu Sapra The interesting thing there was the only closed code for the, the closed open data was just, they used Clip, large Clip, right? 7 times in the scene
[Paper Club] Berkeley Function Calling Paper Club! — Sam Julien, Writer
- ▶ 39:44 unnamed speaker I think they pre-train it from clip.
High Agency Pydantic over VC Backed Frameworks — with Jason Liu of Instructor
- ▶ 5:05 unnamed speaker Uh, one more system that I'm interested in finding out more about is your similarity search system using Clip, uh, and GPT-T-E-Embedding in FICE, um, where you, you said 50, over fifty million dollars in annual revenue.
A Brief History of the Open Source AI Hacker - with Ben Firshman of Replicate
- ▶ 28:39 Ben Firshman People started, OpenAI created this, created Clip, and people started smushing Clip and Gans together to produce image generation models, and this started with, um, you know, it was just a bunch of, like, tinkerers on Discord, basically. 2 times in the scene
The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
- ▶ 22:52 unnamed speaker That's for clip guidance.
RAG is a hack - with Jerry Liu of LlamaIndex
RWKV: Reinventing RNNs for the Transformer Era
- ▶ 51:32 unnamed speaker There's clip-guided autoencoder, but I don't think that's, yeah,