LLaVA, every mention
11 scenes · ← back to LLaVA
tap a year for its mentions
every year anyone Vibhu Sapra 6Peter Robicheaux 4Shawn Wang 3Ethan Sutin 1Ben Firshman 1
Verbatim, from the transcripts: the passages where LLaVA comes up
Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
- ▶ 18:57 Shawn Wang Yeah, so, uh, we've covered Haltien before, who was the author of the Lava paper that did this, uh, a while ago, and I think that that's very foundational work for anyone who hasn't done vision work before.
Moonlake: Interactive, Multimodal World Models — with Chris Manning and Fan-yun Sun
- ▶ 1:00:46 unnamed speaker The early vision language models we saw were like lava style adapters, right?
2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents
- ▶ 30:39 Shawn Wang I, I feel like, um, like, like you didn't last year, you had people like Hau Tien who worked on Lava, uh, which is take Lama and add vision.
Best of 2024 in Vision [LS Live @ NeurIPS]
- ▶ 24:07 Peter Robicheaux And like all these other models, including lava, um, do the same thing, right? 4 times in the scene
[Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz
- ▶ 1:34 Vibhu Sapra They, they label a lot of data and then they, they create a vision language model that's open weight, something like lava, but all they're doing is kind of distilling that proprietary knowledge into an open knowledge. 3 times in the scene
- ▶ 13:23 Vibhu Sapra You train the right data set and it gets good, but, um, it's still cool to see how it's like significantly better than lava.
- ▶ 31:08 Vibhu Sapra Is decent, but, um, I, I guess the interesting differentiation there is how far of a drop other stuff like the lava and whatnot is, uh, 2 times in the scene
The 10,000x Yolo Researcher Metagame — with Yi Tay of Reka
- ▶ 1:30:05 unnamed speaker Basically, the general idea is that most multimodal models, uh, like are, like Lava or Flamingo, which are late fusion, uh, which is you freeze, freeze, and then you sort of join together, um, versus early fusion, where you do it properly,…
How to train a Million Context LLM — with Mark Huang of Gradient.ai
- ▶ 1:01:00 Shawn Wang That was kind of done with, like, you know, what we now call late fusion models with, like, lava and, um, flamingo and, uh, you know, all the others.
Personal AI Meetup - Bee, BasedHardware, LangChain LangFriend, Deepgram EmilyAI
- ▶ 19:45 Ethan Sutin Um, so that's kind of, you know, I think a lot of us here have the same, same idea, and so that's what I've been working on for, since about November, um, when Lava came out, so it first started working on, like, the visual component.
A Brief History of the Open Source AI Hacker - with Ben Firshman of Replicate
- ▶ 1:10:12 Ben Firshman Um, and you can't do that if it's just, like, a hosted API, because it's just, like, it's, you know, it's, it's, you know, you can't, you can't touch the code.