Q So maybe, uh, you mentioned, um, OpenAI Ada. Maybe, like, paint a quick, uh, picture of the space. Like, what are the other, uh, you know, embedding models that, uh, either have existed for a while or are just coming up?
A Yeah, um, so, uh, roughly in the space, uh, there, in the open source side of things, um, there's a bunch of really great models, uh, that have been released, uh, the model weights have been released, and A few of them are being hosted, uh, on, you know, embedding endpoints, um, but the problem with a lot of these models is that they have a small context length, so they're limited to 512 tokens, um, which I think, you know, nets out to maybe a few paragraphs, um, which becomes a problem when you have large, uh, large documents. Maybe you have, like, long financial documents and you want to reason over the whole document versus chunks, um, just because it gets a little confusing and it's, Uh, a little hard to, um, reason about, like, the right way to chunk up a large document if you need to, like, recall something from the first paragraph and the last paragraph. Um, so that's the, the open source side of things. Um, there's, uh, one or two long context models. Um, uh, the ones that are, are good are either too big, they're in, like, the seven billion parameter range, um, or they, uh, don't actually beat OpenAI, uh, Ada's model. Um, and then, you know, for the long contacts on the closed source, uh, uh, ADA is, you know, kind of the de facto for many people's applications today. Um, but, you know, with, with closed source, the, the big problem is, you know, you, you have no idea …
AI assessment note: “in the open source side of things, um, there's a bunch of really great models”