Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q So, so Taubench, uh, I actually didn't know it was from Sierra, uh, but it's, it's actually very, very, uh, influential. Yeah, can you talk more about Taubench?
A Oh yes. So essentially tower bench when it was released, uh, I mean, yes, the top bench is for two, uh, two domains, the test agent capabilities for two domains, which is airline and retail. And the setup is that they have very long prompts in the beginning with lot of requests. And then you have these tools defined and the tools can interact with certain Databases for read and write operations, and it is predefined that once the whole agent tick task is done, then they will validate on the final, uh, edits made on, and they will check for the expected value in the data sets, uh, in the database, which are basically CSV style of files. Uh, and in that fashion, they will validate whether the task was done or not. So it's, it's, it's pretty well defined, such that You get a deterministic evaluation all the times. That's the beauty of this. And it's, it's domain specific, which I think is missing from other evaluation data sets. Uh, and that is also one of the themes for our V two, um, because I think people don't just come from a general mindset, right? People in the industry come from, let's say I'm from healthcare. I want to see if my, if this models work for healthcare or not, people are coming from investment, finance, insurance industry, and they want to know. That is, are the models ready for my domain or not? Because we all know that things work on the general domain side,…
AI assessment note: “top bench is for two, uh, two domains, the test agent capabilities”
Answered raw tape
D 5 · C 4 · P 4 · Cm 4 4.30
Q Um, what are some of the surprising results? For me, it was Mistral Small being the best open source model above Foro and DeepSeq VIII. Um, anything else that jumped out to you?
A Oh yes. So first and foremost, I would say that I was expecting GPTs to perform the best because they kind of pioneered the tool calling, function calling, support, and they have been working on that for so long. Um, there was some criticism around Gemini models, although they were catching up and I was very surprised to see when we released this leaderboard, of course there has And some shaking of the leaderboard also later on, but when we released, we found that Gemini models are performing really good, and they were extremely cost-efficient at the time compared to the second and second model like GPT-IV at the time, and so it became kind of a no-brainer that if anybody wants top performance, they can just directly go and go with the Gemini models, and they just not only released like the pro model, they also released the flash version later on and flashlight, But the flashlight is a very interesting model. Uh, in, in the blog, we have the models when we released this leaderboard, but in the, on the original leaderboard, we also have flashlight, which is, nobody talks about, I think, even now, but I think it's got a very decent score, which is, like, in the top five models. It's extremely cheap model, like, it's so dirt cheap that it makes a lot of sense. Only model that is cheaper than it is, and performance is from Amazon. Uh, which also nobody talks about. So we found this…
AI assessment note: “Then another surprise for me was that the reasoning models were not performing well enough.”