The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Karan Vaidya no published score: only 1 usable exchange on raw tape, and a fair score needs 8+ record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
1exchanges match
1on raw tape
0redirected or not addressed
Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q a lot of talk about how many tools an LLM can support, whether or not, you know, you should give different tools, whether or not you should put all the tools in one MCP, which you guys also offer with the Composio MCP. Can you maybe give people Kind of state of the art today. Where do things go wrong? Uh, and maybe the slope of improvement of the models.

A Sure. So I think, uh, like we have seen over like a lot of practice that like more than like 20, 25, uh, actions like just confuses the server. Like it's not able to kind of like figure out which tool to use. And same goes with like, if the schema of the tools are really complex, that also confuses the agent. So we try to keep our kind of like schema simple. Like some of the things are like flattening the schemas, et cetera. In addition, I think, uh, Like basically I think it's like agents, like because they are LLMs inherently, so they like to think in natural language. So some of the things that we are planning on that front is like a single natural language execution of a tool so that like the agent does just gives like, like agent has just given a single tool and it can call any of the other kind of skills that we have like inside Composio with just natural language. So like create a notion doc and send a meeting invite just in natural language and we'll do that in the backend.

AI assessment note: “more than like 20, 25, actions like just confuses the server”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.