Aug 4, 2025 · 22m · latent-space

⚡️Composio: 10,000+ tools that evolve for Agents — Karan Vaidya and Soham Ganatra

Soham Ganatra · 8m spoken Karan Vaidya · 7m spoken Alessio Fanelli · 3m spoken Shawn Wang · 1m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Composio co-founders Karan Vaidya and Soham Ganatra join the Latent Space podcast to discuss building self-evolving agent skills, scaling tool maintenance autonomously, and leveraging advanced Model Context Protocol (MCP) architectures.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 21.5% of the talking time here. How this is scored →

The hosts as informed peer 5.0 Guest teaching 6.0 Guest disagreement 1.5 The hosts pushing back 1.5
05100:0010:0020:001:23–5:26 · The hosts as informed peer 5/10 Analyzing MCP Capabilities and Protocol Limitations for Clients Alessio prompts the guests on protocol ergonomics and MCP's trade-offs. Soham provides a deep technical breakdown of client-side limitations in MCP, such as rigid tool schemas and lack of output post-processing.5:29–10:13 · The hosts as informed peer 5/10 Agent Evaluation Strategies and MCP-Driven Human Feedback Swyx and Alessio inquire about agent evals and maintenance pipelines. Soham and Karan explain that standard tool-calling evals fail in production and reveal how they autonomously maintain 95% of integrations across 3 million lines of code.10:15–16:18 · The hosts as informed peer 4/10 Managing Tool Limits and Natural Language Execution Alessio and Swyx ask about handling large tool schemas and benchmarking. Soham educates the hosts on how Composio uses MCP's dynamic tool exposure for A/B testing heuristics, leaving Swyx to admit he did not know MCP supported that level of dynamicity.16:19–19:27 · The hosts as informed peer 6/10 MCP Sampling, Event Triggers, and Meta-Tool Architecture Alessio highlights the emergence of MCP sampling and server-side LLM inference from recent conferences. Karan and Soham validate the observation and describe their Agent API and meta-tool vision where an agent passes a single query rather than orchestrating multiple tools.1:23–5:26 · Guest teaching 6/10 Analyzing MCP Capabilities and Protocol Limitations for Clients Alessio prompts the guests on protocol ergonomics and MCP's trade-offs. Soham provides a deep technical breakdown of client-side limitations in MCP, such as rigid tool schemas and lack of output post-processing.5:29–10:13 · Guest teaching 6/10 Agent Evaluation Strategies and MCP-Driven Human Feedback Swyx and Alessio inquire about agent evals and maintenance pipelines. Soham and Karan explain that standard tool-calling evals fail in production and reveal how they autonomously maintain 95% of integrations across 3 million lines of code.10:15–16:18 · Guest teaching 7/10 Managing Tool Limits and Natural Language Execution Alessio and Swyx ask about handling large tool schemas and benchmarking. Soham educates the hosts on how Composio uses MCP's dynamic tool exposure for A/B testing heuristics, leaving Swyx to admit he did not know MCP supported that level of dynamicity.16:19–19:27 · Guest teaching 5/10 MCP Sampling, Event Triggers, and Meta-Tool Architecture Alessio highlights the emergence of MCP sampling and server-side LLM inference from recent conferences. Karan and Soham validate the observation and describe their Agent API and meta-tool vision where an agent passes a single query rather than orchestrating multiple tools.1:23–5:26 · Guest disagreement 2/10 Analyzing MCP Capabilities and Protocol Limitations for Clients Alessio prompts the guests on protocol ergonomics and MCP's trade-offs. Soham provides a deep technical breakdown of client-side limitations in MCP, such as rigid tool schemas and lack of output post-processing.5:29–10:13 · Guest disagreement 2/10 Agent Evaluation Strategies and MCP-Driven Human Feedback Swyx and Alessio inquire about agent evals and maintenance pipelines. Soham and Karan explain that standard tool-calling evals fail in production and reveal how they autonomously maintain 95% of integrations across 3 million lines of code.10:15–16:18 · Guest disagreement 1/10 Managing Tool Limits and Natural Language Execution Alessio and Swyx ask about handling large tool schemas and benchmarking. Soham educates the hosts on how Composio uses MCP's dynamic tool exposure for A/B testing heuristics, leaving Swyx to admit he did not know MCP supported that level of dynamicity.16:19–19:27 · Guest disagreement 1/10 MCP Sampling, Event Triggers, and Meta-Tool Architecture Alessio highlights the emergence of MCP sampling and server-side LLM inference from recent conferences. Karan and Soham validate the observation and describe their Agent API and meta-tool vision where an agent passes a single query rather than orchestrating multiple tools.1:23–5:26 · The hosts pushing back 1/10 Analyzing MCP Capabilities and Protocol Limitations for Clients Alessio prompts the guests on protocol ergonomics and MCP's trade-offs. Soham provides a deep technical breakdown of client-side limitations in MCP, such as rigid tool schemas and lack of output post-processing.5:29–10:13 · The hosts pushing back 2/10 Agent Evaluation Strategies and MCP-Driven Human Feedback Swyx and Alessio inquire about agent evals and maintenance pipelines. Soham and Karan explain that standard tool-calling evals fail in production and reveal how they autonomously maintain 95% of integrations across 3 million lines of code.10:15–16:18 · The hosts pushing back 1/10 Managing Tool Limits and Natural Language Execution Alessio and Swyx ask about handling large tool schemas and benchmarking. Soham educates the hosts on how Composio uses MCP's dynamic tool exposure for A/B testing heuristics, leaving Swyx to admit he did not know MCP supported that level of dynamicity.16:19–19:27 · The hosts pushing back 2/10 MCP Sampling, Event Triggers, and Meta-Tool Architecture Alessio highlights the emergence of MCP sampling and server-side LLM inference from recent conferences. Karan and Soham validate the observation and describe their Agent API and meta-tool vision where an agent passes a single query rather than orchestrating multiple tools.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 25.8% · guest 74.2%0:00 · the hosts 25.8% · guest 74.2%3:00 · the hosts 8.7% · guest 91.3%3:00 · the hosts 8.7% · guest 91.3%6:00 · the hosts 7.7% · guest 92.3%6:00 · the hosts 7.7% · guest 92.3%9:00 · the hosts 38.1% · guest 61.9%9:00 · the hosts 38.1% · guest 61.9%12:00 · the hosts 8.5% · guest 91.5%12:00 · the hosts 8.5% · guest 91.5%15:00 · the hosts 29.9% · guest 70.1%15:00 · the hosts 29.9% · guest 70.1%18:00 · the hosts 30.4% · guest 69.6%18:00 · the hosts 30.4% · guest 69.6%21:00 · the hosts 26.3% · guest 73.7%21:00 · the hosts 26.3% · guest 73.7%
Sharpest disagreement ▶ 8:07 Soham criticizes enterprise tool APIs and documentation

Soham forcefully vents about the 5% of integrations that cannot be autonomously maintained due to terrible documentation and lack of developer sandbox environments in legacy enterprise software.

Hardest push from the hosts ▶ 8:00 Alessio presses on the 5% unmaintained integrations

Alessio immediately challenges the assertion that agents maintain almost everything by asking what constitutes the 5% failure boundary.

Biggest teaching moment ▶ 15:04 Soham teaches Swyx about dynamic tool exposure in MCP

Soham explains how MCP servers can dynamically mutate and expose tools per query, prompting Swyx to openly admit he was unaware the protocol supported that capability.

The host holds their own ▶ 16:19 Alessio brings up MCP sampling and value-chain capture

Alessio demonstrates sharp domain expertise by bringing up the newly introduced MCP sampling spec and analyzing how server-side LLM inference shifts value capture in agent workflows.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Analyzing MCP Capabilities and Protocol Limitations for Clients 5621 Alessio prompts the guests on protocol ergonomics and MCP's trade-offs. Soham provides a deep technical breakdown of client-side limitations in MCP, such as rigid tool schemas and lack of output post-processing.
Agent Evaluation Strategies and MCP-Driven Human Feedback 5622 Swyx and Alessio inquire about agent evals and maintenance pipelines. Soham and Karan explain that standard tool-calling evals fail in production and reveal how they autonomously maintain 95% of integrations across 3 million lines of code.
Managing Tool Limits and Natural Language Execution 4711 Alessio and Swyx ask about handling large tool schemas and benchmarking. Soham educates the hosts on how Composio uses MCP's dynamic tool exposure for A/B testing heuristics, leaving Swyx to admit he did not know MCP supported that level of dynamicity.
MCP Sampling, Event Triggers, and Meta-Tool Architecture 6512 Alessio highlights the emergence of MCP sampling and server-side LLM inference from recent conferences. Karan and Soham validate the observation and describe their Agent API and meta-tool vision where an agent passes a single query rather than orchestrating multiple tools.

Statements from this episode (11)

Disclosure
Vaidya: Composio evolves agent skills by analyzing usage patterns
“We are kind of self evolving skills now for agents. So people building or using agents want to like basically make Agents interact with their apps. They can use Composio. We manage all the authentication and user account related stuff for them so that they don…”
Karan Vaidya Aug 4, 2025 ▶ 0:36
Assertion Supported
Ganatra: MCP Lacks Client Post-Processing for Long Context Responses
“MCP let's just say you have a tool that has really long context, a really long sort of response. Currently, if you use MCP for those kinds of tools, it ends up sort of messing up, right? Because let's say I'm trying to fetch like a hundred emails on Gmail and …”
Soham Ganatra Aug 4, 2025 ▶ 4:04
Assertion Supported
Ganatra: Server-Controlled Schemas in MCP Complicate Client-Side Customization
“Some of the other things that we have observed is like, hey, maybe I want to sort of change the tool descriptions, but like in MCP, the tool descriptions are controlled by the server itself. And so essentially for me as a developer who's sort of using an MCP, …”
Soham Ganatra Aug 4, 2025 ▶ 4:37
Disclosure
Karan Vaidya: Composio runs an internal agent to build and evaluate skills
“We have an agent essentially, which builds a lot of these skills and kind of like does a lot of like testing as well as eval. So it's like, you can say like agent as an eval where our internal agent uses these tools and kind of like we figured out where it is …”
Karan Vaidya Aug 4, 2025 ▶ 5:38
Opinion
Ganatra: Existing AI tool-calling benchmarks are not production-ready
“Tool calling has a lot of different evals. You have seen the function tool calling evals. You have multi-function, like, multi-tune tool calling evals. But, like, none of them are actual, like, production quality. Like, none of them are actually something that…”
Soham Ganatra Aug 4, 2025 ▶ 6:55
Assertion Not checkable as stated
Ganatra: 95% of Composio integrations are built and maintained by agents
“So, 95% of all our integrations are completely built using agents, maintained using agents.”
Soham Ganatra Aug 4, 2025 ▶ 7:40
Assertion Supported
Ganatra: Composio supports more than 500 applications
“So we have at this point more than 500 different applications on top of Composio.”
Soham Ganatra Aug 4, 2025 ▶ 9:25
Assertion Not checkable as stated
Ganatra: Over 100,000 developers have built with Composio
“In terms of developers, we have, like, more than 100,000 developers who have been building on top of course and, like, trying our product out.”
Soham Ganatra Aug 4, 2025 ▶ 10:02
Insight
Vaidya: LLM agents get confused when exposed to over 20 tool actions
“More than like 20, 25 actions like just confuses the server. Like it's not able to kind of like figure out which tool to use. And same goes with like, if the schema of the tools are really complex, that also confuses the agent.”
Karan Vaidya Aug 4, 2025 ▶ 10:46
Assertion Not checkable as stated
Ganatra: Poor API documentation is the main source of tool-calling errors
“Is this Salesforce tool getting transformed into the right request payload so that it is hitting the right endpoint in Salesforce, right? And that is like the major source of error today, because a lot of documentations are not great. A lot of like ad cases ar…”
Soham Ganatra Aug 4, 2025 ▶ 13:29
Assertion Not checkable as stated
Ganatra: Composio's meta-tool became highly used after a year of zero adoption
“We built it, like, a year back. Nobody used it for a very long time. Today, it's probably one of the most used tools.”
Soham Ganatra Aug 4, 2025 ▶ 18:30
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.