Feb 5, 2025 · 31m · latent-space

Why every AI Engineer needs an AI Gateway (ft Portkey.ai CEO)

Rohit Agarwal · 18m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Portkey.ai CEO Rohit Agarwal joins hosts Swix and Alessio to discuss the critical role of AI gateways in managing model routing, production guardrails, OpenTelemetry observability, and multi-step agentic workflows.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 5.9 Guest teaching 3.0 Guest disagreement 1.2 The hosts pushing back 3.7
05100:0010:0020:0030:001:39–3:44 · The hosts as informed peer 7/10 Model Routing for Reasoning LLMs and Agentic Systems Alessio demonstrates strong domain knowledge by challenging the necessity of routing in single-model setups and sharing concrete insights from a Fortune 500 AI council regarding reasoning model latencies. Rohit readily validates Alessio's framing and builds on it with recent DeepSeek data.3:45–7:38 · The hosts as informed peer 6/10 Decoupling Gateway Architecture from Application SDKs Swix expresses strong architectural opinions against putting LLM abstractions in application code while raising valid concerns about proxy latency and downtime. Rohit clarifies his architecture by explaining their 21MB JSON transformer engine, which Swix clarifies isn't an LLM transformer.7:38–10:52 · The hosts as informed peer 4/10 Core Gateway Capabilities: Observability, Limits, and Governance Rohit provides a structured overview of core gateway functions including rate limits, budget enforcement, guardrails, and enterprise chargebacks. The hosts act primarily as facilitators guiding the categorization.10:53–14:39 · The hosts as informed peer 7/10 Production Guardrail Implementation and Error Orchestration Alessio opens with historical context on guardrail evolution, and Rohit explains practical regex versus demo PII guardrails and custom status code 246. Swix directly pushes back on gateway-level dumb retries, arguing that retry prompts require application-level context and tool awareness.14:39–17:57 · The hosts as informed peer 6/10 AI Observability and the Evolution of OpenTelemetry Standards Swix probes Portkey's positioning against traditional observability vendors and standard bodies like OpenLLMetry. Both Swix and Rohit note that standardization efforts often lag behind rapidly shifting API paradigms.17:57–20:33 · The hosts as informed peer 6/10 Observability and Debugging in Multi-Step Agent Architectures Rohit details the debugging requirements and routing opportunities in multi-step agents. Swix counters with skepticism regarding multi-LLM setups, arguing that switching models complicates prompts and accumulates endpoint downtime risks.20:33–23:48 · The hosts as informed peer 5/10 Real-Time Production Evals and Aligning Human Feedback Swix candidly admits that developers and end users rarely provide manual human feedback. Rohit agrees and shares successful alternative patterns where implicit business metrics (like video download rates) serve as positive feedback signals.23:48–26:44 · The hosts as informed peer 7/10 Defining Traces and Sessions in Ambient and Continuous AI Rohit explains the OpenTelemetry model of traces versus sessions, but Swix challenges the paradigm by introducing continuous wearable audio streams and ambient agents where clear session boundaries do not exist.26:44–29:47 · The hosts as informed peer 5/10 Model Context Protocol (MCP) and Gateway Tool Integration Swix questions why MCP is seen as revolutionary compared to OpenAPI specs. Rohit educates Swix on two-way communication capabilities and RPC mechanisms, leading Swix to acknowledge he did not realize two-way agent functionality existed in the spec.1:39–3:44 · Guest teaching 1/10 Model Routing for Reasoning LLMs and Agentic Systems Alessio demonstrates strong domain knowledge by challenging the necessity of routing in single-model setups and sharing concrete insights from a Fortune 500 AI council regarding reasoning model latencies. Rohit readily validates Alessio's framing and builds on it with recent DeepSeek data.3:45–7:38 · Guest teaching 4/10 Decoupling Gateway Architecture from Application SDKs Swix expresses strong architectural opinions against putting LLM abstractions in application code while raising valid concerns about proxy latency and downtime. Rohit clarifies his architecture by explaining their 21MB JSON transformer engine, which Swix clarifies isn't an LLM transformer.7:38–10:52 · Guest teaching 3/10 Core Gateway Capabilities: Observability, Limits, and Governance Rohit provides a structured overview of core gateway functions including rate limits, budget enforcement, guardrails, and enterprise chargebacks. The hosts act primarily as facilitators guiding the categorization.10:53–14:39 · Guest teaching 2/10 Production Guardrail Implementation and Error Orchestration Alessio opens with historical context on guardrail evolution, and Rohit explains practical regex versus demo PII guardrails and custom status code 246. Swix directly pushes back on gateway-level dumb retries, arguing that retry prompts require application-level context and tool awareness.14:39–17:57 · Guest teaching 2/10 AI Observability and the Evolution of OpenTelemetry Standards Swix probes Portkey's positioning against traditional observability vendors and standard bodies like OpenLLMetry. Both Swix and Rohit note that standardization efforts often lag behind rapidly shifting API paradigms.17:57–20:33 · Guest teaching 2/10 Observability and Debugging in Multi-Step Agent Architectures Rohit details the debugging requirements and routing opportunities in multi-step agents. Swix counters with skepticism regarding multi-LLM setups, arguing that switching models complicates prompts and accumulates endpoint downtime risks.20:33–23:48 · Guest teaching 3/10 Real-Time Production Evals and Aligning Human Feedback Swix candidly admits that developers and end users rarely provide manual human feedback. Rohit agrees and shares successful alternative patterns where implicit business metrics (like video download rates) serve as positive feedback signals.23:48–26:44 · Guest teaching 3/10 Defining Traces and Sessions in Ambient and Continuous AI Rohit explains the OpenTelemetry model of traces versus sessions, but Swix challenges the paradigm by introducing continuous wearable audio streams and ambient agents where clear session boundaries do not exist.26:44–29:47 · Guest teaching 7/10 Model Context Protocol (MCP) and Gateway Tool Integration Swix questions why MCP is seen as revolutionary compared to OpenAPI specs. Rohit educates Swix on two-way communication capabilities and RPC mechanisms, leading Swix to acknowledge he did not realize two-way agent functionality existed in the spec.1:39–3:44 · Guest disagreement 1/10 Model Routing for Reasoning LLMs and Agentic Systems Alessio demonstrates strong domain knowledge by challenging the necessity of routing in single-model setups and sharing concrete insights from a Fortune 500 AI council regarding reasoning model latencies. Rohit readily validates Alessio's framing and builds on it with recent DeepSeek data.3:45–7:38 · Guest disagreement 1/10 Decoupling Gateway Architecture from Application SDKs Swix expresses strong architectural opinions against putting LLM abstractions in application code while raising valid concerns about proxy latency and downtime. Rohit clarifies his architecture by explaining their 21MB JSON transformer engine, which Swix clarifies isn't an LLM transformer.7:38–10:52 · Guest disagreement 0/10 Core Gateway Capabilities: Observability, Limits, and Governance Rohit provides a structured overview of core gateway functions including rate limits, budget enforcement, guardrails, and enterprise chargebacks. The hosts act primarily as facilitators guiding the categorization.10:53–14:39 · Guest disagreement 2/10 Production Guardrail Implementation and Error Orchestration Alessio opens with historical context on guardrail evolution, and Rohit explains practical regex versus demo PII guardrails and custom status code 246. Swix directly pushes back on gateway-level dumb retries, arguing that retry prompts require application-level context and tool awareness.14:39–17:57 · Guest disagreement 1/10 AI Observability and the Evolution of OpenTelemetry Standards Swix probes Portkey's positioning against traditional observability vendors and standard bodies like OpenLLMetry. Both Swix and Rohit note that standardization efforts often lag behind rapidly shifting API paradigms.17:57–20:33 · Guest disagreement 2/10 Observability and Debugging in Multi-Step Agent Architectures Rohit details the debugging requirements and routing opportunities in multi-step agents. Swix counters with skepticism regarding multi-LLM setups, arguing that switching models complicates prompts and accumulates endpoint downtime risks.20:33–23:48 · Guest disagreement 1/10 Real-Time Production Evals and Aligning Human Feedback Swix candidly admits that developers and end users rarely provide manual human feedback. Rohit agrees and shares successful alternative patterns where implicit business metrics (like video download rates) serve as positive feedback signals.23:48–26:44 · Guest disagreement 1/10 Defining Traces and Sessions in Ambient and Continuous AI Rohit explains the OpenTelemetry model of traces versus sessions, but Swix challenges the paradigm by introducing continuous wearable audio streams and ambient agents where clear session boundaries do not exist.26:44–29:47 · Guest disagreement 2/10 Model Context Protocol (MCP) and Gateway Tool Integration Swix questions why MCP is seen as revolutionary compared to OpenAPI specs. Rohit educates Swix on two-way communication capabilities and RPC mechanisms, leading Swix to acknowledge he did not realize two-way agent functionality existed in the spec.1:39–3:44 · The hosts pushing back 4/10 Model Routing for Reasoning LLMs and Agentic Systems Alessio demonstrates strong domain knowledge by challenging the necessity of routing in single-model setups and sharing concrete insights from a Fortune 500 AI council regarding reasoning model latencies. Rohit readily validates Alessio's framing and builds on it with recent DeepSeek data.3:45–7:38 · The hosts pushing back 3/10 Decoupling Gateway Architecture from Application SDKs Swix expresses strong architectural opinions against putting LLM abstractions in application code while raising valid concerns about proxy latency and downtime. Rohit clarifies his architecture by explaining their 21MB JSON transformer engine, which Swix clarifies isn't an LLM transformer.7:38–10:52 · The hosts pushing back 1/10 Core Gateway Capabilities: Observability, Limits, and Governance Rohit provides a structured overview of core gateway functions including rate limits, budget enforcement, guardrails, and enterprise chargebacks. The hosts act primarily as facilitators guiding the categorization.10:53–14:39 · The hosts pushing back 6/10 Production Guardrail Implementation and Error Orchestration Alessio opens with historical context on guardrail evolution, and Rohit explains practical regex versus demo PII guardrails and custom status code 246. Swix directly pushes back on gateway-level dumb retries, arguing that retry prompts require application-level context and tool awareness.14:39–17:57 · The hosts pushing back 4/10 AI Observability and the Evolution of OpenTelemetry Standards Swix probes Portkey's positioning against traditional observability vendors and standard bodies like OpenLLMetry. Both Swix and Rohit note that standardization efforts often lag behind rapidly shifting API paradigms.17:57–20:33 · The hosts pushing back 5/10 Observability and Debugging in Multi-Step Agent Architectures Rohit details the debugging requirements and routing opportunities in multi-step agents. Swix counters with skepticism regarding multi-LLM setups, arguing that switching models complicates prompts and accumulates endpoint downtime risks.20:33–23:48 · The hosts pushing back 2/10 Real-Time Production Evals and Aligning Human Feedback Swix candidly admits that developers and end users rarely provide manual human feedback. Rohit agrees and shares successful alternative patterns where implicit business metrics (like video download rates) serve as positive feedback signals.23:48–26:44 · The hosts pushing back 3/10 Defining Traces and Sessions in Ambient and Continuous AI Rohit explains the OpenTelemetry model of traces versus sessions, but Swix challenges the paradigm by introducing continuous wearable audio streams and ambient agents where clear session boundaries do not exist.26:44–29:47 · The hosts pushing back 5/10 Model Context Protocol (MCP) and Gateway Tool Integration Swix questions why MCP is seen as revolutionary compared to OpenAPI specs. Rohit educates Swix on two-way communication capabilities and RPC mechanisms, leading Swix to acknowledge he did not realize two-way agent functionality existed in the spec.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 14:08 Contesting automated gateway retries

Swix directly challenges Portkey's custom status code retry mechanism, arguing that dumb retries fail without the application's unique prompt context and tools.

Hardest push from the hosts ▶ 28:06 Challenging the novelty of MCP

Swix pushes back on the hype around MCP, questioning why it is revolutionary rather than just a minor variation of the existing OpenAPI specification.

Biggest teaching moment ▶ 29:07 Two-way communication in MCP spec

Rohit explains the bidirectional RPC and SSE capabilities of MCP, prompting Swix to openly admit he had missed that entire capability despite browsing the specification.

The host holds their own ▶ 2:12 Reasoning model latency driving routing needs

Alessio demonstrates industry authority by bringing fresh findings from a Fortune 500 enterprise council to illustrate exactly where routing is becoming critical due to reasoning model latency.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Model Routing for Reasoning LLMs and Agentic Systems 7114 Alessio demonstrates strong domain knowledge by challenging the necessity of routing in single-model setups and sharing concrete insights from a Fortune 500 AI council regarding reasoning model latencies. Rohit readily validates Alessio's framing and builds on it with recent DeepSeek data.
Decoupling Gateway Architecture from Application SDKs 6413 Swix expresses strong architectural opinions against putting LLM abstractions in application code while raising valid concerns about proxy latency and downtime. Rohit clarifies his architecture by explaining their 21MB JSON transformer engine, which Swix clarifies isn't an LLM transformer.
Core Gateway Capabilities: Observability, Limits, and Governance 4301 Rohit provides a structured overview of core gateway functions including rate limits, budget enforcement, guardrails, and enterprise chargebacks. The hosts act primarily as facilitators guiding the categorization.
Production Guardrail Implementation and Error Orchestration 7226 Alessio opens with historical context on guardrail evolution, and Rohit explains practical regex versus demo PII guardrails and custom status code 246. Swix directly pushes back on gateway-level dumb retries, arguing that retry prompts require application-level context and tool awareness.
AI Observability and the Evolution of OpenTelemetry Standards 6214 Swix probes Portkey's positioning against traditional observability vendors and standard bodies like OpenLLMetry. Both Swix and Rohit note that standardization efforts often lag behind rapidly shifting API paradigms.
Observability and Debugging in Multi-Step Agent Architectures 6225 Rohit details the debugging requirements and routing opportunities in multi-step agents. Swix counters with skepticism regarding multi-LLM setups, arguing that switching models complicates prompts and accumulates endpoint downtime risks.
Real-Time Production Evals and Aligning Human Feedback 5312 Swix candidly admits that developers and end users rarely provide manual human feedback. Rohit agrees and shares successful alternative patterns where implicit business metrics (like video download rates) serve as positive feedback signals.
Defining Traces and Sessions in Ambient and Continuous AI 7313 Rohit explains the OpenTelemetry model of traces versus sessions, but Swix challenges the paradigm by introducing continuous wearable audio streams and ambient agents where clear session boundaries do not exist.
Model Context Protocol (MCP) and Gateway Tool Integration 5725 Swix questions why MCP is seen as revolutionary compared to OpenAPI specs. Rohit educates Swix on two-way communication capabilities and RPC mechanisms, leading Swix to acknowledge he did not realize two-way agent functionality existed in the spec.

Statements from this episode (15)

Insight
Agarwal: AI gateways act as operational platforms offering governance beyond basic proxies
“So I think AI gateways are essentially the operational platforms that enable teams to connect to LLMs more efficiently. They help you improve cost, performance, and accuracy by not having you to build individual connections to all of these different AI service…”
Rohit Agarwal Feb 5, 2025 ▶ 1:04
Assertion Not checkable as stated
Agarwal: 90% of production LLM use cases do not use automatic routing
“In fact, I would say in production, 90% of the use cases do not use automatic routing. What they want is deterministic flows. As long as the gateway manages authentication authorization for them, it's perfectly fine. The request hitting A specific model that t…”
Rohit Agarwal Feb 5, 2025 ▶ 1:53
Insight
Agarwal: Multi-step agentic systems are driving adoption of LLM routing
“With the rise of agentic systems where people are building multi-step agents, that's also where routing is becoming really popular because you might want to have different tasks point to different LLMs, and then you don't want to build authorization and all of…”
Rohit Agarwal Feb 5, 2025 ▶ 3:14
Assertion Partly supported
Agarwal: Portkey open source gateway has a 21 MB memory footprint
“So the gateway, when it runs, has a memory footprint of like, 21 MB, and the way we've been able to do it is just using a very interesting architecture where we're using transformers in JSON to do everything out of the box.”
Rohit Agarwal Feb 5, 2025 ▶ 6:14
Insight
Agarwal: Common AI guardrails focus on format and topics over PII
“When I talk about guardrails, the first thing that comes to mind is, oh yeah, we have to do PII redaction and sensitive data redaction. But somehow we've seen that the more common use cases for guardrails end up being, you know, is it the right length? Is it u…”
Rohit Agarwal Feb 5, 2025 ▶ 9:19
Assertion Not checkable as stated
Agarwal: Gateway-based guardrails reduce latency versus application-level checks
“And again, you can decrease latency massively by building this on the gateway. Rather than building this within your application.”
Rohit Agarwal Feb 5, 2025 ▶ 9:44
Assertion Not checkable as stated
Agarwal: Regex rules are the most deployed LLM guardrails in production
“In reality, we've seen what gets deployed the most is regex-based guardrails to say, you know, I want to catch for specific words, and then take certain actions based on that. Or I want to catch for empty outputs.”
Rohit Agarwal Feb 5, 2025 ▶ 12:03
Insight
Agarwal: Gateways produce the best LLM metrics without manual codebase instrumentation
“We're saying that the best LLM metrics can be produced directly on the gateway. Without every team having to instrument their own code bases to pick out these metrics.”
Rohit Agarwal Feb 5, 2025 ▶ 15:01
Insight
Agarwal: AI telemetry cannot standardize while underlying APIs change monthly
“Standards usually evolve when there's some amount of coherence and stability in an API. I think the APIs themselves change every month. So I'm not sure how there's going to be standardization that occurs.”
Rohit Agarwal Feb 5, 2025 ▶ 17:28
Assertion Not checkable as stated
Agarwal: AI teams rarely use human feedback, despite high satisfaction when used
“I think it's not used as much. I agree. And I think that's something I tell every customer that you need to close the loop and feedback is going to help you. I think the teams that are doing it are really happy. They're not doing it.”
Rohit Agarwal Feb 5, 2025 ▶ 21:59
Prediction Not checkable as stated
Agarwal: LLM teams will adopt human feedback only after stabilizing applications
“I think the output side of things will come. Once we've reached stable state, where there are production applications that are stable, and then they want to start optimizing, is when they'll start looking at human feedback as well.”
Rohit Agarwal Feb 5, 2025 ▶ 22:26
Insight
Agarwal: Successful LLM feedback loops map directly to business metrics
“I think the teams that have been successful is when they're mapping these feedback metrics to a business metric that they track.”
Rohit Agarwal Feb 5, 2025 ▶ 22:53
Prediction Not checkable as stated
Agarwal: Ambient AI agents will break traditional fixed-boundary telemetry traces
“I would imagine like with more ambient agents coming in, there might not be a clear start and a clear end. So, it'll be interesting to see how we start solving this problem. Maybe the whole concept of, you know, there's a typical start and a typical end goes a…”
Rohit Agarwal Feb 5, 2025 ▶ 26:18
Prediction Not checkable as stated
Agarwal: MCP will become standard way agents connect to services
“I think the future of agents, the way they connect to different services is going to be MCP.”
Rohit Agarwal Feb 5, 2025 ▶ 26:58
Prediction Not checkable as stated
Agarwal: AI production will require operational tools just like DevOps enabled cloud
“The more they want platforms like this, and I feel this is similar to what happened with the cloud era back in 2012. You will not have cloud adoption till the time DevOps did not take off, and companies like Datadog, Cloudflare, et cetera, didn't help you buil…”
Rohit Agarwal Feb 5, 2025 ▶ 30:03
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.