← Back to Insights
Guide Series AI gateway and do you need one
Integration & Security

What Is an AI Gateway, and When Do You Need One

A guide to policy enforcement for enterprise model traffic
Integration & Security 12 min read September 30, 2026 Duczer East Insights
Part of the Insights topic: Agentic AI Governance & Model Risk →

This guide is for the architect or platform lead who has been asked whether the enterprise needs an AI gateway, and for the compliance or risk leader who will be asked to sign off on one. It sets out what an AI gateway does that an API gateway does not, where it sits in a regulated architecture, which products occupy the space and how they differ, what fails in production, and when the honest answer is that you do not need one yet. It is written for people who already run an API gateway and already call language models, and who now have to decide whether those two facts require a third thing.

An AI gateway is a policy enforcement point between applications and the language models they call. It authenticates the caller, applies limits on tokens and cost, inspects prompts and responses, routes requests to the right model, and records what was sent and returned. It does for model traffic what an API gateway does for API traffic, with controls that only make sense when the payload is a prompt and the bill is measured in tokens. The rest of this guide is about the decisions behind that sentence.

Eight Things That Change

The first question a CIO asks is reasonable: the enterprise already owns an API gateway, so why is this not another route on it? The answer is that eight things change when the thing being called is a model rather than a service. A product that does none of them is an API gateway with a new label.

Token accounting and budgets come first. An API call costs roughly the same whether the payload is ten bytes or ten kilobytes. A model call is billed by the token, in and out, and a long document or a verbose agent can cost a hundred times more than a short question. The gateway counts tokens per request, attributes them to an application, a team or a project, and enforces a budget before the model is called rather than after the invoice arrives. Prompt and response inspection is second. The gateway reads what goes in and what comes out. On the way in it can block or flag a prompt that matches a policy, such as an attempt to place a customer record into a prompt that should not contain one. On the way out it checks the response before the caller sees it.

Model routing and fallback is third. A request that needs a small, fast model should not go to a large, expensive one, and a request that carries regulated data should not leave the building. The gateway sends each request to the model that fits its cost, latency and data-handling requirement, and fails over when a provider is down or rate-limited. Semantic caching is fourth: two prompts close enough to mean the same thing can share an answer, which cuts cost and latency on repeated questions and introduces a class of bug covered later. Redaction on the way out is fifth. When the model is not inside your boundary, the gateway strips or masks personal and regulated data before the prompt leaves.

Tool-call authorization is sixth, and it is the one most vendor demonstrations skip. Agents do not only ask questions. They call tools: search a system, open a ticket, move money. The gateway decides which agent may invoke which tool, under whose authority, before the call is made. This is the control that separates agent traffic from chat traffic. Per-model rate limits are seventh, by model, by caller and by time window, so one runaway agent in a loop cannot exhaust a quota shared by forty applications. Audit of prompts and completions is eighth. The gateway records the prompt, the response, the model and version, the caller, and every policy decision, in a form a model-risk reviewer can read months later. Not a log line. A record.

The Response Path and Architecture Placement

The gateway sees every request twice, and the second pass is the one that gets dropped. On the way to the model it authenticates the caller, checks the budget, inspects the prompt, picks the model and forwards the call. On the way back it inspects the response, redacts what must not leave, records the exchange and returns the answer. A gateway that only governs the outbound half will return a response containing an account number the model reconstructed from context, and log nothing about it. The first thing to check in any gateway design is whether the response path has a policy on it at all. Draw the five stops on a whiteboard, application, gateway, model, gateway, application, and mark which ones run inside your boundary. That drawing is the design.

Where the AI gateway sits relative to the API gateway is a secondary question with three defensible answers. It can sit in front, so model traffic is governed before it reaches the estate, which fits when model traffic comes from many sources and the API gateway is not the natural chokepoint. It can sit behind, so the API gateway handles identity and coarse routing and hands model-bound traffic to an AI gateway for model policy, which fits when the API gateway is already the front door for everything. Or it can be the same product, with the API gateway vendor's AI policies applied as a second policy set on model routes. The third arrangement is the most common in practice, because the enterprise already owns an API gateway and every major vendor has shipped AI capabilities into it. The arrangement matters less than one question. Your model traffic is only as governed as the component that sees the prompt in the clear. Identify that component and where it runs, and the rest of the architecture follows.

Regulated Requirements and Policy Ownership

Vendor pages describe the gateway as a way to call several hosted providers under one key, with a dashboard. That is a real use, and it is not the use a bank or a hospital has. In a regulated enterprise the gateway is the control that lets you say, in writing, what a model was allowed to see, what it was allowed to do, and who authorized it. Three requirements change the design. The prompt must not leave the boundary unless a policy says it may, which means either the gateway runs inside the boundary and routes to in-perimeter models, or it redacts before egress under a written policy that someone has reviewed. The audit record must be complete enough for a model-risk reviewer to reconstruct a decision months later: prompt, response, model and version, caller, and every policy applied, retained as long as the decision it supported. A dashboard of request counts does not meet this; a queryable record does. And tool-call authorization must be enforced by identity. An agent that calls a core banking system does so as a named principal with scoped permissions, not through a shared key the gateway happens to hold. The gateway enforces the policy; the identity layer defines who the agent is and what it may touch, a design covered in AI agent governance on WSO2. Where the tools behind the gateway speak the Model Context Protocol, the injection surface moves into tool results, which is the the subject of an earlier piece on MCP and authorization.

There is a fourth requirement that is not technical. Someone has to own the policies. Gateway deployments stall far more often on unassigned routing rules, budgets and redaction patterns than on the software. Budget for that role before the licence.

Vendor Landscape and Common Failures

The landscape is a map, not a ranking, and it changes quarterly, so verify any capability with the vendor before a design depends on it. WSO2's AI Gateway covers token and cost tracking, guardrails including PII masking, semantic caching, multi-provider proxying and MCP tool exposure, and ships both as the Bijira SaaS and as self-hosted software (the deployment trade-off between the two is is set out in Bijira vs API Manager). Kong's AI plugins add add token-based rate limiting, semantic caching and load balancing across models to the Kong gateway, and are strongest where Kong already fronts the APIs. Apigee's AI gateway capabilities bring model routing, token quotas and Model Armor prompt and response sanitization to Google Cloud API management, and fit naturally in a Google Cloud estate. Cloudflare AI Gateway is edge-hosted caching, logging and rate limiting for calls to hosted models. Portkey is a dedicated LLM gateway focused on routing, observability and cost across providers. LiteLLM is an open-source proxy that presents many providers behind one interface, light and self-hostable, and a common first step. AWS, Azure and Google each offer model-traffic controls that work best inside their own platform and least well across platforms. If your compliance regime requires the control plane inside your boundary, the list shortens to the self-hostable options. If it can accept a hosted control plane with in-perimeter traffic, the list is longer and the decision moves to operational fit.

Four failures account for most gateway incidents, and none shows up in a demonstration. Streaming and timeouts: model responses stream for seconds or minutes, and a gateway tuned for millisecond API calls will cut them off, buffer them so the client sees nothing until the end, or time out the connection. Every timeout between application and model has to be raised deliberately, and streaming has to be tested through the gateway, not around it. Token accounting drift: the gateway counts tokens one way and the provider bills another, because tokenizers differ by model and cached or truncated responses are counted inconsistently. Budgets that look enforced are not, until the invoice shows the gap. Reconcile the gateway's count against the provider's bill monthly and treat a persistent gap as a defect. Cache invalidation: a semantic cache keyed on the user prompt keeps serving old answers after the system prompt or the model changes, so the cache key has to include the system prompt, the model and the version, or the cache has to be flushed on every change to any of them. Most teams learn this the day a policy change appears not to take effect. Retries multiplying spend: the client retries, the gateway retries, the provider SDK retries, and one failed call becomes nine billed calls while a provider outage becomes a spend spike. Retries belong in one place, and the gateway is that place, with the client and SDK retries turned off.

A proxy is a weekend and a gateway is a program. Routing and logging can be built in a few hundred lines, and many teams start there. The problem is what comes next. Token budgets that reconcile to the invoice, semantic caching that invalidates correctly, tool-call authorization tied to an identity provider, and an audit store a reviewer will accept are each a project of their own, and a team that builds them all ends up owning a product. If the enterprise already runs an API gateway with AI policies, extend it. If not, adopt an open-source gateway or buy one, and spend the engineering on the policies, because the policies are the part no vendor can write. They encode what your enterprise is willing to let a model see and do.

“Your model traffic is only as governed as the component that sees the prompt in the clear.”

The half nobody writes is when you do not need one. One application, one hosted model, one team, no regulated data in the prompt, and spend small enough that the provider's own dashboard is enough control. That describes most first projects, and a gateway would add latency and an operations burden for nothing. Add it when any of these arrives: the second team, the second provider, the first in-perimeter model, the first agent that calls a tool, or the first question from an auditor. At that point it stops being optional, and it is far easier to add before the fifth application than after. The migration cost is not the gateway. It is moving five applications' worth of keys, retries and error handling behind it.

An AI gateway is worth having when model traffic has more than one source, more than one destination, or more than one person who will be asked about it later. In a regulated enterprise that is nearly always, and the deciding question is not which product but which component sees the prompt in the clear and where it runs. Answer that, assign an owner to the policies, test the response path as carefully as the request path, and the product choice becomes the easy part.

Frequently asked questions

What is an AI gateway?

An AI gateway is infrastructure that intercepts every call from an application or agent to a language model and applies policy before the call goes out and again when the response comes back. The policies cover who may call which model, how many tokens they may spend, what the prompt may contain, what the response may reveal, and which tools an agent may invoke. Each exchange is recorded with the model version and the policy decisions, so a reviewer can reconstruct it later. It is a distinct layer from the model and from the application, which is what lets one set of rules apply across many of both.

What is the difference between an AI gateway and an API gateway?

An API gateway assumes the payload is structured data, the cost of a call is roughly fixed, and the risk lives in who is calling. An AI gateway assumes the payload is natural language, the cost is proportional to the length of the exchange, and the risk lives in the content itself: what the prompt exposes and what the response discloses. That difference produces the extra controls: token budgets, semantic inspection of both directions, routing by model capability and data sensitivity, caching by meaning rather than by URL, and authorization of tool calls made by agents. Many vendors ship both as one product with two policy sets.

When does an enterprise need an AI gateway?

The trigger is multiplicity. Once a second team is calling models, a second provider is in use, a model is running inside the perimeter, an agent is invoking tools, or an auditor has asked what was sent to a model, the per-application controls stop scaling and a shared enforcement point becomes necessary. Before any of those, a single application calling one hosted model through the provider's SDK is usually well served by the provider's own dashboard and keys. The cost of adding a gateway later is not the gateway itself but rewiring every application's keys, retries and error handling to sit behind it.

How much does an AI gateway cost to run?

Licensing is the smaller part. Open-source gateways cost engineering time to deploy and patch; commercial ones price by request volume, by tokens processed, or as a module on an existing API management subscription. The larger and less visible cost is the operating role: someone has to define and maintain the routing rules, the per-team budgets, the redaction patterns and the retention policy for the audit store, and reconcile the gateway's token counts against provider invoices. Deployments that budget for the software and not for that role tend to stall with the software installed and the policies unwritten.

Can an AI gateway run entirely on-premises?

Yes, and for many regulated deployments it must. Several gateways install as self-hosted software inside the perimeter and route to models that also run inside it, so no prompt or response leaves. Several SaaS gateways offer a private data plane, where the traffic and logs stay inside the perimeter while the control plane that holds policies and configuration is vendor-hosted. Whether that split is acceptable depends on whether the compliance regime treats API definitions, policies and lifecycle metadata as in-scope data. Air-gapped environments and no-third-party-control-plane policies rule the hosted option out and shorten the product list considerably.

Does an AI gateway prevent prompt injection?

It reduces the damage rather than eliminating the attack. A gateway can match prompts and tool results against known payload patterns, block obvious attempts, and refuse tool calls that the calling agent is not authorized to make regardless of what the prompt says. It cannot reliably detect every injection, because the attack is written in the same language the model is designed to follow. The durable defense is limiting what an injected instruction can achieve: scoped tool authorization, budgets, and a response path that redacts before anything leaves. The gateway is the natural place to enforce all three.

How does token-based budgeting work in an AI gateway?

The gateway tokenizes the prompt on the way in, receives the completion token count on the way back, and attributes both to whatever key identifies the caller: an application, a team, a project or an individual agent. Budgets are set per attribution and per window, hourly, daily or monthly, and a request that would exceed its budget is rejected or routed to a cheaper model before it reaches the provider. The counts should be reconciled against the provider's invoice, because tokenizers differ by model and cached or truncated responses are counted differently on each side. A persistent gap means the budgets are not doing what they appear to do.

What should a bank look for in an AI gateway?

Three properties, checked before any feature comparison. Prompts, responses and logs stay inside the perimeter, or leave only under a written redaction policy. The audit record captures prompt, response, model version, caller and every policy applied, and is retained as long as the decision it supported. Tool-call authorization is tied to the identity system, so an agent acts as a named principal with scoped permissions rather than through a shared credential the gateway holds. A product that satisfies those three is a candidate; one that satisfies two is a risk item. The vendor list comes last.

Would you like to discuss AI gateway design for your enterprise?

Duczer East architects model governance structures that align with enterprise compliance and operational boundaries, including AI agent governance on WSO2 and policy enforcement design.

Prefer email? info@duceast.com
Duczer East — Where Data Engineering Meets Agentic AI

The Practitioner's Briefing

Senior-level insights on agentic AI, data engineering, and enterprise integration — delivered to your inbox.