MCP gateways turn connecting agents to third-party data into a platform job
Nolan Sullivan
September 23, 2026 · 11 min read
_Model Context Protocol (MCP) gateways are becoming an incredibly useful piece of infrastructure, but use cases vary company to company. This post covers one of the most common use cases: enablement. MCP gateways are the most efficient way to connect agents to the data employees need to use. AI enablement teams rolling out agents run into three problems:
- Every server has a different auth model
- Nobody knows which servers are approved for which team
- Every server added makes the agent's context window worse.
A gateway fixes all three in one place, for every client at once. This post expands on the enablement half of what an MCP gateway is, the reference that also covers the architecture and the security side.
Authentication is the first hurdle
Without a gateway, authentication happens per person, per client, per server. Take a rollout of 300 engineers who each use two clients (say Claude Code and Cursor) against eight remote servers. That is 4,800 OAuth flows before anyone does any work, and each one repeats when its token expires or a client is reinstalled. Servers without OAuth are worse, because the fallback is a personal API key pasted into a JSON config file on a laptop, and the enablement team fields the tickets when it stops working.
A gateway splits the problem in two. A person signs in once through the company's identity provider (IdP), so 300 engineers on two clients is 600 sign-ins. Toward the servers, the gateway acts as the OAuth client to each upstream service on the person's behalf, stores and refreshes their token, and holds a single shared credential for services that have no per-user OAuth. Because the gateway presents OAuth to the client, agents that expect modern auth can reach an upstream that only issues API keys.
And with Enterprise Managed Auth (EMA) the authentication process is getting even smoother. Under EMA, the IdP decides who may reach a server and hands the client a token for it during single sign-on, so the per-server consent screen goes away and an admin revokes access from the IdP console. The catch is that EMA has to be supported at both ends of every connection, and each client and server vendor ships on its own schedule. A gateway implements EMA once, at the point every client and server already pass through, so the whole catalog gets IdP-driven access, including servers that will never add support themselves.
A curated catalog with roles beats a KB page of server URLs
Without a gateway, the catalog at most companies is a knoweldge base page, or a Slack channel where people share configs. The enablement team can't say which servers are in use, and a new hire's first week is self-assembling their agent setup.
A gateway replaces that with one list of vetted servers, third-party and internal, and a view of that list per role. Sales sees Salesforce and Gong, engineering sees GitHub, Linear, and Sentry, and support sees Zendesk. When roles map to groups in the directory, a new hire gets their tools on the first day and a leaver loses them on their last, without a ticket. Access can narrow below the server, so a role might get the read-only tools on a server and nothing that writes. Distribution matters as much as the list. One gateway URL, or a plugin that installs a role's servers into Claude Code, Cursor, or Codex, makes the approved path shorter than the unapproved one. The agent access management page covers the access model in more depth.
The cost of curation is that the platform team becomes a gatekeeper. If adding a server to the catalog takes two weeks of review, people go back to pasting URLs into config files, and the catalog describes a smaller world than the one employees work in. A curated catalog only holds if new servers get in quickly, including the company's own APIs, which need a path to becoming tools without someone writing and hosting a server for each one.
Context cost depends on the client, a gateway evens it out
How much context MCP eats varies a lot from client to client. Every client loads tool definitions into the context window before the model sees the request, but they don't all load them the same way. Claude defers MCP tool definitions by default and pulls them in when a task calls for them. Other clients load every tool from every connected server up front, on every request.
So the same eight servers are cheap in one client and crippling in another, and which one an engineer gets depends on which agent they opened that morning. A gateway takes the guessing game out by deciding how tools are presented once, upstream of every client, so a server costs the same context in Claude Code, Cursor, and Codex. Two patterns do the work, and a gateway is where an enablement team can apply them to servers it didn't build:
- Dynamic tools. The gateway exposes a small set of meta-tools, such as
search_toolsandexecute_tool, and the model discovers real tools as it needs them. In the same benchmark, a 400-tool server started at about 2,500 tokens with progressive search and 1,300 with semantic search, flat as the toolset grew. - Code mode. The model writes a short program against the tools and runs it in a sandbox, so intermediate results stay in variables instead of passing back through the context. Anthropic reported one workflow dropping from 150,000 tokens to 2,000.
Both have costs. Dynamic tools add at least one round trip before the first real call, and semantic search can miss a tool the model didn't know to ask for. Our benchmark was a handful of runs on simple tasks, and we labeled it preliminary. Code mode needs a sandbox, retries when the generated code fails, and turns a log of discrete tool calls into a log of scripts, which makes debugging and audit harder. It also adds overhead to the simple one-call interactions that make up most MCP traffic, a trade-off we covered in comparing MCP server generators. A gateway lets the enablement team choose per server: dynamic tools on the 300-tool CRM, static definitions on the five-tool internal service.
Enablement is the day-zero job, and the same layer does the next two
Enablement is usually the first reason a company puts a gateway in the path. Buyers we've talked to this year start with the IdP connected to a gateway and a catalog for a few teams, and add inspection and deeper observability later. Security and governance are separate jobs, covered in companion posts in this series. Security covers inspecting traffic, stopping data leaks, and finding shadow MCP servers the gateway never sees. Governance covers who is acting on each call, what policy applies, and what the audit log records, described in agent governance.
The jobs overlap because they depend on the same data. The identity a person uses to sign in is the identity policy is evaluated against, and the catalog is the inventory a security review asks for. That is the reason to pick an enablement gateway with the later jobs in mind: replacing the layer that holds everyone's upstream tokens is a larger project than adding policy to it.
Who builds MCP gateways for enablement
Speakeasy ships its MCP gateway as one part of the AI Control Plane. People sign in once through the company's IdP, and the control plane holds a token per user for services like Linear, GitHub, and Slack, encrypted at rest and refreshed silently. The catalog covers common SaaS tools, internal APIs become servers generated from the existing API, and role-based access applies at the server, toolset, and tool level, so each team gets its own sub-catalog. Plugins publish role-based bundles to Claude Code, Cursor, and Codex, servers behind the gateway support dynamic tools through search_tools and execute_tool, and private servers connect over outbound-only tunnels. The control plane runs on Speakeasy-managed infrastructure and doesn't route model traffic, so it is the wrong pick for an air-gapped deployment or for a team that needs model failover from the same product.
Microsoft has two separate products. Microsoft MCP Gateway is an open-source reverse proxy and management layer for MCP servers on Kubernetes, with session-aware routing, a control plane for deploying and updating servers, and bearer-token RBAC (docs). Azure API Center can act as an MCP registry, an inventory where teams register and discover servers. The registry is a catalog and the gateway is a proxy, so a team on Azure that wants a curated list and a single path to its servers pairs the two.
IBM ContextForge (repo, docs) is an open-source registry and proxy that federates MCP servers, A2A agents, and REST or gRPC services behind one endpoint. It suits a platform team that wants to run the gateway itself and bring existing non-MCP services into the same catalog.
Commercial options beyond those include MintMCP, Runlayer, Kong (AI MCP Proxy), Bifrost, TrueFoundry, Willow (formerly Webrix), Docker MCP Gateway, and AWS AgentCore Gateway. We compared eleven of them in the best MCP gateways for enterprise, and which MCP gateway architecture do I need covers the build-or-buy question.
Where to start
- Put the five most-requested servers behind a gateway, with sign-in through your IdP, and count how many OAuth flows and config tickets disappear.
- Map two or three roles from existing directory groups and give each its own view of the catalog, read-only by default.
- Measure tool-definition tokens on your largest servers, then switch those servers to dynamic tools and compare cost and task completion before rolling it wider.
If you're an enablement team working through the first of these and want to compare notes, we'd like to talk.
What is an MCP gateway?
An MCP gateway is a proxy that sits between AI agents and the Model Context Protocol (MCP) servers they call, so every tool call passes through one point for identity, access policy, inspection, and audit. It serves two jobs on the same traffic: enablement, which this post covers, and governance. The MCP gateway reference covers the full definition and architecture.
What does an MCP gateway do for AI enablement?
For an AI enablement team, an MCP gateway handles the work of connecting agents to third-party data. People sign in once through the company's identity provider, and the gateway authenticates to each upstream MCP server on their behalf. It gives each role a curated catalog of approved servers, and it reduces context bloat with dynamic tools or code mode, so agents stay usable as more servers are connected.
How does an MCP gateway simplify MCP authentication?
Without a gateway, every person completes an OAuth flow for every server in every client, and repeats it when tokens expire. An MCP gateway signs the person in once through the IdP, acts as the OAuth client to each upstream service, stores and refreshes per-user tokens, and holds a shared credential for services with no per-user OAuth. Each person still consents once per upstream service.
What is MCP RBAC in a gateway?
MCP role-based access control (RBAC) in a gateway decides which servers and tools each role can see and call. Roles usually map to directory groups in Okta or Entra, so access follows people as they join, move teams, or leave. Gateways differ in how finely access can narrow: some stop at the server, others go down to individual tools or read-only annotations.
What is MCP code mode, and how is it different from dynamic tools?
MCP code mode has the model write a short program against the available tools and run it in a sandbox, so intermediate results don't pass back through the context window. Dynamic tools keep discrete tool calls but hide tool definitions behind meta-tools such as search_tools and execute_tool, loading definitions only when the model asks. Code mode saves more on multi-step workflows but needs a sandbox and makes calls harder to audit; dynamic tools add a discovery round trip.
Do I need an MCP gateway if my AI client already has a connector directory?
Not for one team on one client with a few servers, where the client's own connectors and OAuth support are enough. A gateway pays off when an enablement team supports several clients, dozens of servers, and hundreds of people, because sign-in, the role-based catalog, and context controls apply once at the gateway for every client instead of separately inside each one.
Last updated on