B2B AI agent API access: governing your partners' agents without a second perimeter
How to govern B2B AI agent API access when partners' agents start calling your APIs: scoped MCP credentials, per-partner rate limits, and one audit trail.
- mcp
- ai-agents
- partners
- governance
- access-control
The email arrives from your most technical partner. Their integration team has rebuilt their side of the connection as an AI agent, and they want to know if you expose an MCP endpoint. Or the email never arrives, and you find out another way: a partner's traffic pattern changes overnight from steady scheduled batch calls to bursty, exploratory request sequences that read like something probing your API surface. Either way, B2B AI agent API access has stopped being a future planning item. Your partners are already building agents, and those agents want to call your APIs.
This is a different problem from governing your own internal agents. You control your own agents' code, prompts, and deployment. You control none of that for a partner's agent. The only place you can govern it is the boundary where its requests enter your infrastructure, which means your gateway either handles agentic traffic as a first-class consumer type or you lose visibility into a growing share of partner traffic.
The uncomfortable part: most partner agents today authenticate with credentials issued for the partner's application, wired into an agent framework the partner adopted after you issued them. On your side, nothing distinguishes the agent from the app. Same key, same scope, same rate limits, one blended line in your logs.
Why the standard answers fail
The first standard answer is to refuse: partner contracts permit application integration, not agents, so agents are prohibited. This fails quietly. The partner's agent still calls your API with the application's credentials, because from the partner's point of view the agent is just their new client implementation. Prohibition without a technical control does not stop the traffic; it only guarantees the traffic is unlabeled. When a regulator or your own security team later asks which partner requests were agent-initiated, you have no way to answer.
The second standard answer is to stand up separate AI infrastructure. This is the route the incumbent vendors point to: Kong sells an AI gateway alongside its API gateway, and teams on Apigee or AWS API Gateway, which expose no MCP surface at all, end up building custom shim services per partner. Now you run two perimeters. Partner applications authenticate through one stack, partner agents through another, and every control you care about, credential issuance, scoping, rate limits, audit, exists twice and drifts apart. The two-gateway problem is bad enough for internal agents; for partner-facing access it doubles your external attack surface.
The third answer is to treat the agent as just another API client and hope for the best. This gets closer, but generic client records are not enough either. An agent's failure mode is not an app's failure mode. A buggy app repeats one bad call; a misdirected agent explores. It will discover endpoints you never mentioned in the integration guide, retry in loops, and chain calls in orders no human integration developer would. Governing that takes per-identity discovery scope, method restrictions, and limits tuned to agentic behavior, not just a shared API key with a global rate limit.
What B2B AI agent API access needs from your gateway
The requirements come straight from the failure modes. A partner's agent needs its own identity, separate from the partner's application, so the two are distinguishable in every log line. It needs a discovery surface, because agents work by finding out what they can call rather than reading your PDF integration guide. That discovery must be scoped, so the agent sees exactly the endpoints this partner is contracted for and nothing else. Writes need to be restrictable independently of reads. And all of it must land in the same audit trail as the partner's application traffic, because your compliance obligations do not care which kind of software made the request.
Zerq handles this with the same primitives that govern every other consumer: clients, profiles, policies, and the Gateway MCP surface. A client represents the consumer, a profile is its runtime access contract, a policy caps its throughput, and the MCP endpoint gives the agent a governed way to discover and execute your published APIs. No separate deployment, no second credential store.
Step by step: provisioning a partner's agent
Here is the full provisioning flow in the management UI. Assume the partner is Northline, an existing partner whose application already calls your payments collection.
-
Create a dedicated client for the agent. Go to Clients in the sidebar, click New Client, and fill in the name (
northline-agent), a description noting this is agentic traffic, the partner's contact email, and the collections this client may access. Assign only the collections in Northline's contract. Zerq auto-creates a default profile with token authentication when you save. -
Scope the profile. Open the client's profile and set Allowed Methods to
GETif the agent only reads, orGET, POSTif it creates resources. Requests using any other method return405 Method Not Allowed. Add the partner's egress ranges under IP Restrictions in CIDR notation, for example198.51.100.0/24. IP checks run before authentication, so a request from outside the partner's network is rejected before the token is even verified. -
Attach a conservative policy. Create a policy with a rate limit like 120 requests per
1msliding window and a quota such as 100,000 requests per30d. Agents retry and loop in ways applications do not, and the policy is what turns a runaway agent loop into a stream of429responses instead of load on your backends. Quotas reset on the 1st of each month at 00:00 UTC, which maps cleanly onto contractual monthly volume. -
Hand the partner their connection block. The profile detail page has copy buttons for the
X-Client-ID,X-Profile-ID, and the fullAuthorizationheader. The credential is displayed once at creation, so copy it into your secrets handover process immediately. The partner drops the values into their MCP client configuration:
{
"mcpServers": {
"acme-payments": {
"url": "https://gateway.example.com/mcp", // your Gateway MCP endpoint
"transport": "streamableHttp", // MCP over streamable HTTP
"headers": {
"Authorization": "Bearer <token>", // profile credential, shown once at creation
"X-Client-ID": "<client-id>", // identifies the partner agent client
"X-Profile-ID": "<profile-id>" // selects the scoped agent profile
}
}
}
}
- Verify the scope before go-live. Run one allowed call and one denied call with the new identity. The denied call, a
POSTon a read-only profile or a request for a collection outside the assignment, must come back403or405with the deny reason recorded in the request logs. The security governance guidance for MCP access is to repeat that denied-path test on every policy release, so scope regressions surface as a failed test instead of an incident.
What the partner's agent actually sees
Once connected, the agent initializes an MCP session and receives an Mcp-Session-Id header to carry on subsequent calls. From there it has exactly four tools: list_collections, list_endpoints, endpoint_details, and execute_endpoint. That fixed surface is itself a control. The partner's agent cannot be handed a poisoned tool catalog or talked into calling something outside the gateway, because the only tools that exist are the ones the gateway serves, and list_collections returns only the collections assigned to this client.
A real call looks like this:
{
"jsonrpc": "2.0",
"id": 3,
"method": "tools/call",
"params": {
"name": "execute_endpoint",
"arguments": {
"endpoint_id": "ep_charges_list",
"query_params": { "limit": 10, "status": "settled" }
}
}
}
The result carries the upstream status, the parsed body, response headers, a latency_ms timing, and isError: true whenever the upstream returns a 4xx or 5xx, so the agent gets an unambiguous failure signal instead of hallucinating around an HTML error page. Each execute_endpoint call has a 60-second timeout and is subject to the profile's method restrictions and the client's policy, exactly like a REST request. Everything the agent needs to plan a call, path parameters, query parameters, request and response JSON Schemas, worked examples, comes from endpoint_details, which means your published proxy definitions are the integration documentation.
One audit trail, and how you query it
Every tool call the agent makes creates a request log entry with the same fields as the partner application's REST traffic: request ID, timestamp, method, path, target endpoint, status code, latency, client ID, profile ID, matched collection, client IP, and full request and response headers and bodies. Because the agent has its own client, separating the two is a filter, not a forensic project. Filter request logs by the northline-agent client ID for everything the agent did; filter by the application's client ID for everything the app did; filter by status 4xx on the agent's profile ID to watch what the agent attempted and was denied. Filters are URL-synced, so the compliance view for a partner's agentic traffic is a bookmarkable link you can hand to an auditor.
That denied-request view deserves a standing place in your review cadence. A rising 403 trend on an agent profile means the agent is repeatedly reaching for scope it does not have, which is either an integration bug on the partner's side or the early signature of a compromised or misdirected agent. Either way you want the conversation with the partner before the pattern grows, and the log gives you request IDs to attach to it.
Offboarding and incidents use the same lifecycle controls as any consumer. Toggle the profile inactive to pause the agent instantly without touching the partner's application access. Rotate the credential with an expiry period in hours when a contract requires time-boxed access. Deactivate the whole client and every request answers 403.
A checklist before you say yes
Before enabling a partner's agent, answer these concretely:
- Which collections is this partner contractually entitled to, and is the agent client assigned to exactly those and no more?
- Does the agent need write methods at all? If yes, which ones, on which endpoints, and did the partner state the business action behind each?
- What are the partner's egress IP ranges, and are they pinned on the profile?
- What request volume does the partner project for the agent, and is the policy set near that projection rather than at your global default?
- Who at the partner owns the agent's behavior, and what is the escalation path when you see a denied-request spike?
- Have you executed and logged one allowed and one denied test call for the new identity?
If a partner cannot answer the write-method or ownership questions, issue a GET-only profile and revisit after the first month of logs.
What this looks like in practice
A payments platform serves forty B2B partners through its gateway. One partner, a mid-size lender, replaces its nightly reconciliation batch job with an agent that queries settlement data on demand and drafts exception reports. Before the change, the lender's traffic was a predictable 2 a.m. burst under one application credential. The platform team creates a second client for the lender's agent, GET-only, pinned to the lender's egress CIDR, assigned solely to the settlements collection, with a 120-per-minute rate limit. The lender's agent config points at the gateway's /mcp endpoint with its own headers.
The result is that nothing about the platform's open banking compliance posture had to change. The same gateway enforces the same per-partner scope on the new traffic. When the lender's agent misbehaved in week two, a planning bug drove it to re-request the same settlement window in a tight loop, the policy answered with 429s, the backends never felt it, and the platform team sent the partner a filtered log link showing exactly what the agent did, down to request IDs. The quarterly access review now lists the lender twice, app and agent, each with its own scope, its own limits, and its own line of evidence.
The alternative timeline is worth stating: without the dedicated identity, that loop would have burned the lender's shared quota, throttled their production application, and shown up in the logs as the partner's own app misbehaving.
Partner AI agents are not a reason to build a second perimeter, and they are not safe to absorb invisibly into existing application credentials. Zerq gives you a governed middle path: one gateway where a partner's agent gets its own identity, a discovery surface scoped to its contract, method and network restrictions enforced before your backends see a byte, throughput limits tuned to agentic behavior, and an audit trail where agent and application traffic are one filter apart. The full capability set runs in your own infrastructure, so the governance boundary for your partners' agents stays inside yours.
Zerq is an enterprise API gateway built for regulated industries — one platform for API management, AI agent access, compliance audit, and developer portal, running entirely in your own infrastructure. See how it works or request a demo to walk through your specific requirements.