Agents MCP tool authorization

MCP

The catalog of MCP servers your agents are allowed to call: a name, a tool list and tags, never a connection.

Forgebench never dials a registered server and holds no credential for one. The agent's own MCP client talks to the real server directly, exactly as it does today. A registered server is a name an operator has agreed to govern, plus an optional hand-typed tool catalog and tags: no endpoint, no auth, no OAuth connect, nothing to refresh.

What gets checked, and when

The model's response can include a tool_calls entry naming an MCP tool, by convention, a qualified name like mcp_time_get_current_time. Forgebench compares that name, verbatim, against the bindings the agent has been given. If it matches a live binding, the agent is allowed to see that instruction and act on it, outside Forgebench, with its own client, its own credentials. If it doesn't, the agent never learns "call X" in the first place, because this is its only channel for that instruction.

Registering a server, and the server's own detail page: its tools, and each one's kill switch

Register a server

  1. Name it

    How agents and the audit log refer to this server. Not editable after creation: every binding and audit entry already reference it.

  2. List its tools

    Required. One qualified_tool_name: description per line, using the exact name that will appear in a governed tool_calls response, not a bare tool name. This is the list a binding is validated against; an agent can only be bound to a name entered here.

  3. Tag it

    Comma-separated. How you group and filter a catalog once it outgrows one screen.

Editing a server's tool list or description takes effect on the next call: no deploy, no agent-side change.

A server's own detail page

Opening a server from the catalog gives you four tabs: Overview, Tools, Agents, and Ceilings.

TabWhat it shows
ToolsEvery tool declared for this server, each individually enabled or disabled. Disabling one here is a kill switch: it stops every bound agent from calling that tool on the next attempt, independent of any agent's own binding, still an operator-entered hint, not a live discovery, since Forgebench never connects to the server to check.
AgentsEvery agent currently bound to this server.
CeilingsA per-tool call ceiling, added from the Tools tab with Add ceiling. Blank means no ceiling on that tool.

Binding servers to agents

Binding happens from an agent's own detail page, not from here: the MCP catalog is the shared inventory; the allowlist is a property of the agent.

  1. Open the agent, then its Tools tab

    Choose the MCP Tools sub-tab. Registry Tools next to it is the separate Tool Registry allowlist, a different gate entirely.

  2. Suggest tools, or bind manually

    Suggest tools ranks the catalog for this agent. Accepting a suggestion is what actually creates a binding; suggesting alone grants nothing. Bind tools skips the ranking: pick a server from the dropdown, then tick as many of its declared tools as this agent needs; ticks are kept if you switch servers, so one dialog can bind tools from several servers in one pass.

  3. Confirm

    Only a confirmed selection becomes a live allowlist entry, whichever path you took to get there. Each binding can be revoked individually from the same MCP allowlist table.

Binding two of a server's tools to an agent from its MCP Tools sub-tab

Set a call-volume ceiling

Binding decides whether an agent may relay a tool call at all; a ceiling is a separate, optional cap on how often: a rate limit (per minute/hour) or a quota (per day/month), same mechanism either way.

Add one from a server's own Tools tab (Add ceiling on a tool's row, scoped to that server + qualified tool name) or from the server's Ceilings tab, which lists and manages every ceiling already set on it. Leave the agent field blank to cap every agent calling that tool combined, or pick one to scope the ceiling to just that agent. Tripping a ceiling refuses with 429, a different outcome from an unbound tool's 403.

Handle it in your code

Nothing changes about how your agent calls the model or dials the MCP server: this gate sits entirely inside the governed chat completion. Bind the tool here, then read tool_calls off the response exactly as you already do, and dispatch to your own MCP client only for entries whose name matches what you bound.

tools/tool_choice aren't first-class parameters on either SDK's chat call: the control plane accepts and forwards them untouched (extra=ignore), so send them through Python's extra_body passthrough, or a widened type on TypeScript (tools isn't declared on ChatCompletionRequest yet). On Python, tool_calls comes back as a list of plain dicts (OpenAI-shaped on the wire), not attribute-access objects: index into it (call["function"]["name"]), don't dot into it.

A tool call this agent isn't bound to never reaches this loop at all: the whole response is blocked before your code sees it, raised as a PermissionDeniedError with .code set to mcp_tool_not_authorized (unlike model_not_allowed, this one's identifier does live under .code, not just .body):

curl -X POST https://api.forgebench.ai/v1/chat/completions \
-H "Authorization: Bearer sk_..." \
-H "Content-Type: application/json" \
-d '{
  "model": "gpt-4o",
  "messages": [{"role": "user", "content": "What time is it in Tokyo?"}],
  "tools": [{
    "type": "function",
    "function": {
      "name": "mcp_time_get_current_time",
      "description": "Current time in a timezone",
      "parameters": {"type": "object", "properties": {"tz": {"type": "string"}}}
    }
  }]
}'

A 403 mcp_tool_not_authorized means the model proposed an mcp_... name this agent has no binding for: bind it here, or fix the tool schema you sent, not the calling code. Against curl, that refusal is a plain JSON body, no typed exception to catch:

{"detail":{"code":"mcp_tool_not_authorized","tool":"mcp_time_get_current_time","reason":"not_permitted","message":"this agent may not call that tool"}}

Bulk CSV import

For registering many servers at once, import a CSV with columns name, description, tags, tools: tools holds the whole catalog as one cell, one qualified_tool_name: description line per row, standard CSV quoting around the embedded newlines. A duplicate name or a row missing its tool list is skipped, not fatal; the rest of the file still imports.

Revoke vs delete

ActionWhat happens
RevokeEvery agent bound to this server stops being able to call it on its next attempt: no deploy, and no window where some agents have lost access and others have not. The server row stays, now marked revoked, so the catalog keeps telling the truth about what it once allowed.
DeleteOnly a revoked server can be deleted. This permanently removes the row, including its now-revoked bindings. The audit log still records that it existed, was revoked, and is now deleted, but the catalog itself can no longer show that. There is no undo.
  1. Open the ⋯ menu on the server's row, choose Revoke

    Only live (non-revoked) servers offer it: a revoked server's menu only offers Delete.

  2. Read the confirmation, then confirm

    Revoke <name>? explains what's about to happen: every bound agent loses access on its next attempt, existing bindings are marked revoked rather than deleted, and there's no partial window where some agents still have access and others don't.

Delete only becomes available once a server is already revoked, and asks you to confirm separately: it has no undo, unlike revoke.

MCP vs Tool Registry

These two pages look similar and govern completely different things. MCP is a catalog of servers your agents already talk to directly: it authorizes a relay instruction, nothing executes here. Tool Registry is a response gate for tools with no server at all: it authorizes the LLM's tool_calls entry itself, for a tool your own code executes after the call returns.

The two gates share the same response but never the same name: a tool_call starting with mcp_ belongs exclusively to this page's allowlist; every other name belongs exclusively to the Tool Registry. Confusing the two means configuring the wrong gate for a tool that then never gets authorized by either.

Next