Developers TypeScript SDK

TypeScript SDK

A typed, runtime-agnostic client for the governed chokepoint. Built on the platform fetch, so it runs on Node ≥18, Deno, browsers and edge runtimes with no extra runtime dependencies.

What this actually is, in plain terms

Forgebench sits in front of real model providers (OpenAI, Anthropic, Gemini, ...). Instead of your code holding an OpenAI key and calling OpenAI directly, it holds a Forgebench key (sk_...) and calls Forgebench, which forwards the request to whichever model you asked for. The forgebench object below is your one doorway into that: every method on it is a different thing you can ask Forgebench to do on your behalf.

Doing it this way, instead of just calling OpenAI yourself, is what buys you the governed chokepoint: budget checks before money is spent, a tamper-evident audit row for every call, and a provider credential your code never has to hold. See Core concepts for the vocabulary (Agent, Budget, Policy, etc.) this whole SDK is built around.

Install

npm install "github:seedlinglabs/forgebench-sdk#path:/sdk-ts"

Connect to the gateway

Every call (chat, agents, runs, tools) rides through one control-plane chokepoint: auth → tenant resolution (Postgres RLS) → budget pre-gate → per-tenant provider key injection → the gateway call → one metering event + one hash-chained audit row. You never send a provider credential (api_key/api_base for OpenAI, Anthropic, etc.): the SDK never has one to send, and the chokepoint strips a client-supplied one anyway.

import { Forgebench } from "@seedlinglabs/forgebench-sdk";

const forgebench = new Forgebench({
  baseUrl: "https://api.forgebench.ai",
  apiKey: process.env.FORGEBENCH_API_KEY!,
});
OptionWhat it is
baseUrlThe control-plane URL. Required: the client has no built-in default, unlike the Python SDK's localhost fallback.
apiKeyAn sk_... secret: the programmatic path.
tokenA control-plane session JWT, as an alternative to apiKey: the console's own credential.
timeoutMsPer-request timeout. Default 60_000.
maxRetriesRetries on transient (429/5xx/network) failures, with backoff. Default 2. Streaming requests are never auto-retried.
defaultHeadersMerged into every request.
fetchOverride, for tests or non-global-fetch runtimes.

Point baseUrl at https://api.forgebench.ai for the hosted control plane, or your own domain for a self-hosted / airgapped install; FORGEBENCH_BASE_URL below is your own fallback convention, not something the SDK provides:

const forgebench = new Forgebench({
  baseUrl: process.env.FORGEBENCH_BASE_URL ?? "http://localhost:8000",
  apiKey: process.env.FORGEBENCH_API_KEY!,
});

whoami

Confirms the credential is valid and tells you what it resolves to: call it once at startup to fail fast on a bad key rather than discovering it on your first real request:

const me = await forgebench.whoami();
console.log(me.tenant_id, me.roles, me.auth_method);

await forgebench.health(); // unauthenticated liveness

Chat completions

const res = await forgebench.chat.create({
  model: "mock-gpt",
  messages: [{ role: "user", content: "Hello, Forgebench." }],
});

console.log(res.choices[0]?.message.content);
console.log("tokens:", res.usage.total_tokens);
console.log("trace:", res.trace_id, "call:", res.call_id);

model can be omitted to use the server default. trace_id ties the response to its audit entry, its metering record and its trace. call_id is this call's own row on the ledger; pass it as parentCallId on the calls it causes (see Call lineage below). See Make your first call for the full response shape.

Streaming

for await (const chunk of forgebench.chat.stream({
  messages: [{ role: "user", content: "Stream me a reply." }],
})) {
  process.stdout.write(chunk.choices[0]?.delta.content ?? "");
}

Or accumulate straight to a string:

const text = await forgebench.chat.streamToText(
  { messages: [{ role: "user", content: "hi" }] },
  { onToken: (t) => process.stdout.write(t) },
);

The client parses the SSE stream and consumes the terminal [DONE] sentinel for you. The budget gate still runs before the stream opens, so a 402 throws immediately, never mid-stream. Streamed requests are never auto-retried.

Call lineage

Every governed response carries a call_id. Pass it as parentCallId on the calls it causes (a sub-agent's turn, the next step of a tool loop), and the ledger records them as children: cost rolls up to the root, and the console draws the run as a tree instead of a flat set of rows sharing a trace. trace_id groups; parentCallId structures. It's correlation only; it never affects whether a call is allowed or what it costs.

const plan = await forgebench.chat.create({ messages: [...] });

const step = await forgebench.chat.create(
  { messages: [...] },
  { parentCallId: plan.call_id },
);

Tools

There are two entirely separate gates for "may this agent call this tool," and which one applies depends on the tool's name in the model's tool_calls response. Sending the right shape matters more than anything else here: get the naming wrong and the gate that should authorize the call never sees it as one of its own.

Sending a tool schema

Below, search is a Tool Registry entry (bare name) and mcp_time_get_current_time is an MCP tool (mcp_ prefix). A tool call this agent isn't bound to never reaches this loop at all: the whole response is refused with 403 tool_not_authorized (Tool Registry) or 403 mcp_tool_not_authorized (MCP); see Errors below:

import { Forgebench, PermissionDeniedError, type ChatCompletionRequest } from "@seedlinglabs/forgebench-sdk";

try {
  const res = await forgebench.chat.create({
    model: "gpt-4o",
    messages,
    tools: [
      {
        type: "function",
        function: {
          name: "search",
          description: "Search the KB",
          parameters: { type: "object", properties: { q: { type: "string" } } },
        },
      },
      {
        type: "function",
        function: {
          name: "mcp_time_get_current_time",
          description: "Current time in a timezone",
          parameters: { type: "object", properties: { tz: { type: "string" } } },
        },
      },
    ],
    tool_choice: "auto",
  } as ChatCompletionRequest & Record<string, unknown>);

  for (const call of res.choices[0]?.message.tool_calls ?? []) {
    if (call.function.name.startsWith("mcp_")) {
      const result = await yourMcpClient.call(call.function.name, call.function.arguments);
    } else {
      const result = await yourToolDispatch(call.function.name, call.function.arguments);
    }
  }
} catch (e) {
  if (e instanceof PermissionDeniedError) {
    console.error("tool not authorized", e.code, e.body);
  } else {
    throw e;
  }
}

tool_calls on the response is OpenAI-shaped ({ id, type, function: { name, arguments } }); arguments is a JSON string, same as the raw API, not pre-parsed. Bind the tool on the Tool Registry or MCP page first; no amount of retrying the identical call fixes an unbound tool.

Other fields only reachable through a widened type

tools/tool_choice aren't the only fields missing from ChatCompletionRequest: the control plane's request model accepts several more that only reach it through the same as ChatCompletionRequest & Record<string, unknown> cast:

FieldWhat it's for
response_formatJSON mode, e.g. {"type": "json_object"} for structured output.
metadataThe dimensions tagging; see Metadata.
modalities["text", "image"], for image-generating models.
reasoning_effortA hint for reasoning-capable models (low/medium/high).
request_idYour own idempotency/correlation id.
sourceA free-text tag for where the call came from.

The agentTools resource (agent credentials only)

If the API key this client was built with was issued to a registered agent (not a human/developer key), forgebench.agentTools gives you the agent's live allowlist and a ready-made dispatch loop: the API key resolves the tool list server-side, so a binding revoked centrally drops out of the very next turn without you hardcoding or re-deploying anything. openaiSchema() builds the model's tool list from the live allowlist (GET /v1/agent-tools); dispatch then runs every tool_call the gate let through, reports each outcome, and returns the role: "tool" messages to append before the next turn:

const tools = await forgebench.agentTools.openaiSchema();

const turn = await forgebench.chat.create({
  model: "gpt-4o",
  messages,
  tools,
} as ChatCompletionRequest & Record<string, unknown>);

messages.push(
  ...(await forgebench.agentTools.dispatch(
    turn.choices[0].message.tool_calls,
    (name, args) => yourMcpClient.call(name, args),
    { callId: turn.call_id! },
  )),
);

dispatch never aborts the loop on a failing tool: a throwing executor is reported as that tool's error and surfaced to the model as { error }, so the model gets to decide what to do about a failed call, same as any other tool result. Pass typed argument schemas per tool with openaiSchema({ parameters: { search: {...json schema...} } }); the control plane only stores a tool's name and description, not its input shape.

agentTools.report(...) is the lower-level primitive dispatch calls internally; reach for it directly only if you're executing tools outside a simple executor callback.

Agents, runs, deployments

const agent = await forgebench.agents.create({
  name: "support-bot",
  model: "mock-gpt",
  system_prompt: "You are concise.",
});

const run = await forgebench.runs.create({ agent_id: agent.id, input: { topic: "refunds" } });
const finished = await forgebench.runs.waitForCompletion(run.id, { intervalMs: 1000 });
console.log(finished.status, finished.output);

const deployment = await forgebench.deployments.create({ agent_id: agent.id, name: "prod" });

Agent → agent (A2A tasks)

An agent bound to another by an operator can open a task on it through the door; the callee runs wherever it runs, and its own governed calls hang under the task in the ledger:

const task = await forgebench.agents.call("qp-research", {
  text: "10 MCQs on photosynthesis",
  data: { grade: 9 },
  parentCallId: plan.call_id,
});
const notes = artifactText(task);

To BE a callee with no inbound port (pull delivery: this agent's own credential long-polls the door):

await forgebench.agents.serve(async (task, reply) => {
  const turn = await forgebench.chat.create(
    { model: "mock-gpt", messages: [{ role: "user", content: messageText(task.input.message) }] },
    { parentCallId: task.call_id },
  );
  return reply.done({ text: turn.choices[0].message.content ?? "" });
});

The handler's return value can be a TaskReply (reply.done(...), reply.ask(...), reply.fail(...)), or just a plain string / object / array of artifacts, which is coerced into a completed reply automatically. A thrown error fails the task with its message rather than crashing the loop.

Budget enforcement (HTTP 402)

The chokepoint rejects over-budget tenants before any provider call, so no spend occurs on a rejected request:

import { BudgetExceededError } from "@seedlinglabs/forgebench-sdk";

try {
  await forgebench.chat.create({ messages: [{ role: "user", content: "again" }] });
} catch (err) {
  if (err instanceof BudgetExceededError) {
    console.error("Tenant is over budget:", err.code, err.requestId);
  } else {
    throw err;
  }
}

Errors

All non-2xx responses map to a specific subclass of ForgebenchAPIError, which carries status, code, body and requestId:

HTTPErrorRetry?
400BadRequestErrorNo: fix the request
401AuthenticationErrorNo: fix the credential
402BudgetExceededErrorNo: degrade deliberately
403PermissionDeniedErrorNo: fix scope/allowlist/binding
404NotFoundErrorNo
409ConflictErrorNo: fix the resource's state first
422UnprocessableEntityErrorNo: fix the request
429RateLimitErrorYes: auto-retried up to maxRetries
5xxInternalServerErrorYes: auto-retried

Network/abort failures raise ForgebenchConnectionError; per-request timeouts raise ForgebenchTimeoutError (a subclass of it). Transient failures (429/5xx/network) are retried with exponential backoff + jitter (maxRetries, default 2) before the SDK ever throws them to your code; by the time you catch a RateLimitError or InternalServerError, the SDK has already given up. Every other error means "this exact call will fail again unchanged": retrying it in a loop just produces the identical refusal, and for a tool call specifically, three repeated unauthorized attempts inside ten minutes escalate to their own 429 (tool_repeatedly_unauthorized) precisely to stop that pattern before it burns another provider call.

PermissionDeniedError carries which specific cause fired, but not always in the same place: tool refusals populate .code directly (e.g. .code === "tool_not_authorized"), while model_not_allowed nests its identifier under .body.detail.error instead; .code is null for that one. Check .body when .code comes back empty:

import {
  BudgetExceededError,
  PermissionDeniedError,
  AuthenticationError,
  NotFoundError,
  ConflictError,
  UnprocessableEntityError,
  RateLimitError,
  InternalServerError,
  ForgebenchConnectionError,
  ForgebenchTimeoutError,
} from "@seedlinglabs/forgebench-sdk";

try {
  const res = await forgebench.chat.create({ model: "gpt-4o", messages, tools } as ChatCompletionRequest & Record<string, unknown>);
} catch (e) {
  if (e instanceof BudgetExceededError) {
    console.warn("budget gate refused the call", e.body);
  } else if (e instanceof PermissionDeniedError) {
    console.error("not authorized", e.code, e.body);
  } else if (e instanceof AuthenticationError) {
    console.error("auth failed", e.body);
  } else if (e instanceof UnprocessableEntityError) {
    console.error("request rejected", e.body);
  } else if (e instanceof NotFoundError) {
    console.error("not found", e.body);
  } else if (e instanceof ConflictError) {
    console.error("conflicting state", e.body);
  } else if (e instanceof RateLimitError) {
    console.warn("rate limited after retries", e.body);
  } else if (e instanceof InternalServerError) {
    console.error("control-plane error after retries", e.body);
  } else if (e instanceof ForgebenchTimeoutError) {
    console.error("request timed out");
  } else if (e instanceof ForgebenchConnectionError) {
    console.error("could not reach the control plane", e);
  } else {
    throw e;
  }
}

Governance: audit + metering

const summary = await forgebench.metering.summary();
console.log(`spent $${summary.spent_usd} of $${summary.monthly_limit_usd}`);

const entries = await forgebench.audit.list({ limit: 20 });

entries[i].prev_hash / row_hash form a per-tenant tamper-evident hash chain.

API keys

const created = await forgebench.keys.create({ name: "ci", scopes: ["chat:write"] });
console.log(created.api_key); // shown ONCE, store it now

const keys = await forgebench.keys.list();
await forgebench.keys.revoke(created.id);

keys.list() returns the safe view: no secrets.

Next