Serving tasks
An agent becomes a callee the moment another agent has an approved binding to call it. What that callee actually does with an incoming task is entirely up to its own process — Forgebench hands the task over, either by letting the callee pull it or by pushing it to a URL, and waits for a reply in the same shape either way.
Pull: the callee claims its own work
If the callee has no endpoint_url set, it's pull-only: its own process
claims work by long-polling GET /v1/agents/me/tasks/next with its own
agent credential, then posts the result back. Nothing has to be reachable
from the outside — no inbound port, no public IP, no firewall rule to open.
That's the same shape a Temporal worker or a CI runner already uses to pull
jobs instead of accepting them.
The claim query uses FOR UPDATE SKIP LOCKED, so running several replicas
of the same callee polling at once is safe: each replica gets a different
task, and none is ever handed out twice.
import os
from forgebench import Forgebench
client = Forgebench(api_key=os.environ["CALLEE_AGENT_KEY"],
base_url="https://api.forgebench.ai")
def handle(task):
invoice_id = task.data.get("invoice_id")
if invoice_id is None:
return task.ask("which invoice_id should I classify?")
try:
category = classify(invoice_id, parent_call_id=task.call_id)
except ClassificationError as exc:
return task.fail(str(exc))
return task.done(data={"category": category})
# blocks; returns the number of tasks handled
client.agents.serve(handle, max_tasks=None)import { Forgebench } from "@seedlinglabs/forgebench-sdk";
const forgebench = new Forgebench({
baseUrl: "https://api.forgebench.ai",
apiKey: process.env.CALLEE_AGENT_KEY!,
});
await forgebench.agents.serve(async (task, reply) => {
const invoiceId = task.input.message?.parts.find((p) => p.data)?.data?.invoice_id;
if (!invoiceId) return reply.ask("which invoice_id should I classify?");
try {
const category = await classify(invoiceId, { parentCallId: task.call_id ?? undefined });
return reply.done({ data: { category } });
} catch (err) {
return reply.fail(err instanceof Error ? err.message : String(err));
}
});serve() claims a task, hands it to your handler, and posts the reply,
looping until max_tasks/maxTasks is reached or stop()/an aborted
signal says to quit. The handler can return:
- a
TaskReplybuilt fromtask.done(...)/task.ask(...)/task.fail(...)(Python) or thereplybuilder (TypeScript) - a plain string, dict/object, or list of artifacts — converted into a
completedreply automatically, so a handler that just wants to answer doesn't have to build the envelope by hand - nothing at all — raise/throw, and the exception's message fails the task instead of crashing the whole loop
Finishing a claimed task manually
Without serve(), claim and reply are two separate calls — reach for this
if you already have your own polling loop or worker framework and serve()
would just be fighting it:
task = client.agents.next_task(wait=20)
if task is not None:
reply = task.done(text="done") # or task.ask(...) / task.fail(...)
client.agents.reply(task.id, reply)const task = await forgebench.agents.nextTask({ wait: 20 });
if (task) {
await forgebench.agents.reply(task.id, { state: "completed", artifacts: [] });
}A reply is only accepted while the task is working, and only from the
callee that claimed it — if the caller cancels the task in the meantime,
the reply is refused with 409 rather than silently applied over a
canceled task.
Push: a webhook instead of polling
Setting an endpoint_url on the callee switches delivery from pull to
push: the control plane POSTs every task to that URL instead of the callee
polling for one. This trades "no inbound port" for "no polling loop to
run" — worth it once the callee already has a always-on HTTP service
anyway (a FastAPI app, a Lambda behind a URL, and so on).
Configure it from the callee's own Calls tab, or the same call via the API:

curl -X PUT https://api.forgebench.ai/v1/agents/$CALLEE_ID/endpoint \
-H "Authorization: Bearer $ADMIN_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://callee.example.com/a2a",
"auth": {"type": "bearer", "secret": "s3cret"}
}'Notes on the endpoint config:
auth.secretis sealed in the vault and never returned by any API response or shown again in the console. Omitauthon a laterPUTto keep the current credential, or send"clear_auth": trueto drop it."url": nullreverts the agent to pull-only.- A private-range URL (RFC 1918, etc.) requires
"allow_private_network": trueplus admin privilege to set. Loopback and cloud-metadata addresses are always blocked, both when the URL is saved and on every dial after that — a hostname that gets re-pointed at a private address later doesn't get a free pass just because it passed the check once.
The control plane POSTs the task to the configured URL:
POST https://callee.example.com/a2a
Content-Type: application/json
X-Forgebench-Task: <task_id>
X-Parent-Call: <task's own call_id>
Authorization: Bearer s3cret # or your configured header
{"task_id": "...", "call_id": "...", "message": {"role": "user", "parts": [...]}}Your handler answers in the same shape a pull reply uses (TaskResultIn):
from fastapi import FastAPI, Request
app = FastAPI()
@app.post("/a2a")
async def handle_task(request: Request):
body = await request.json()
parts = body["message"]["parts"]
text = next((p["text"] for p in parts if "text" in p), "")
# pass this to parent_call_id on this agent's own governed calls
parent_call_id = body["call_id"]
return {
"state": "completed",
"artifacts": [{"name": "response", "parts": [{"text": f"handled: {text}"}]}],
}The response body doesn't have to match TaskResultIn exactly — this is
what makes a plain webhook with no A2A awareness at all still work:
- a JSON object without a
state/artifactskey is wrapped as adataartifact holding the whole body - a non-JSON (plain text) response becomes a
textartifact {"state": "working"}tells the control plane this callee will finish asynchronously; it later posts the real result toPOST /v1/agents/me/tasks/{id}/resultwith its own agent credential, the same route a pull callee uses to finish up
Lineage: the call tree
Every governed call returns call_id in its response body (and
X-Call-Id on the streaming path). Send X-Parent-Call: <call_id> on a
call that another call caused, and the control plane stamps
parent_call_id/root_call_id on it. This is bookkeeping, not policy — no
gate reads it, and nothing is refused because of it.
A task participates in the same lineage:
- the caller's
parent_call_id/root_call_idare stamped on the task row when it's opened - the task's own
call_idbecomes the parent for every governed call the callee makes while working on it
Chain that through a few hops — agent A calls agent B, which calls a model and then agent C — and the whole workflow becomes one tree on the ledger, with no orchestrator anywhere having to declare the relationship up front. One query gets its total cost:
SELECT sum(cost_usd) FROM metering_events WHERE root_call_id = :root;Pass parent_call_id (Python) / parentCallId (TypeScript) to
agents.call(...) with the governed call that led to opening the task. A
Task object exposes its own call_id/parent_call_id/root_call_id so a
serve() handler can continue the chain downward.

