Routing

Which model should serve a given prompt, and what that would have cost against what an agent is actually spending on real traffic — with nothing dispatched until you decide to act on it.

Routing answers two different questions, and keeps them apart on purpose:

  • Before you ship: given one prompt, which model in the whole benchmark catalog is the best pick — on capability, cost or speed, by the policy you choose?
  • After you've shipped: on this agent's actual traffic, would a different model have done as well for less? Answered by watching, not by guessing.

Routing is not Gateway. Gateway owns reliability — retries and fallbacks at call time, when a provider is down. Routing owns capability — recommending which model, based on evidence, and never dispatching anything on its own without an explicit action from you.

The three tabs

TabWhat it does
Shadow routingPer-agent: start an observation window, watch what routing would have chosen on real traffic, and act on the recommendation once it is ready.
Recommend a modelPaste a prompt, get the best model for it right now — a dry run. No call is made and nothing is billed.
DecisionsThe ledger of every shadow decision recorded so far, across agents.

Recommend a model — a dry run

Paste a prompt, pick what to optimise for (capability, cost, speed — or a balanced default), and get back a ranked shortlist from the full benchmark catalog, with the evidence behind the pick: the classified task type, the difficulty band, and which of a model's benchmark scores drove the ranking.

This never dispatches a call. The recommendation ranks the whole catalog — a model is only marked as servable in your workspace, never filtered out — so the answer is the best model that exists for the task, not just whichever ones your gateway happens to be configured to reach.

A recommendation with its shortlist and evidence
A recommendation with its shortlist and evidence

Shadow routing — watching real traffic

For each agent, Routing can watch its real calls and work out, after the fact, what it would have picked instead — for free. The classifier used for this is the cheap deterministic one, never a model call, so watching a busy agent costs nothing and never taxes the calls it's watching.

  1. Start an observation window

    One to sixty days, per agent. Builder role or higher.

  2. Traffic accumulates a counterfactual

    Every governed call the agent makes is scored the same way the "Recommend a model" tab scores a pasted prompt — quietly, alongside the real call, never blocking or slowing it down.

  3. The window resolves to a recommendation, with a confidence tier

    Any traffic that carried a recorded routing decision produces a real verdict — it is labelled by confidence rather than withheld. 30 calls reaches provisional, 200 reaches confident; below 30 the verdict is still given, just marked low confidence. Only a window with no decisioned calls at all (or an anchor model with no benchmark coverage) resolves as "insufficient data".

  4. Request the switch

    Filing a request to switch an agent's base model is its own approval-gated action — always reviewed by an admin or owner before anything changes.

A window can optionally sample a small share of real calls for paired live calls: the same prompt, re-sent in the background to each candidate model under consideration, so you can see how they'd actually have answered side by side. This is the one place real money moves in Routing — a bounded, configurable percentage of traffic, and the candidate answers are never served to anyone; there is no automatic judge grading them, so it answers "how did it respond," not "which one was better."

Decisions

Every shadow decision recorded across every agent, in one ledger: which model was actually served, which model routing would have picked, and the cost gap between them on that call's own token counts — so "routing would have saved X" is a number built from real traffic, not an estimate.

The bars a switch has to clear

A recommendation is not made just because a cheaper model exists. The thresholds are deliberate, and worth knowing before you read a verdict as timid:

BarValueWhy
Capability tolerance5%A candidate may not be more than 5% worse on its worst task class — not its average. An option that is better on average and much worse at one thing is not a saving.
Minimum saving1% and $0.01Both must be cleared. A switch that saves a rounding error is churn, not value.
Confidence tiers30 / 200 callsProvisional, then confident. A verdict below 30 is still shown, marked low.
Measured-cost weighting10 live-test pairsAt 10 pairs, measured cost and list-price modelling carry equal weight; more pairs shift it toward what was actually observed.
Live error ceiling20% over 5+ samplesA candidate erroring above this is dropped regardless of price.

Next