Guardrail findings
Requires the Pro plan or higher; see Billing.
A post-fact scan of already-answered traces for PII, secrets and denylisted terms. This is evidentiary, not preventive — it never blocks a response. What it gives you is the record: what got said, whether it looked like a leak, and — if the pattern keeps repeating — a way to stop the agent before the next one, via auto-pause.
The checks
Three checks, each turned on per agent, per check:
| Check | What it flags |
|---|---|
| PII detection | Emails, phone numbers, card numbers and similar identifiers found in a scanned trace. |
| Secrets & credentials | API keys, bearer tokens and other credential-shaped strings found in a scanned trace. |
| Denylist | Any of a set of banned terms or phrases — your own words, chosen per agent or tenant-wide, as a plain phrase or a regex. |
Each check has two modes:
- Flag — counts toward auto-pause.
- Log — keeps a record, does not count toward auto-pause.
There is no block mode. A scan runs after the response is already delivered, so there is nothing left to block.
Set up guardrails for an agent
Create guardrail opens a 5-step wizard. The steps you see depend on what you enable as you go — skip the denylist check and its step skips too.
Choose the agent
Which agent's traces this guardrail scans.
Choose checks, and a mode for each
Turn on any of the three checks (PII, Secrets & credentials, Denylist) and set each one to Flag or Log. Nothing is saved yet — this step is still a draft.
Add denylist terms (only if Denylist is on)
Type a term or phrase, mark it a regex if it is one, and choose its scope — this agent only, or every agent in the tenant. Each Add takes effect immediately, on its own; it isn't part of the final Save.
Turn on auto-pause, optionally
Off by default. Turn it on and set Pause after a number of findings to have the agent stop accepting calls once its Flag-mode finding count reaches that number — Log-mode findings never count toward it.
Review, then save
A table shows exactly what goes live — each check's enabled/disabled state and mode, the denylist term count already saved, and the auto-pause setting. Saving applies the check policy and auto-pause setting together; denylist terms are already live from step 3.
Editing an already-configured agent (Edit checks on its row) opens the same wizard straight on the checks step, skipping agent selection.
Reading findings
Selecting a row's finding count opens a paginated table of every recent match for that agent: when it happened, which check caught it, whether it was in the input or the output, its mode, the matched value (masked, never shown in full), and a link to open that call's trace.
Resume a paused agent
Auto-pause stops new calls, not the agent's row. Find it in the table — a paused agent's name carries a red In review — paused line right under it, separate from the Scanning column (which still just reads On/Off).
Open its row menu and choose Resume from review. That resets its finding count against the threshold to zero — the total findings figure itself is untouched, and none of what piled up while paused is deleted or retroactively uncounted, it simply stops blocking new calls.
Two independent switches
An agent's guardrail state has two independent axes:
- Scanning — whether the poller looks at this agent's traces at all.
- Check policy — which of the three checks are on, and in what mode.
An agent can have scanning on with every check off (nothing gets flagged, but you'd catch a misconfiguration if you looked), or scanning off with checks configured underneath (nothing runs, but the setup survives for when you flip it back on).
Each row's ⋮ menu, shown above, turns either axis off without touching the other, on top of the same Edit checks wizard used to set them up: Disable agent scanning stops the poller for this agent while its checks stay configured, and Disable auto-pause leaves findings still getting flagged but never pauses the agent. Both read Disable only while the respective setting is on; once turned off, the same slot flips to Enable agent scanning / Enable auto-pause.
Denylist terms
Terms are added and removed immediately, one at a time — each Add or Remove call takes effect on its own, not batched into a save you might forget to click. A term is either this agent only or every agent in this tenant, and either a plain phrase or a regex.
Auto-pause
Set a threshold on an agent, and once its Flag-mode finding count reaches it, the agent is paused: its calls are refused until someone reviews it and resumes it. Off by default.
The count is derived live from the Flag-mode findings recorded since the last reset, not a stored counter — recomputing on every read means there is nothing to drift out of sync.
Resuming an agent from review resets the count to zero. Turning auto-pause back on after it was off does not retroactively count whatever piled up while it was off.
Scan limits
- The poller runs on an interval, not inline with the call — a finding, and any auto-pause it triggers, typically lands within about a minute of the call, not immediately.
- Scanned text is capped at roughly 50 KB of prompt plus completion combined.
A finding from a truncated scan is stamped
scanned_truncated. - A hard timeout wraps the whole scan for one trace, not each check individually.
- A check that raises an exception is caught and logged as a failure for that one check. It never takes the other checks — or the poller — down with it.

