Skip to main content
When an agent’s spec doesn’t behave the way you expect, the explain endpoint runs the policy engine in dry-run mode against the agent’s composed alignment card and reformats the structured findings into care-framed prose with sub-resource verb hints for each remediation.

What explain tells you

explain evaluates one specific tool call against the agent’s composed card, so the request names the tool: pass request_snippet.tool_name. A request that names no tool is rejected with 400 missing_tool_name — a toolless call has nothing to check, so it isn’t offered a verdict.
The response has three useful sections:
  1. trace — the raw EvaluationResult from @mnemom/policy-engine. Carries verdict (pass / warn / fail), violations[], warnings[], card_gaps[], coverage report, policy_id, policy_version.
  2. reasoning — care-framed prose summarizing the verdict + counts.
  3. suggested_remediations[] — per-finding structured hints. Each carries:
    • for: violation / warning / card_gap
    • index: position in the corresponding trace array
    • remediation: a care-framed prose remediation
    • method + url (optional): the sub-resource verb that would resolve it

The four violation types

The policy engine surfaces four classes of PolicyViolation.type. Each maps to a deterministic remediation template.

forbidden

A tool the agent is using matches a pattern in enforcement.forbidden_tools[].
Two paths forward: edit the forbidden list (if the rule was over-broad), or grant a documented exemption via POST /v1/agents/{agent_id}/exemptions (if the agent legitimately needs this action).

capability_exceeded

The agent invoked a tool that isn’t mapped to any capability’s tools[] list.
Apply the suggested PATCH:
PATCH here is a JSON merge patch (RFC 7396): top-level capability keys you omit are left alone, but an array field you do send (like tools) replaces the existing array wholesale rather than appending to it — include every tool the capability should end up with, not just the new one.

unmapped_denied

The agent used a tool that has no capability mapping, AND enforcement.allow_unmapped_tools: false (the default).
Two surfaces would fix this:
  • The principled path: add the tool to an appropriate capability via PATCH /capabilities.
  • The escape hatch: set enforcement.allow_unmapped_tools: true via PATCH /enforcement. This loosens the default surface — only use it for sandbox / development agents.

gateway_hook_missing_receipt

A tool the agent invoked matches a catalog gateway hook (e.g., policy_attentiveness binds on financial tools), but the required receipt (e.g., think) wasn’t present in the conversation history.
This one isn’t a spec edit — it’s a runtime pattern. The agent’s prompt or scaffolding would benefit from a think consultation before invoking the gated tool. See governance signals and the policy engine for the receipt mechanism.

Warnings

Warnings don’t block; they surface drift the operator may want to address.

Card gaps

card_gaps[] flag capabilities that exist but whose action surfaces aren’t fully declared.

LLM enrichment

If you want long-form prose explanations (useful for non-technical reviewers), pass "enrich": true:
The endpoint calls Claude via the LLM substrate and expands the structured remediations into 4-6 sentences of care-framed prose. The pure-sync path is the floor; the LLM enrichment is additive. LLM failure doesn’t break explain — the structured reasoning always renders. Enrichment consumes the per-principal LLM budget (10/hour/user; 100/hour/org). Off by default.

Hypothetical explain — what would happen?

Explain a hypothetical tool call against the agent’s current composed card by naming the tool:
The request body also accepts turn_id, but the endpoint does not yet use it to load historical spec state or replay a specific past evaluation — every explain call evaluates against the agent’s current composed card. Only request_snippet.tool_name affects the verdict; a request_snippet.tool_args value is accepted but not evaluated.
For “would this specific call pass?” use simulate instead — it runs both gateway and observer contexts and combines them into an allowed-verdict.

The full debug-and-fix flow

  1. Run explain against the agent.
  2. For each suggested_remediation, check the method + url hint.
  3. Apply the sub-resource PATCH with the suggested change.
  4. Re-run explain to confirm the finding is gone.
  5. Repeat until verdict: "pass".

What explain doesn’t do

  • It doesn’t write anything. It’s a read-style POST.
  • It doesn’t replace the simulate endpoint — explain is about the current spec; simulate is about a hypothetical call against that spec.
  • It doesn’t run the protection card’s policy. V1 returns a Phase 5 deferred shape on POST /v1/protection/agent/<id>/explain until the dedicated protection evaluator ships.