When to simulate
- Before a manifest change goes live — verify the new spec accepts the calls you want and rejects the ones you don’t.
- Before an agent prompt change — check whether a new tool the agent might call would clear policy.
- During post-incident review — replay a problematic tool call to understand why the gateway flagged it.
- In CI — wire simulate into your test suite to gate spec changes on hypothetical-call coverage.
The request shape
Two body fields, either or both:candidate_tool_call is set without candidate_input, the platform synthesizes an assistant message carrying the toolUses block. If both are set, the explicit candidate_input.messages wins for what gets evaluated as conversation context, but tool identifiers are collected from both: candidate_tool_call.tool_name and any candidate_input.messages[].toolUses[].name entries.
The response shape
tools_evaluated counts how many caller-supplied tool identifiers were actually checked. gateway_decision.verdict and observer_assessment.verdict are the raw per-context engine output, and the engine only produces a finding while iterating over that list — so when tools_evaluated is 0, both verdicts read "pass" with no violations because nothing was checked, not because anything was approved. allowed is the field to trust; it’s derived to fail closed on that case (see below).
The allowed verdict
allowed is one of three values:
The derivation table:
Common patterns
Test a manifest before committing
When iterating on a new manifest, PUT it to a sandbox agent, simulate the calls you expect the agent to make, then promote to production once they all pass.Discover required receipts
The gateway hooks bind on catalog values likepolicy_attentiveness and deliberation_before_action. They require specific consultation receipts (typically think) before invoking certain tool classes. Simulate surfaces the requirement:
think consultation to the agent’s prompt scaffolding, then re-simulate to confirm the receipt closes the requirement.
Wire simulate into CI
Themnemom/cards-action drives simulate per-PR — every spec change runs a configurable set of hypothetical calls and posts the verdicts as a PR comment. You can also wire it manually:
Pure-sync only
Simulate calls@mnemom/policy-engine directly twice (once with context: "gateway", once with context: "observer") using dryRun: true. It never crosses process boundaries — no real gateway state change, no observer ledger entry. The evaluation is deterministic for a given (spec, candidate) pair.
Protection scope
POST /v1/protection/agent/<id>/simulate exists so the URL surface is symmetric with alignment, but the dedicated protection-policy evaluator isn’t shipped yet: it always returns allowed: "conditional" with a protection_evaluator_pending_phase_5 condition, gateway_decision/observer_assessment both verdict: "warn", and tools_evaluated: 0 — it never actually calls the policy engine. The protection card itself still composes through the same scope cascade as alignment, surfacable via GET /v1/protection/agent/<id>/effective.
Rate limits + cost
Simulate is pure-sync — no LLM call, no rate-limit budget consumed. You can simulate freely in CI loops, test suites, and pre-commit hooks.Related reading
- AI helpers — overview of the four AI-forward verbs.
- Explain and remediate — for understanding violations on the current spec.
- Sub-resource verbs — for applying the fix once simulate identifies what’s needed.
- Policy engine — the evaluator simulate drives.