> ## Documentation Index
> Fetch the complete documentation index at: https://docs.mnemom.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Enforcement Modes

> Configure how the Mnemom Gateway responds to alignment and integrity violations

The Mnemom Gateway supports three enforcement modes that control how the system responds when alignment or integrity violations are detected. You can choose the level of intervention appropriate for your use case — from passive observation to active blocking.

## Modes overview

<CardGroup cols={3}>
  <Card title="Observe" icon="eye">
    Detect violations, record them, take no action. This is the **default mode**. Ideal for initial deployment and monitoring.
  </Card>

  <Card title="Nudge" icon="comment">
    Detect violations and inject feedback into the agent's next request via system prompt. The agent sees it and can self-correct.
  </Card>

  <Card title="Enforce" icon="ban">
    Gates same-turn on both streaming and non-streaming requests (HTTP 403 for non-streaming); on streaming this adds latency because the response is evaluated before it is delivered.
  </Card>
</CardGroup>

## Mode details

<Tabs>
  <Tab title="Observe">
    ### Observe mode (default)

    In observe mode, the gateway detects and records all violations but takes no action to modify agent behavior. This is the default mode for all new agents.

    **Behavior:**

    * All API calls pass through unchanged
    * Violations are detected and recorded in the trace database
    * Integrity checkpoints are created for every interaction
    * Drift alerts are generated when behavioral patterns shift
    * No modification to agent requests or responses

    **When to use:**

    * During initial deployment to establish behavioral baselines
    * When you want to monitor without affecting agent behavior
    * For compliance auditing where you need a record but not intervention
    * When evaluating whether to enable more active enforcement

    **Configuration:** set `integrity_mode: observe` (and/or `autonomy_mode: observe`) on the agent's [alignment card](/specifications/alignment-card-schema). See [Setting enforcement mode](#setting-enforcement-mode) below for the three paths (dashboard / CLI / API).
  </Tab>

  <Tab title="Nudge">
    ### Nudge mode

    In nudge mode, when a violation is detected, the gateway creates a pending nudge record. On the agent's next request, the gateway injects an integrity notice into the system prompt. The agent sees the notice, can review its approach, and self-correct.

    **Behavior:**

    * Violations are detected and recorded
    * A pending nudge record is created for each violation
    * On the next API request, an integrity notice is injected into the system prompt
    * The agent receives the nudge and can self-correct
    * Nudge delivery is tracked in the conscience timeline
    * The original request passes through (not blocked)

    **How nudges work:**

    1. The gateway detects a violation via background streaming analysis (\~3s latency)
    2. A pending nudge record is created immediately — ready before the agent's next turn
    3. On the agent's next API call, the gateway injects a system prompt addition:
       ```
       [INTEGRITY NOTICE] Your previous action was flagged for:
       UNBOUNDED_ACTION (HIGH) — action "delete_all_records" is not
       in your declared bounded_actions. Please review your approach.
       ```
    4. The agent processes the notice alongside its normal instructions
    5. The nudge is marked as delivered in the conscience timeline

    **When to use:**

    * When you want the agent to self-correct without hard blocking
    * For agents in production where availability is critical
    * When building trust in the enforcement system before enabling hard blocks
    * For `advisory` or `delegated_authority` principal relationships

    **Configuration:** set `integrity_mode: nudge` (and/or `autonomy_mode: nudge`) on the agent's [alignment card](/specifications/alignment-card-schema). See [Setting enforcement mode](#setting-enforcement-mode) below.
  </Tab>

  <Tab title="Enforce">
    ### Enforce mode

    In enforce mode, the gateway actively blocks when critical violations are detected. It gates the same turn on both streaming and non-streaming requests — the guarantee is identical; only the latency differs.

    **Behavior:**

    * **Non-streaming requests**: Hard block with HTTP 403 response. The violating response is not delivered.
    * **Streaming requests**: The response is evaluated before it is delivered, so a violating response is gated the same turn. This adds latency on streaming, but preserves the same-turn guarantee — enforce does **not** fall back to advisory nudging.

    **When to use:**

    * For high-security environments where violations must be prevented
    * When operating under `fail_closed` AIP failure policy
    * For agents handling sensitive data or financial transactions
    * When regulatory requirements demand active intervention

    **Configuration:** set `integrity_mode: enforce` (and/or `autonomy_mode: enforce`) on the agent's [alignment card](/specifications/alignment-card-schema). See [Setting enforcement mode](#setting-enforcement-mode) below.

    <Warning>
      Enforce mode gates both streaming and non-streaming API calls when violations are detected — HTTP 403 for non-streaming, and same-turn gating before delivery for streaming (which adds latency). Ensure your application handles 403 responses gracefully before enabling this mode.
    </Warning>
  </Tab>
</Tabs>

## Grace period (read this before enabling enforce)

Enforce mode does **not** block every violation immediately. By default, the policy engine grants a **24-hour grace period** to any tool the agent has not been seen using before — during that window, what would otherwise be a `fail` verdict is silently downgraded to a `warn`, the gateway forwards the request, and the model gets to call the tool. This is intentional ergonomics for production rollouts (operators have a day to amend the alignment card after adding a new MCP server) but it has a real security trade-off you must understand.

<Warning>
  **Brand-new tools — including ones an attacker introduces — get a 24h pass on enforce by default.**

  If an attacker injects an unmapped tool through prompt injection, an upstream MCP server compromise, or a tool-name overlap, the gateway will let the call through for the first 24 hours after it's first observed for that agent. Mature agents with stable tool inventories aren't exposed; agents that add tools dynamically, run untrusted MCP servers, or accept tool definitions from user input absolutely are.

  If your threat model includes adversarial tool introduction, set `grace_period_hours: 0` on the agent's [alignment card](/concepts/agent-cards) to disable the grace path entirely.
</Warning>

The setting lives in `enforcement.grace_period_hours` on the alignment card and accepts any non-negative number. `0` disables grace and makes enforce strict from the moment a card is published.

```yaml theme={null}
# Strict enforce — every unmapped/forbidden tool blocks immediately,
# from the first time it's seen.
enforcement:
  grace_period_hours: 0
```

```yaml theme={null}
# Default enforce — 24h grace per new tool. Operationally forgiving;
# weakest under adversarial tool introduction.
enforcement:
  grace_period_hours: 24
```

```yaml theme={null}
# Long grace — deploy-on-weekly-cycles teams that need extra time
# to update cards. Wider exposure window; only appropriate for
# trusted environments.
enforcement:
  grace_period_hours: 168   # 7 days
```

| Use case | Recommended `grace_period_hours` |
| - | - |
| Adversarial threat model (untrusted user input, untrusted MCP servers, agents accepting tool defs from users) | **`0`** |
| Test harness / CI matrix | `0` |
| Production with stable tool inventory + mature card | `0` (no operational benefit when the tool list is stable) |
| Production with frequent tool additions, trusted operators | `24` (default) |
| Slow-cadence operational teams | `48`–`168` |

The grace decision is per-(agent, tool) and is tracked in the platform's `tool_first_seen` table — first observation timestamps the row, subsequent calls within the window pass, calls after the window enforce. There is no API to back-date a `tool_first_seen` row, so disabling the grace period is the only way to make enforce strict from the moment a card is published. See [Policy Engine § Grace period](/concepts/policy-engine#grace-period) for the full mechanism.

## Setting enforcement mode

Enforcement mode is a top-level field on the alignment card — `autonomy_mode` for the action-policing pipeline (CLPI tool-use policy + bounded actions) and `integrity_mode` for the values/conscience pipeline (AIP + drift). The legacy `PUT /v1/agents/{id}/enforcement` endpoint was retired 2026-05-14.

Three paths, pick the one that fits your workflow:

**Dashboard.** Open `https://mnemom.ai/dashboard/agents/{your-agent-id}/card`, toggle `autonomy_mode` and/or `integrity_mode`, save. Easiest path for one-off changes.

**CLI.**

```bash theme={null}
mnemom card edit    # opens current alignment-card YAML in $EDITOR
# change autonomy_mode and/or integrity_mode (off | observe | nudge | enforce)
# save and exit; CLI publishes + triggers recompose
```

**Programmatic.** [`PUT /v1/alignment/agent/{agent_id}`](/api-reference/endpoint/put-alignment-agent-agent-id) with the full canonical card, or [`PUT /v1/alignment/agent/{agent_id}/modes`](/api-reference/endpoint/put-alignment-agent-agent-id-modes) to update just `autonomy_mode`/`integrity_mode` without a full read-modify-write. See [Card Management](/guides/card-management) for the read-modify-write flow and the [alignment-card schema](/specifications/alignment-card-schema) for the field requirements.

## Nudge strategy

When enforcement mode is `nudge` or `enforce`, you can further control *when* nudges are created using the nudge strategy setting:

| Strategy | Behavior |
| - | - |
| `always` | Every boundary violation creates a nudge (default) |
| `sampling` | Nudge on a percentage of violations (uses `proof_rate` or dedicated `nudge_rate`) |
| `threshold` | Only nudge after N violations in the current session |
| `off` | No nudging — violations are recorded but no correction is injected |

**Configuration:**

```bash theme={null}
curl -X PUT https://api.mnemom.ai/v1/agents/:id/settings \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"nudge_strategy": "always"}'
```

<Tip>
  Use `threshold` mode to avoid alert fatigue. The agent only receives a nudge after repeated violations in the same session, giving it a chance to self-correct naturally first.
</Tip>

## Per-agent feature toggles

<Note>
  Team operations require the `team_reputation` entitlement, enabled on enterprise agreements. See [Pricing](/pricing/overview) for how plans and entitlements work.
</Note>

Each agent has independent controls for the transparency and integrity pipeline:

| Setting | Type | Default | Description |
| - | - | - | - |
| `aap_enabled` | boolean | `true` | Enable AAP action traces |
| `aip_enabled` | boolean | `true` | Enable AIP integrity analysis |
| `proof_enabled` | boolean | `true` | Enable cryptographic attestation (Ed25519 + Merkle) |
| `proof_rate` | integer | `100` | % of non-violation checkpoints that get full attestation |
| `proof_boundary_cap` | integer | `100` | Max % of boundary violations to prove (when card-gap detection is unavailable). Boundary violations with detected card gaps are automatically [deferred](/protocols/aip/verifiable-verdicts#deferred-proofs). |
| `nudge_strategy` | string | `always` | When to create nudges (`always`, `sampling`, `threshold`, `off`) |

**Configuration:**

```bash theme={null}
# Disable AIP for an agent (AAP traces still flow)
curl -X PUT https://api.mnemom.ai/v1/agents/:id/settings \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"aip_enabled": false}'

# Set proof sampling to 50%
curl -X PUT https://api.mnemom.ai/v1/agents/:id/settings \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"proof_rate": 50}'
```

These settings are also available in the **Agent Settings** panel on the web dashboard for claimed agents.

## Violation types and enforcement

Enforcement applies to all violation types detected by the AAP verification engine, AIP integrity checks, and the policy engine:

### Alignment violations

| Violation Type | Severity | Enforcement Behavior |
| - | - | - |
| `FORBIDDEN_ACTION` | CRITICAL | Blocked in enforce mode; nudged in nudge mode |
| `CARD_MISMATCH` | CRITICAL | Blocked in enforce mode; nudged in nudge mode |
| `UNBOUNDED_ACTION` | HIGH | Blocked in enforce mode; nudged in nudge mode |
| `MISSED_ESCALATION` | HIGH | Blocked in enforce mode; nudged in nudge mode |
| `CARD_EXPIRED` | HIGH | Blocked in enforce mode; nudged in nudge mode |
| `UNDECLARED_VALUE` | MEDIUM | Nudged in nudge/enforce mode (not blocked) |

### Policy violations

The policy engine checks tool usage against the card's `capabilities` + `enforcement` sections, but it does **not** have its own independent enforce switch — whether a violation is logged or blocks the request is governed by the same `autonomy_mode` field documented above (`off`/`observe`/`nudge` all evaluate-and-log without blocking; `nudge` and `observe` behave identically for policy today; `enforce` blocks on `critical`/`high` violations).

| Violation Type | Severity | Enforcement Behavior |
| - | - | - |
| Forbidden-tool match | Per-rule (`critical`/`high`/`medium`/`low`) | Blocked in `autonomy_mode: enforce` when severity is `critical`/`high`; logged otherwise |
| Unmapped tool, `allow_unmapped_tools: false` (default) | `high` | Blocked in `autonomy_mode: enforce`; logged in `observe`/`nudge` |
| Unmapped tool, `allow_unmapped_tools: true` | `low` | Never blocks; informational only |

The `X-Policy-Verdict` response header is always present when the canonical card carries a policy section:

| Header Value | Meaning |
| - | - |
| `pass` | All tools mapped and permitted |
| `warn` | Violations detected but not blocking |
| `fail` | Violations detected and request blocked (`autonomy_mode: enforce` only) |

See [Policy Engine](/concepts/policy-engine) for full details on how policies are evaluated, including the `allow_unmapped_tools` field and the grace-period mechanics above.

<Note>
  In `enforce` mode, only `CRITICAL`/`critical` and `HIGH`/`high` severity violations trigger a hard block. `MEDIUM`/`medium` (and lower) severity violations are logged/nudged instead, even in enforce mode. This applies to both alignment violations and policy violations.
</Note>

## Safe House front-door enforcement

The modes above govern the agent's *own* behaviour — alignment, integrity, and policy. [Safe House](/concepts/safe-house) is a third, parallel enforcement system, and it acts on the request on the way **in** rather than on the response on the way out. Its enforcement point is worth stating explicitly here, because it is the one place where the gateway intervenes between the provider response and the next inbound request.

**The front door does not fire once per turn.** It fires once per inbound surface a request carries. When `screen_surfaces.tool_responses` is enabled on the [protection card](/concepts/protection-card), a request that hands tool results back to the model is screened for the message *and* for each of those tool results — synchronously, in the request path, before anything is forwarded upstream.

Enforcement on a flagged tool result is turn-internal and does not fail the request:

| Per-tool verdict | Enforcement action, applied to the request body before dispatch |
| - | - |
| `pass` | Tool result forwarded unchanged. |
| `warn` | Tool result **decorated** — delivered to the model, annotated as untrusted. |
| `quarantine` / `block` | Tool result **withheld** — content replaced with a quarantine notice; the model never sees the payload. The rest of the request still goes through, so the caller gets a 200 with `front=enforced` on `X-Mnemom-Verdict`, not a 403. |

This is a front-door check, not an AIP check, and it is not deferred to the agent's next turn. AIP (the `integrity_mode` pipeline documented above) runs afterwards on the model's *reasoning*; by then the model has already read whatever the front door let through.

## Response headers

The gateway reports every enforcement decision on the response, whether or not anything fired:

| Header | Meaning |
| - | - |
| `X-Mnemom-Verdict` | Structured, present on every response except a long non-streaming one the gateway committed early (see [Quickstart: response headers](/quickstart/gateway)): `front=<v>; autonomy=<v>; integrity=<v>; back=<v>`, one entry per phase — Safe House front door, action-policing (policy + bounded actions), AIP integrity, Safe House back door. Each `<v>` is one of `pass`, `observed`, `nudged`, `enforced`, `unverified`. Absent phases default to `pass`. |
| `X-Mnemom-Advisory` | A JSON array of `{source, text, severity?, id?}` entries surfacing what a `nudged`/`enforced` phase actually did, capped at 5 entries. Omitted entirely when there's nothing to advise. |
| `X-Mnemom-Request-Id` | A per-request UUID for support correlation. |
| `X-AIP-Verdict` + `X-AIP-Checkpoint-Id` | The two AIP-specific headers that still exist standalone alongside the structured `X-Mnemom-Verdict.integrity` field. |
| `X-Policy-Verdict` | `pass` / `warn` / `fail` from the policy engine specifically — see [Policy violations](#policy-violations) above. |
| `X-Mnemom-Effective-Mode` | Present only when a turn ran in a different mode than configured. Today that is `passthrough; reason=balance_depleted; configured=enforce`: the agent's µ balance was depleted, so the turn was forwarded without enforcement or analysis, and its `X-AIP-Verdict: clear` / `integrity=pass` do not mean the turn was checked. |
| `X-Mnemom-Deferred` | On a long non-streaming response committed early: names what the early headers could not carry. `X-AIP-Verdict` is `pending` and `X-Mnemom-Verdict` / `X-Policy-Verdict` are absent on such a response. |

<Note>
  A number of legacy per-phase headers — `X-Mnemom-Autonomy-Verdict`, the `X-Safe-House-*` family (`Verdict`, `Quarantine-Id`, `Session-Risk`, `Mode`, `Simulated-Verdict`, `Canary-Triggered`, `DLP`, `Advisory`, `Event`), and most of the old `X-AIP-*` family (`Action`, `Proceed`, `Synthetic`, `Source`, `Analysis-Scope`, `Certificate-Id`, `Chain-Hash`, `Reason`, `Nudge-Count`, `Enforcement`) — were retired in a clean break in favor of the structured `X-Mnemom-Verdict`/`X-Mnemom-Advisory` pair above. The gateway actively strips any of these names from every outbound response as defense-in-depth. If you integrated against one of them before this cutover, migrate to the fields above — the old names will never reappear.
</Note>

For the front-door table above, a `pass`/`warn` per-tool verdict never raises `front` past `observed`; a `quarantine`/`block` verdict is what produces `front=enforced` on `X-Mnemom-Verdict`.

## Conscience timeline

All enforcement actions are tracked in the conscience timeline, accessible via the API and the web dashboard at [mnemom.ai](https://mnemom.ai). The timeline records:

* When a violation was detected
* What type and severity
* What enforcement action was taken (observed, nudged, blocked)
* Whether the agent self-corrected after a nudge
* Drift patterns across enforcement events

## Provider compatibility

Enforcement works across all providers where AIP is supported:

| Provider | Observe | Nudge | Enforce |
| - | - | - | - |
| Anthropic | Yes | Yes | Yes |
| OpenAI | Yes | Yes | Yes |
| Gemini | Yes | Yes | Yes |

Enforce gates same-turn on both streaming and non-streaming across all three providers; on streaming this adds latency because the response is evaluated before it is delivered.

## See also

* [Mnemom Gateway Overview](/gateway/overview) — Architecture and how it works
* [CLI Reference](/gateway/cli) — CLI commands for monitoring enforcement
* [Observability](/guides/observability) — Export enforcement data to your observability platform
* [Safe House](/concepts/safe-house#when-the-front-door-runs) — When the front door runs, and why it runs more than once per turn

***

## Agent containment

<Note>
  Agent containment is a separate enforcement layer that operates above the per-request enforcement modes. While enforcement modes (observe/nudge/enforce) control individual request handling — including the per-surface front-door screening described above, which can fire several times within one request — containment controls whether the agent can make requests at all.
</Note>

### Containment states

| State | Meaning | Gateway Behavior |
| - | - | - |
| `active` | Normal operation (default) | Requests proceed normally |
| `paused` | Temporarily stopped | All requests blocked with HTTP 403 |
| `killed` | Permanently stopped | All requests blocked with HTTP 403 |

**Paused** agents can be resumed by an org owner or admin. **Killed** agents require explicit reactivation by an owner only. The distinction matters for audit: pause means "we need to investigate," kill means "this agent is compromised."

### Containment API

```bash theme={null}
# Pause an agent
curl -X POST https://api.mnemom.ai/v1/orgs/{org_id}/agents/{agent_id}/pause \
  -H "Authorization: Bearer $TOKEN" \
  -d '{"reason": "Investigating boundary violations"}'

# Resume a paused agent
curl -X POST https://api.mnemom.ai/v1/orgs/{org_id}/agents/{agent_id}/resume \
  -H "Authorization: Bearer $TOKEN"

# Kill an agent (owner only)
curl -X POST https://api.mnemom.ai/v1/orgs/{org_id}/agents/{agent_id}/kill \
  -H "Authorization: Bearer $TOKEN" \
  -d '{"reason": "Agent compromised"}'

# Reactivate a killed agent (owner only)
curl -X POST https://api.mnemom.ai/v1/orgs/{org_id}/agents/{agent_id}/reactivate \
  -H "Authorization: Bearer $TOKEN"

# Get containment status and audit log
curl https://api.mnemom.ai/v1/orgs/{org_id}/agents/{agent_id}/containment \
  -H "Authorization: Bearer $TOKEN"
```

### Error response

When a contained agent attempts an API request through the gateway, it receives:

```json theme={null}
{
  "error": {
    "code": "containment_error",
    "message": "Agent contained",
    "details": { "reason": "agent_paused" }
  }
}
```

`details.reason` is `agent_paused` or `agent_killed` depending on the containment state. The HTTP status code is `403 Forbidden` (distinct from `402 Payment Required` used for billing enforcement — see [Errors](/api-reference/errors#402-payment-required--billing-and-depletion)).

### Auto-containment

Agents can be configured to automatically pause after consecutive boundary violations:

```bash theme={null}
# Enable auto-containment after 3 consecutive violations
curl -X PUT https://api.mnemom.ai/v1/orgs/{org_id}/agents/{agent_id}/containment-policy \
  -H "Authorization: Bearer $TOKEN" \
  -d '{"auto_containment_threshold": 3}'

# Disable auto-containment
curl -X PUT https://api.mnemom.ai/v1/orgs/{org_id}/agents/{agent_id}/containment-policy \
  -H "Authorization: Bearer $TOKEN" \
  -d '{"auto_containment_threshold": null}'
```

When auto-containment triggers, it:

* Sets the agent status to `paused` with actor `system`
* Logs the action in the containment audit log
* Emits an `agent.paused` webhook event
* Purges the gateway cache so the block takes effect immediately

### RBAC requirements

| Action | Required Role |
| - | - |
| Pause | `owner` or `admin` |
| Resume | `owner` or `admin` |
| Kill | `owner` only |
| Reactivate | `owner` only |
| View status | Any org role |
| Set auto-containment policy | `owner` or `admin` |

### Webhook events

Three webhook events are emitted for containment actions:

* `agent.paused` — Agent was paused (manually or automatically)
* `agent.resumed` — Agent was resumed or reactivated
* `agent.killed` — Agent was killed

Each event includes:

```json theme={null}
{
  "agent_id": "mnm-550e8400-e29b-41d4-a716-446655440000",
  "org_id": "org-xxx",
  "action": "pause",
  "actor": "user-xxx",
  "reason": "Investigating boundary violations",
  "previous_status": "active",
  "new_status": "paused"
}
```

### Containment audit log

Every containment action is recorded in a tamper-evident audit log. Each entry includes:

* The action taken (pause, resume, kill, reactivate, auto\_pause)
* Who triggered it (user ID or `system` for auto-containment)
* The reason provided
* Previous and new containment states
* Timestamp


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.