> ## Documentation Index
> Fetch the complete documentation index at: https://docs.mnemom.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Protection Card

> The YAML document that tells Safe House how to defend an agent at runtime. Mode, thresholds, screen surfaces, trusted sources — all edited, composed, and audited like an alignment card.

The **protection card** is the YAML document that tells [Safe House](/concepts/safe-house) how to defend an agent at runtime. It is one of the [two cards](/concepts/agent-cards) every Mnemom agent has — the alignment card declares *who the agent is*; the protection card declares *how the agent is defended*.

The protection card elevates what used to be an ad-hoc Safe House JSON config into a first-class YAML card with composition across up to four scopes (platform > org > team *(optional)* > agent), granular exemptions, and the same audit + amendment semantics as the alignment card.

## Structure

A protection card has four top-level sections. The complete schema is at [/specifications/protection-card-schema](/specifications/protection-card-schema); this page describes each section's purpose.

### `mode`

Top-level action policy for Safe House on this agent. Shares the four-mode enum (`off | observe | nudge | enforce`) with the alignment card's two master switches (`autonomy_mode` and `integrity_mode`) — same words, same semantics, same UI picker.

* **`off`** — detection skipped entirely (cost / latency / non-applicability cases). No telemetry.
* **`observe`** — all detectors run, signals are logged asynchronously to the trace + reputation pipeline, no action is taken on the agent's request.
* **`nudge`** — detectors run synchronously; matches attach an advisory annotation to the agent's prompt context (and an `X-Mnemom-Advisory` response header — see the [Headers reference](/api-reference/headers)) but the request proceeds. The model sees the advisory as part of its context.
* **`enforce`** — detectors run synchronously; matches block the request (quarantine ≥ quarantine threshold, hard block ≥ block threshold).

`enforce` is implicitly synchronous — to block a request, the gateway must wait for the verdict. There is no separate `enforce_sync` mode.

**Composition: strictest wins** across `enforce > nudge > observe > off`. An agent cannot drop below the platform/org floor.

```yaml theme={null}
mode: enforce
```

[AEGIS Managed Rules](/concepts/managed-rules) add detection content to every gateway. When a reviewed rule is promoted at platform scope, the gateway loads it from the Ed25519-signed rule set, after verifying the signature, and adds the rule's detection thresholds to the screening pipeline your protection card configured. Your card's `mode` (`off / observe / nudge / enforce`) determines what action the gateway takes when a Managed Rule matches.

### `thresholds`

Three-band escalation ladder for Safe House detector scores. All values are floats in `[0, 1]`.

```yaml theme={null}
thresholds:
  warn: 0.60         # informational annotation in observe/nudge; soft annotation in enforce
  quarantine: 0.80   # quarantine in enforce; informational in observe/nudge
  block: 0.95        # hard block in enforce
```

Three bands map cleanly onto the SOC severity ladder familiar to most operators. Per-detector tuning is an internal calibration concern and not exposed in the schema. Validation enforces `warn ≤ quarantine ≤ block`.

### `screen_surfaces`

Which request surfaces Safe House inspects:

* `incoming` — the user/principal prompt entering the agent
* `outgoing` — the agent's response leaving the agent
* `tool_calls` — tool-use invocations the agent makes
* `tool_responses` — responses to those tool calls, screened at the front door **in the request that carries them back to the model**, before that request is forwarded upstream

`incoming` and `tool_responses` are both front-door surfaces, and the front door runs on each of them independently — enabling `tool_responses` means the detector pipeline runs again, inside the same request, once per tool result. See [When the front door runs](/concepts/safe-house#when-the-front-door-runs).

```yaml theme={null}
screen_surfaces:
  incoming: true
  outgoing: true
  tool_calls: true
  tool_responses: false      # don't inspect tool return values
```

Default is all-true (scan everything). Turning off a surface is a deliberate performance/coverage trade-off and is logged in every detection event so auditors can see what was *not* inspected.

### `trusted_sources`

Per-bucket allowlist of upstream sources whose content Safe House skips detection for. Buckets are typed so the validator can apply per-bucket deny-lists and the composer can apply per-bucket intersection rules.

```yaml theme={null}
trusted_sources:
  domains:
    - internal.mnemom.ai
    - vendor-api.example.com:8080
  agent_ids:
    - mnm-aabbccdd-eeff-0011    # agent-to-agent pass-through
  ip_ranges:
    - 10.0.0.0/8                # RFC1918 internal space
```

The validator deny-lists public LLM endpoints (`api.openai.com`, etc.), public DNS resolvers (`8.8.8.0/24`, `1.1.1.0/24`), and any-host CIDRs (`0.0.0.0/0`, `::/0`). Adding a publicly-routable IP range or a customer-controllable domain is a critical misconfiguration even if it passes the deny-list.

Trusted sources cause Safe House to skip detection (no detector cycles spent), and every match emits an `sh_trusted_source_skip` audit trace so reviewers can see what was waved through.

## Composition across scopes

Like the alignment card, the protection card composes across up to four scopes — platform > org > team *(optional)* > agent:

| Section | Composition rule |
| - | - |
| `mode` | **Strictest wins** (`enforce > nudge > observe > off`). An agent cannot drop below the platform/org/team floor. |
| `thresholds.*` | **Min across scopes** — lowest = strictest wins. An agent can tighten further than the platform/org/team but not loosen. |
| `screen_surfaces.*` | **OR per field — true wins.** If any scope requires scanning a surface, it's scanned. |
| `trusted_sources` | **Platform ceiling** (intersection); **org + every team + agent union** within that ceiling. No downstream scope can widen trust beyond what the platform allows. |

See [Card Composition](/concepts/card-composition) for the full rules + worked examples.

### Exemptions

A protection-card exemption waives a specific threshold or surface for a specific agent with a stated reason, expiry, and audit trail.

```yaml theme={null}
# Granted via:
#   POST /v1/agents/{agent_id}/exemptions
# Body:
#   exempt_section: "protection.thresholds.canary_match"
#   reason: "Agent requires canary in prompts for debugging red-team scenarios"
#   granted_by: "org-admin@example.com"
#   expires_at: "2026-07-17T00:00:00Z"
```

Exemptions replace the legacy boolean `org_card_exempt` flag — which waived the whole card at once — with section-specific, audit-logged, time-bounded grants.

## How the protection card is used

### Gateway (request path)

Every request hits:

1. `canonical_protection_cards` read (cached, 5-min TTL) to get the composed protection card.
2. Detectors inspect enabled `screen_surfaces` with the composed `thresholds`. This is **per surface, not per request** — the inbound message runs the pipeline once, and each tool result carried in the same request body runs it again on its own `tool_result` surface.
3. `mode` determines the action: observe (log only), nudge (advisory), enforce (block, or — for a flagged tool result — withhold it from the body and forward the rest).
4. `trusted_sources` short-circuits detectors for allowlisted upstreams (with trace entry).

Step 3's rewrites land on the outbound body *before* it is forwarded to the provider. Nothing here is deferred to a background pass or to the agent's next request.

The protection card is never fetched from the agent-scope row on the request path — it's always the canonical (pre-composed) version.

### Observer (trace analysis)

The observer pipeline does **not** read the protection card. Protection is inline at the gateway; the observer's role is to reconcile and enrich traces after the fact. Detector signals produced at the gateway are written into the trace itself.

### Website + CLI

```bash theme={null}
mnemom protection show                         # canonical protection card (YAML)
mnemom protection edit                         # open in $EDITOR
mnemom protection publish protection.card.yaml # validate + publish + recompose
mnemom protection validate protection.card.yaml
```

The website agent detail → Security tab shows both raw agent-scope and canonical-composed views.

## Modes across both cards

All three master switches — the alignment card's `autonomy_mode` and `integrity_mode`, and the protection card's `mode` — share the same four-mode enum and the same UI picker:

| Value | Alignment (`integrity_mode`) | Protection (`mode`) |
| - | - | - |
| `off` | Skip AIP + drift detection | Detection skipped entirely |
| `observe` | Integrity checkpoints run; violations logged | Safe House detectors run; signals logged |
| `nudge` | Violations inject advisory on next request | Detectors attach advisory; request proceeds |
| `enforce` | Violations hard-block (auto-pause on boundary) | Detectors may block (quarantine / block) |

A well-run fleet has both dimensions aligned across all agents. The [fleet coherence v2 scorer](/concepts/fleet-coherence) checks integrity uniformity as a first-order structural invariant.

## See also

* [Agent Cards](/concepts/agent-cards) — the two-card model (alignment + protection)
* [Alignment Card (protocol surface)](/concepts/alignment-cards) — the AAP alignment card
* [Protection Card Schema](/specifications/protection-card-schema) — normative YAML schema
* [Card Composition](/concepts/card-composition) — the full platform → org → team → agent composition rules
* [Safe House](/concepts/safe-house) — the runtime detection pipeline this card configures
* [Safe House Threat Model](/guides/safe-house-threat-model) — what Safe House defends against


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.