> ## Documentation Index
> Fetch the complete documentation index at: https://docs.mnemom.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Protection Card Schema

> Normative YAML schema for the protection card, including every section, field type, required/optional status, and composition semantic.

Normative reference for the **protection card** — the YAML document that configures [Safe House](/concepts/safe-house) for a specific agent, and one half of every Mnemom agent's [two cards](/concepts/agent-cards). This page specifies every section, field, required/optional status, type, and composition semantic.

Conceptual overview: [/concepts/protection-card](/concepts/protection-card). Alignment-card spec: [/specifications/alignment-card-schema](/specifications/alignment-card-schema). Card composition rules across platform/org/agent scopes: [/concepts/card-composition](/concepts/card-composition).

## Top-level structure

```yaml theme={null}
card_version: protection/2026-04-26   # required; string; schema version
agent_id: mnm-<uuid>                  # required; string
expires_at: null                      # optional; protection cards rarely expire

mode: enforce                         # required; see §mode
thresholds: { ... }                   # optional; see §thresholds
screen_surfaces: { ... }              # optional; see §screen-surfaces
trusted_sources: { ... }              # optional; see §trusted-sources
protected_surface: { ... }            # optional to send; composer always emits it — see §protected_surface
review: { ... }                       # optional; see §review
scopes: { ... }                       # optional; see §scopes
extensions: { ... }                   # optional; not composed — see §extensions

# Server-assigned response fields — do not send on PUT:
card_id: pc-<uuid>                    # server-assigned
issued_at: 2026-04-26T12:00:00Z       # response-only
content_hash: sha256:<hex>            # response-only; also the ETag value
version: 4                            # response-only; monotonic card version
_composition: { ... }                 # response-only, and only when `?include_composition=true`; see §composition-metadata
```

`thresholds`, `screen_surfaces` and `trusted_sources` may each be omitted from a raw agent-scope PUT — the composer fills the platform default for whichever bucket is absent. `thresholds` requires all three of `warn`/`quarantine`/`block` when the block is present at all; `screen_surfaces` and `trusted_sources` validate whichever individual keys you send and leave the rest to the composer — see each section below.

## §mode

Top-level action policy for Safe House on this agent. The same `off | observe | nudge | enforce` enum as the alignment card's `autonomy_mode` / `integrity_mode` master switches — see [Master switches](/specifications/alignment-card-schema#master-switches).

```yaml theme={null}
mode: enforce   # "off" | "observe" | "nudge" | "enforce"
```

| Value | Behavior |
| - | - |
| `off` | Detection skipped entirely. No telemetry. For cost / latency / non-applicability cases. |
| `observe` | Detectors run, signals logged asynchronously, no request-path action. |
| `nudge` | Detectors run synchronously; matches attach an advisory annotation to the agent's prompt context (and an `X-Mnemom-Advisory` response header) but the request proceeds. |
| `enforce` | Detectors run synchronously; matches block the request (quarantine ≥ quarantine threshold, hard block ≥ block threshold). |

`enforce` implies synchronous verdict — to block a request, the gateway must wait for the verdict before delivering the message. There is no separate `enforce_sync` mode.

**Composition: strictest wins** across `enforce > nudge > observe > off`. An agent cannot drop below the platform/org floor.

`nudge` is the load-bearing middle ground: the model receives the advisory as part of its prompt context, so the *security signal* reaches the model without blocking the request. Customers running long-tail-confidence detectors typically run `nudge` rather than `enforce` until thresholds settle.

## §thresholds

Three-band escalation ladder for Safe House detector scores. All values are floats in `[0, 1]`.

```yaml theme={null}
thresholds:
  warn: 0.60         # required when thresholds is present
  quarantine: 0.80   # required when thresholds is present
  block: 0.95        # required when thresholds is present
```

| Field | Range | Meaning |
| - | - | - |
| `warn` | `[0, 1]` | Score at-or-above triggers a warn-level annotation in observe/nudge mode (and a soft annotation in enforce). |
| `quarantine` | `[0, 1]` | Score at-or-above triggers a quarantine in enforce mode (message held for review); informational in observe/nudge. |
| `block` | `[0, 1]` | Score at-or-above triggers a hard block in enforce mode. |

**Validation:** `warn ≤ quarantine ≤ block`. The validator rejects any out-of-order combination at write time.

**Composition: min across scopes.** The lowest threshold wins, since lower = stricter (matches sooner). An agent cannot loosen a stricter platform/org threshold; it can only tighten further.

Three bands map cleanly onto the SOC severity ladder familiar to most operators. Per-detector tuning is an internal calibration concern and is not exposed in the schema.

## §screen\_surfaces

Which request surfaces Safe House inspects, named by direction (incoming/outgoing) and tool relationship.

```yaml theme={null}
screen_surfaces:
  incoming: true         # the prompt/message reaching the agent
  outgoing: true         # the agent's generated response
  tool_calls: true       # arguments to tool invocations
  tool_responses: true   # values returned by tools
```

| Field | Default | Meaning |
| - | - | - |
| `incoming` | `true` | Inbound prompts: user messages, webhook triggers, queue messages, API calls — anything entering the agent. |
| `outgoing` | `true` | The agent's generated response leaving the agent. |
| `tool_calls` | `true` | Arguments the agent sends to tool invocations (outbound tool side). |
| `tool_responses` | `true` | Return values from tool calls reaching the agent (inbound tool side). Screened at the front door **in the request that carries them back to the model**, before that request is forwarded upstream — not on the agent's next turn. |

**Surfaces are screening units, not a per-turn budget.** `incoming` and `tool_responses` are both front-door surfaces and the detector pipeline runs on each independently, within the same request: once for the inbound message, once more for each tool result the request carries. See [When the front door runs](/concepts/safe-house#when-the-front-door-runs).

**Validation:** Only the four named keys are accepted. Unknown keys are rejected at write time.

**Composition: OR per field — true wins.** If any scope sets a surface to `true`, it's scanned. Agents cannot disable scanning that org or platform requires. Phrased in alignment-card vocabulary: **strictest wins** (with `true = scan` being the more restrictive choice).

Direction-based naming is durable across transport changes: an agent receiving a webhook trigger is "incoming" whether it's a user message, an API event, or a queue payload. Differentiating `tool_calls` from `tool_responses` reflects that they have different threat models — outgoing tool args may exfiltrate; incoming tool responses may inject. It also reflects a different enforcement shape: an `incoming` finding can fail the request, whereas a `tool_responses` finding is applied *inside* the request — the offending tool result is withheld or decorated in the body and the rest is forwarded.

Turning off a surface emits a low-priority audit trace so reviewers can see what was *not* scanned. If you need to disable a surface for a specific agent, the recommended path is an [exemption](/concepts/card-composition#exemptions) with a documented reason rather than a raw `false` in the agent card.

## trusted\_sources

Per-bucket allowlist of upstream sources whose content Safe House skips detection for. The buckets are typed so the validator can apply per-bucket deny-lists and the composer can apply per-bucket intersection rules.

```yaml theme={null}
trusted_sources:
  domains:
    - internal.acme.com
    - vendor-api.example.com:8080
  agent_ids:
    - mnm-aabbccdd-eeff-0011         # agent-to-agent pass-through
  ip_ranges:
    - 10.0.0.0/8                      # RFC1918 internal space
    - 172.16.0.0/12
```

| Field | Type | Validation |
| - | - | - |
| `domains` | `string[]` | DNS name (or `host:port`); deny-listed against public LLM endpoints (`api.openai.com`, `api.anthropic.com`, etc.) and public DNS-over-HTTPS providers. |
| `agent_ids` | `string[]` | Mnemom agent IDs (`mnm-*` format). No wildcards. |
| `ip_ranges` | `string[]` (CIDR) | IPv4 or IPv6 CIDR; deny-listed against `0.0.0.0/0`, `::/0`, and public DNS resolver ranges (`8.8.8.0/24`, `1.1.1.0/24`, `9.9.9.0/24`). |

**Composition:**

* **Platform → agent: intersection.** The platform list is the compliance ceiling — downstream scopes (org, agent) cannot widen trust beyond what the platform allows. If the platform sets `ip_ranges: [10.0.0.0/8]`, an agent cannot add `192.168.0.0/16` to its own list and have it take effect.
* **Org + agent: union within the ceiling.** Either scope can add trust within the platform-imposed ceiling.
* **Empty platform list = unconstrained ceiling.** When the platform doesn't specify a bucket, downstream entries pass through without intersection.

Trusted sources cause Safe House to skip detection for matching content (no detector cycles spent), but every match emits a low-priority `sh_trusted_source_skip` audit trace so reviewers can see what was waved through.

**Security note:** the validator's deny-list is non-exhaustive — adding a publicly-routable IP range or a customer-controllable domain is a critical misconfiguration even if it passes the deny-list. Treat `trusted_sources` as a sharp tool.

## §protected\_surface

Org-declared assets and mechanical operations the substrate must protect, enforced independent of the agent's own declared intent (the alignment card's `autonomy.forbidden_actions` is advisory; this is enforced). The composer always emits this block — a caller may omit it entirely to inherit the composed floor.

```yaml theme={null}
protected_surface:
  assets:
    - kind: table
      selector: customers
      reason: "PII table"
  forbidden_operations:
    - pattern: "unscoped UPDATE/DELETE"
      applies_to: ["table:customers"]   # empty/absent = GLOBAL (every asset)
      severity: critical
  escalation_required:
    - pattern: "TRUNCATE"
      reason: "Requires human approval"
```

| Field | Type | Required | Composition |
| - | - | - | - |
| `assets[].kind` / `.selector` | string | Yes | Together form the intrinsic identity `${kind}:${selector}`; the composer unions by this identity across Platform → Org → Team → Agent |
| `assets[].label` / `.reason` | string | No | Display-only; resolved by the composer when scopes disagree — not a value to key logic on |
| `forbidden_operations[].pattern` | string | Yes | Intrinsic identity (normalized); **union** across scopes |
| `forbidden_operations[].applies_to` | string\[] of asset identities | No | Empty/absent = GLOBAL; global beats any scoped list on collision |
| `forbidden_operations[].severity` | enum `low\|medium\|high\|critical` | No | **Max across scopes** |
| `escalation_required[].pattern` / `.applies_to` / `.reason` | — | pattern required | Same identity + union rules as `forbidden_operations`, minus severity |
| `*.source_scope` | string | Server-assigned | Do not send on a PUT |

Precedence when a request matches more than one: `forbidden_operations` > `escalation_required` > allow.

## §review

Review-hold policy (Safe House Review). Optional; a raw card may set any subset of fields and the composer fills the rest.

```yaml theme={null}
review:
  enabled: true
  gate_on:
    incoming: quarantine
    outgoing: quarantine
    tool_calls: warn
    tool_responses: warn
    integrity: review_needed       # "off" | "review_needed" | "boundary_violation"
  sla_seconds: 300
  on_timeout: reject                # "reject" | "release"; default "reject" (fail-closed)
  notify:
    sse: true
    webhooks: true
  reviewer:
    kind: builtin_opus               # "builtin_opus" | "endpoint" | "human"
    dual_control: false              # human-tier only: require two distinct resolvers
  quarantine_notice: "This request is held for review."
```

| Field | Type | Notes |
| - | - | - |
| `enabled` | boolean | Required when `review` is present |
| `gate_on.{incoming,outgoing,tool_calls,tool_responses}` | enum `off\|warn\|quarantine\|block` | Minimum Safe House band that escalates to a hold |
| `gate_on.integrity` | enum `off\|review_needed\|boundary_violation` | AIP integrity plane gate |
| `sla_seconds` | number (≥ 1) | Hold timeout |
| `on_timeout` | enum | `reject` (fail-closed) or `release` |
| `notify.sse` / `notify.webhooks` | boolean | Delivery channels for hold events |
| `reviewer.kind` | enum `builtin_opus\|endpoint\|human` | `endpoint` (BYO reviewer) is schema-accepted but not yet consumed by the gateway. `human` routes to `POST /v1/safe-house/reviews/{review_id}/resolve` |
| `reviewer.endpoint_url` | string (https URL) | Required when `kind` is `endpoint` |
| `reviewer.dual_control` | boolean | When `true`, a human-tier resolution requires two distinct resolvers before it finalizes |
| `quarantine_notice` | string (≤ 2000 chars) | No |

**Composition:** strictest-wins per field, same ladder as `mode`.

## §scopes

Agent-declared named scopes used for grounding checks (agent scope only — org and platform scopes do not contribute in the current version).

```yaml theme={null}
scopes:
  finance_qa:
    grounding:
      mode: enforce          # "off" | "shadow" | "enforce"
      corpus_ref: finance-kb
```

| Field | Type | Validation |
| - | - | - |
| `<name>` | object key | `^[a-z0-9][a-z0-9_-]{0,63}$`; at most 32 scope names |
| `<name>.grounding.mode` | enum `off\|shadow\|enforce` | Unrecognized value defaults to `off` (fail closed) at the gateway |
| `<name>.grounding.corpus_ref` | string | Same `^[a-z0-9][a-z0-9_-]{0,63}$` pattern |

## §extensions

Free-form `Record<string, unknown>`. Accepted on a raw agent-scope PUT, but the composer does **not** copy `extensions` onto the canonical/composed card — it is stored on the raw card only and never appears in a `GET`/`preview-compose` response. Do not use it to carry data a runtime consumer needs to read.

## §\_composition (canonical-only)

Only returned from `GET`/`preview-compose` when the caller passes `?include_composition=true`; absent on raw agent-scope cards written by `PUT /v1/protection/agent/{agent_id}`.

```yaml theme={null}
_composition:
  canonical_id: cp-44ee22bb
  composed_at: 2026-04-26T18:23:41Z
  scopes_applied:
    - scope: platform
      version: 1
    - scope: "org:acme"
    - scope: "agent:mnm-patch-001"
      card_id: pc-88ccdd11
  exemptions_applied: []
  source_card_id: pc-88ccdd11
```

Same shape as the [alignment card's `_composition`](/specifications/alignment-card-schema#_composition-canonical-only) — see that page for the field table. `_composition` is read-only on the wire.

## YAML safe schema

All `yaml.load()` calls use `{ schema: yaml.CORE_SCHEMA }` — Node-specific tags are rejected. Plain scalars, maps, and sequences only.

## Body-size limits

Full protection card payload: **64 KB** max, enforced at the API boundary from both `Content-Length` and the actual body size. Oversize bodies get `413 Payload Too Large`.

## Versioning

`card_version` is a required, non-empty string. The composer stamps every canonical card with `protection/<YYYY-MM-DD>` using the composition date — treat the prefix (`protection/`) as the stable signal, not the date suffix, which changes on every recompose.

## See also

* [Protection Card](/concepts/protection-card) — conceptual overview
* [Safe House](/concepts/safe-house) — the detection pipeline this card configures
* [Alignment Card Schema](/specifications/alignment-card-schema) — companion spec
* [Card Composition](/concepts/card-composition) — three-scope composition rules + exemptions


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.