Skip to main content
Normative reference for the protection card — the YAML document that configures Safe House for a specific agent, and one half of every Mnemom agent’s two cards. This page specifies every section, field, required/optional status, type, and composition semantic. Conceptual overview: /concepts/protection-card. Alignment-card spec: /specifications/alignment-card-schema. Card composition rules across platform/org/agent scopes: /concepts/card-composition.

Top-level structure

thresholds, screen_surfaces and trusted_sources may each be omitted from a raw agent-scope PUT — the composer fills the platform default for whichever bucket is absent. thresholds requires all three of warn/quarantine/block when the block is present at all; screen_surfaces and trusted_sources validate whichever individual keys you send and leave the rest to the composer — see each section below.

§mode

Top-level action policy for Safe House on this agent. The same off | observe | nudge | enforce enum as the alignment card’s autonomy_mode / integrity_mode master switches — see Master switches.
enforce implies synchronous verdict — to block a request, the gateway must wait for the verdict before delivering the message. There is no separate enforce_sync mode. Composition: strictest wins across enforce > nudge > observe > off. An agent cannot drop below the platform/org floor. nudge is the load-bearing middle ground: the model receives the advisory as part of its prompt context, so the security signal reaches the model without blocking the request. Customers running long-tail-confidence detectors typically run nudge rather than enforce until thresholds settle.

§thresholds

Three-band escalation ladder for Safe House detector scores. All values are floats in [0, 1].
Validation: warn ≤ quarantine ≤ block. The validator rejects any out-of-order combination at write time. Composition: min across scopes. The lowest threshold wins, since lower = stricter (matches sooner). An agent cannot loosen a stricter platform/org threshold; it can only tighten further. Three bands map cleanly onto the SOC severity ladder familiar to most operators. Per-detector tuning is an internal calibration concern and is not exposed in the schema.

§screen_surfaces

Which request surfaces Safe House inspects, named by direction (incoming/outgoing) and tool relationship.
Surfaces are screening units, not a per-turn budget. incoming and tool_responses are both front-door surfaces and the detector pipeline runs on each independently, within the same request: once for the inbound message, once more for each tool result the request carries. See When the front door runs. Validation: Only the four named keys are accepted. Unknown keys are rejected at write time. Composition: OR per field — true wins. If any scope sets a surface to true, it’s scanned. Agents cannot disable scanning that org or platform requires. Phrased in alignment-card vocabulary: strictest wins (with true = scan being the more restrictive choice). Direction-based naming is durable across transport changes: an agent receiving a webhook trigger is “incoming” whether it’s a user message, an API event, or a queue payload. Differentiating tool_calls from tool_responses reflects that they have different threat models — outgoing tool args may exfiltrate; incoming tool responses may inject. It also reflects a different enforcement shape: an incoming finding can fail the request, whereas a tool_responses finding is applied inside the request — the offending tool result is withheld or decorated in the body and the rest is forwarded. Turning off a surface emits a low-priority audit trace so reviewers can see what was not scanned. If you need to disable a surface for a specific agent, the recommended path is an exemption with a documented reason rather than a raw false in the agent card.

trusted_sources

Per-bucket allowlist of upstream sources whose content Safe House skips detection for. The buckets are typed so the validator can apply per-bucket deny-lists and the composer can apply per-bucket intersection rules.
Composition:
  • Platform → agent: intersection. The platform list is the compliance ceiling — downstream scopes (org, agent) cannot widen trust beyond what the platform allows. If the platform sets ip_ranges: [10.0.0.0/8], an agent cannot add 192.168.0.0/16 to its own list and have it take effect.
  • Org + agent: union within the ceiling. Either scope can add trust within the platform-imposed ceiling.
  • Empty platform list = unconstrained ceiling. When the platform doesn’t specify a bucket, downstream entries pass through without intersection.
Trusted sources cause Safe House to skip detection for matching content (no detector cycles spent), but every match emits a low-priority sh_trusted_source_skip audit trace so reviewers can see what was waved through. Security note: the validator’s deny-list is non-exhaustive — adding a publicly-routable IP range or a customer-controllable domain is a critical misconfiguration even if it passes the deny-list. Treat trusted_sources as a sharp tool.

§protected_surface

Org-declared assets and mechanical operations the substrate must protect, enforced independent of the agent’s own declared intent (the alignment card’s autonomy.forbidden_actions is advisory; this is enforced). The composer always emits this block — a caller may omit it entirely to inherit the composed floor.
Precedence when a request matches more than one: forbidden_operations > escalation_required > allow.

§review

Review-hold policy (Safe House Review). Optional; a raw card may set any subset of fields and the composer fills the rest.
Composition: strictest-wins per field, same ladder as mode.

§scopes

Agent-declared named scopes used for grounding checks (agent scope only — org and platform scopes do not contribute in the current version).

§extensions

Free-form Record<string, unknown>. Accepted on a raw agent-scope PUT, but the composer does not copy extensions onto the canonical/composed card — it is stored on the raw card only and never appears in a GET/preview-compose response. Do not use it to carry data a runtime consumer needs to read.

§_composition (canonical-only)

Only returned from GET/preview-compose when the caller passes ?include_composition=true; absent on raw agent-scope cards written by PUT /v1/protection/agent/{agent_id}.
Same shape as the alignment card’s _composition — see that page for the field table. _composition is read-only on the wire.

YAML safe schema

All yaml.load() calls use { schema: yaml.CORE_SCHEMA } — Node-specific tags are rejected. Plain scalars, maps, and sequences only.

Body-size limits

Full protection card payload: 64 KB max, enforced at the API boundary from both Content-Length and the actual body size. Oversize bodies get 413 Payload Too Large.

Versioning

card_version is a required, non-empty string. The composer stamps every canonical card with protection/<YYYY-MM-DD> using the composition date — treat the prefix (protection/) as the stable signal, not the date suffix, which changes on every recompose.

See also