https://api.mnemom.ai
Configuration
Safe House behavior is configured through the protection card — a per-agent manifest of mode, score thresholds, screened surfaces, and trusted sources, composed across platform → org → agent scopes. Control it globally for the org (via the org protection template), per-agent, or in bulk.
The granular sub-resource paths (
/mode, /thresholds, /screen_surfaces, /trusted_sources) accept PUT or PATCH and exist at every scope (/protection/agent/:agent_id/..., /protection/org/:org_id/..., /protection/team/:team_id/..., /protection/platform/:scope_id/...). PUT on these requires an Idempotency-Key and If-Match header.
Retrieve the org-scope manifest:
PUT /v1/protection/agent/:agent_id with a full protection card.
Bulk-apply a config to many agents:
Quarantine management
Quarantined items are held pending human review. Reviewers can release (with or without a false-positive flag) or confirm as a genuine threat. The unit is an evaluation, not a turn. The front door runs once per inbound surface a request carries, so a single turn can produce several evaluations and several quarantine records — one for the inbound message, and one for each tool result screened in the same request. See When the front door runs.
List open quarantine items:
Query & observability
Query the full evaluation history, aggregate metrics, and access a live SSE stream for real-time monitoring.
Query evaluations with filters:
data: line is a JSON evaluation event (agent_id, verdict, overall_risk, top_threat,
…). Filter with agent_id, org_id, surface, a comma-separated verdict list, and min_risk.
The stream polls internally every 2 seconds, heartbeats roughly every 15 seconds, and closes
after 5 minutes — reconnect on close.
Patterns & intelligence
Manage the threat pattern library and retrieve adaptive threshold recommendations.
List active patterns for a threat type:
threat_type, scope (agent or org),
current vs. suggested threshold values, a rationale, and a confidence level.
Submit a candidate pattern:
A submission is a labeled example message, not a regex — the arena evaluation pipeline derives
detection logic from a labeled corpus rather than taking a hand-written pattern directly.
label is malicious or benign — submitting confirmed benign examples that resemble an
attack pattern helps reduce false positives too. Submitted content enters candidate status.
The arena evaluation pipeline tests candidates against the labeled message set, and patterns
that exceed precision/recall thresholds are promoted to active.
Canary credentials
Canary credentials are honeypot API keys, tokens, or other secrets deliberately planted in the agent’s context. If an attacker extracts and uses them, Safe House detects the use and fires ansh.canary.triggered webhook event.
Create a canary:
canary_type is one of api_key, password, database_url, ssh_key, oauth_token.
canary_value is returned only at creation time. Safe House monitors for its appearance in outbound requests or inbound message content.
Check canary status:
Special endpoints
Cross-Agent campaign detection
List detected attack campaigns — groups of related attacks targeting multiple agents from the same infrastructure.EU AI Act compliance export
Export Safe House evaluation data in EU AI Act Article 50 compliance format.Accept: text/csv for spreadsheet-compatible export.
Error responses
All Safe House endpoints return standard Mnemom error objects:See also
- Safe House Threat Model — What each threat type means and how detection works
- Webhook Notifications — React to Safe House events in real-time
- Safe House Monitoring — Security Observatory and alert management
- Policy Overview — Policy enforcement runs alongside Safe House