> ## Documentation Index
> Fetch the complete documentation index at: https://docs.mnemom.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Risk Assessment

> Context-aware, real-time risk scoring for individual agents and teams — with optional zero-knowledge proofs

Risk assessment answers the question every platform operator asks before letting an agent act: **how dangerous is this, right now, in this context?**

Unlike static reputation scores, risk assessments are dynamic. The same agent can be low-risk for a data access request and high-risk for a financial transaction — because different actions stress different capabilities. Risk assessments incorporate the agent's reputation, recent violation history, the specific action being attempted, and the operator's risk tolerance.

For teams, the engine goes further: it evaluates whether a group of individually acceptable agents might still be dangerous when combined — through correlated failure, value divergence, or contagion dynamics.

## Individual risk

Each individual risk assessment produces a score between 0 and 1, computed as:

```
R = 0.60 × R_context + 0.30 × R_recency + 0.10 × R_confidence
```

### Context-aware component risk (60%)

Five reputation components are weighted differently depending on the action type:

| Component | What It Tracks |
| - | - |
| **Integrity Ratio** | How often behavior matches declared values |
| **Compliance** | Adherence to organizational rules |
| **Drift Stability** | Whether behavior is changing over time |
| **Trace Completeness** | Whether full reasoning traces are provided |
| **Coherence Compatibility** | How well the agent works with the fleet |

Different actions weight these differently. A `financial_transaction` emphasizes compliance (0.30) and integrity (0.30). A `task_delegation` emphasizes coherence (0.35) — can this agent hand off work reliably? A `multi_agent_coordination` action weights coherence highest (0.40).

<Info>
  Six action types are supported: `financial_transaction`, `data_access`, `task_delegation`, `tool_invocation`, `autonomous_operation`, and `multi_agent_coordination`. Each has a distinct weight profile tuned to the risks specific to that action.
</Info>

### Recency penalty (30%)

Recent violations count more than old ones. The engine uses exponential decay with a 30-day half-life:

```
R_recency = sum(severity_weight x e^(-lambda x days_ago))
```

A critical violation yesterday contributes nearly 1.0. The same violation 30 days ago contributes 0.5. After 90 days, it is negligible. Severity weights range from 0.1 (low) to 1.0 (critical).

### Confidence penalty (10%)

Agents with limited behavioral history receive an uncertainty premium: `insufficient` data adds 0.30, `low` adds 0.20, `medium` adds 0.10, and `high` confidence adds nothing.

### Amount scaling

When `context.amount` is supplied (e.g. for a `financial_transaction`), the composite `R` is multiplied by a log10-based scale that grows with transaction size: roughly 1.0x at $100, 1.1x at $1,000, 1.2x at $10,000, 1.3x at $100,000, and 1.4x at \$1,000,000, capped at 1.5x. No `amount` (or an amount ≤ 0) leaves the score unscaled.

### Risk levels and recommendations

The composite score maps to four risk levels. Thresholds shift based on the caller's risk tolerance:

| Tolerance | Low | Medium | High | Critical |
| - | - | - | - | - |
| Conservative | \< 0.15 | \< 0.35 | \< 0.55 | >= 0.55 |
| Moderate | \< 0.25 | \< 0.50 | \< 0.75 | >= 0.75 |
| Aggressive | \< 0.35 | \< 0.60 | \< 0.85 | >= 0.85 |

Each level maps to a recommendation: `approve` (low/medium), `review` (high, requires human approval), or `deny` (critical, block the action).

## Team risk

Team risk assessment evaluates whether a group of agents is safe to operate together. A team of individually low-risk agents can still be dangerous.

### Three-pillar model

```
TeamCoherence = 0.30 x (1 - AQ) + 0.45 x CQ + 0.25 x (1 - SR)
TeamRisk = 1 - TeamCoherence
```

**Aggregate Quality (AQ)** uses tail-risk weighting inspired by CoVaR (Conditional Value at Risk): agents with higher individual risk get exponentially more weight. One bad agent drags the score down far more than one good agent lifts it up.

**Coherence Quality (CQ)** evaluates pairwise compatibility across four dimensions — value overlap, priority alignment, behavioral correlation, and boundary compatibility. CQ penalizes high variance: a uniformly moderate team beats a volatile one where some pairs are excellent and others are terrible.

**Structural Risk (SR)** models failure contagion. For each pair of agents, SR estimates how much damage one agent's failure would cause to the other, based on their coherence gap and the "value at risk." The team's SR combines the worst single-agent vulnerability with the fleet average.

### Shapley attribution

After computing team coherence, the engine attributes each agent's marginal contribution using leave-one-out (LOO) Shapley values:

```
MC_i = TeamCoherence(all) - TeamCoherence(all without agent i)
```

Positive values mean the agent improves the team. Negative values mean they drag it down. This tells operators exactly which agents to swap, add, or remove to optimize team composition.

Historical team risk assessments are a primary input to the [Team Trust Rating](/concepts/team-reputation) — specifically the Coherence History and Operational Record components. Teams that consistently receive low-risk assessments build stronger reputation scores over time.

### Circuit breakers

Hard safety floors override the continuous score when conditions are extreme:

* Any agent with reputation below 200 forces the team to critical/deny
* Any pairwise boundary compatibility below 100 forces critical/deny

### Additional analytics

| Analysis | What It Detects |
| - | - |
| **Outlier Detection** | Agents whose risk exceeds the fleet mean by > 1 standard deviation |
| **Cluster Detection** | Groups of agents with correlated risk (shared failure modes) |
| **Value Divergence** | Values declared by some agents but missing in others |
| **Synergy Detection** | Whether the team score is better or worse than the average of individual scores |

### Team recommendations

| Score Range | Recommendation | Meaning |
| - | - | - |
| \< 0.20 | `approve_team` | Team operates as a unit |
| \< 0.40 | `approve_team` | Team operates with monitoring |
| \< 0.60 | `approve_individuals_only` | Individual actions allowed, joint operations blocked |
| >= 0.60 | `deny` | All team operations blocked |

## Zero-knowledge proofs

Every risk assessment can optionally be backed by a cryptographic proof. Immediately after `/risk/assess` or `/risk/assess/team` computes a score, the risk service fires an asynchronous, fire-and-forget request to a separate proving service and links the resulting `proof_id` to the assessment. The risk score and recommendation are returned synchronously and are always valid regardless of proof status -- an absent or failed proof never invalidates the assessment.

### Proof lifecycle and fail-open behavior

`proof_status` (on the assessment) and `status` (on the proof object returned by `GET /v1/risk/proofs/:proof_id`) share the same lifecycle:

| Status | Meaning |
| - | - |
| `none` | No proof was dispatched for this assessment (`proof_id` is unset). |
| `pending` | A proof request was dispatched to the proving service; computation has not started. |
| `proving` | The proving service is actively computing the proof. |
| `verified` | The proof completed and was verified. Retrieve the full record via `GET /v1/risk/proofs/:proof_id`. |
| `failed` | Proof generation failed. The risk score remains valid; only the proof could not be produced. |

Proofs typically complete within tens of seconds. Poll `GET /v1/risk/proofs/:proof_id` to follow a proof from `pending` through to `verified` or `failed`; see [Get proof](/api-reference/risk-overview#get-proof) in the API reference.

<Note>
  Proof availability depends on the feature entitlements active on the caller's account -- see [Feature gating](/api-reference/risk-overview#feature-gating). Current μ-based usage pricing has no fixed plan tiers; check your account settings for which entitlements are active.
</Note>

## Authorization model

All Risk API endpoints require authentication. Pass your credentials as either:

* **Bearer token**: `Authorization: Bearer <token>`
* **API key**: `X-Mnemom-Api-Key: <key>`

### Account-scoped GET endpoints

GET endpoints that retrieve a stored resource by ID — `GET /v1/risk/assessments/:assessment_id`, `GET /v1/risk/team-assessments/:assessment_id`, and `GET /v1/risk/proofs/:proof_id` — are account-scoped. A request with a valid token from a **different account** returns `404` rather than `403`. This prevents leaking resource existence to authenticated callers who do not own the resource.

List endpoints (`GET /v1/risk/history/:agent_id`, `GET /v1/risk/team-history/:team_id`) filter implicitly to the calling account and return an empty list rather than `404` when the requested ID belongs to a different account.

### Which agents you can assess

Assessing another organization's public or unlisted agent is allowed — that counterparty check is what the engine is for. An assessment never carries reputation data that the reputation API itself withholds, so both `POST /v1/risk/assess` and `POST /v1/risk/assess/team` check every agent first:

* **Private reputation** — an agent whose reputation visibility is `private` can be assessed only by a member (`owner`, `admin`, `member` or `viewer`) of the organization that governs it. Anyone else gets `403` with code `reputation_private`.
* **Deleted agents** — a deleted (tombstoned) agent returns `404` with code `agent_not_found`.
* **Teams by `team_id`** — a team assessment by `team_id` requires membership in the team's organization. A team in an organization you don't belong to returns the same `404` as a team that doesn't exist.

If the membership check itself cannot complete, the request fails closed with `503` and code `membership_unresolved`; retry it.

## See also

* [Reputation Scores](/concepts/reputation-scores) — the input data that feeds risk assessments
* [Team Trust Rating](/concepts/team-reputation) — team-level reputation built from team risk assessments
* [Fleet Coherence](/concepts/fleet-coherence) — the pairwise coherence data used for team risk
* [Integrity Checkpoints](/concepts/integrity-checkpoints) — how violations are detected
* [Risk Engine Guide](/guides/risk-engine) — step-by-step usage with SDK examples
* [Security & Trust Model](/guides/security-trust-model) — the broader cryptographic verification pipeline


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.