> ## Documentation Index
> Fetch the complete documentation index at: https://docs.mnemom.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Verifiable Verdicts

> Zero-knowledge proofs for verdict derivation, enabling cryptographic verification that an integrity verdict was computed correctly from its inputs.

Verifiable Verdicts extend the AIP attestation layer with **zero-knowledge proofs** that cryptographically demonstrate a verdict was derived correctly from its inputs. While [signatures](/protocols/aip/certificates) prove *who* signed a verdict and [hash chains](/protocols/aip/specification#16-attestation) prove *when* it was signed, ZK proofs prove *how* it was computed -- closing the remaining trust gap.

***

## What gets proven

The ZK proof covers the **verdict derivation rules**: the deterministic mapping from analysis outputs to the final verdict. Specifically, the proof demonstrates:

1. **Category-to-severity mapping**: Each concern category (e.g., `boundary_violation`, `prompt_injection`, `value_misalignment`) is mapped to the correct severity level (`low`, `medium`, `high`, `critical`) according to the protocol rules.

2. **Severity-to-verdict escalation**: The most severe concern determines the checkpoint verdict:
   * No concerns or all `low` severity results in `clear`
   * `medium` severity results in `review_needed`
   * `high` or `critical` severity results in `boundary_violation`

3. **Concerns hash integrity**: The hash of the concerns array in the proof journal matches the hash in the certificate, proving the verdict was derived from the exact concerns that were reported.

4. **Card constraint evaluation**: The verdict accounts for the agent's declared autonomy envelope -- bounded actions, forbidden actions, and escalation triggers -- as specified in the Alignment Card.

## What does NOT get proven

The ZK proof intentionally excludes:

* **LLM inference**: The analysis LLM's reasoning process is not proven. LLM inference is non-deterministic and computationally infeasible to prove inside a zkVM. The proof covers only the deterministic post-processing of the LLM's structured output.

* **Thinking block content**: The raw thinking block is never included in the proof. Only the SHA-256 hash (`thinking_block_hash`) is referenced, preserving the agent's reasoning privacy.

* **Analysis model selection**: The choice of which LLM to use for analysis is an operational decision outside the proof scope.

<Warning>
  The ZK proof guarantees that the verdict was correctly derived from the analysis output, not that the analysis itself was correct. The analysis quality remains bounded by the analysis LLM's capability. See [Limitations](/protocols/aip/limitations).
</Warning>

***

## Guest program

The verdict derivation logic runs inside an [SP1](https://docs.succinct.xyz/) zkVM guest program. The guest program is a deterministic Rust binary that takes structured inputs, applies the verdict rules, and produces public outputs.

### Inputs (private)

These are provided to the guest program but are **not** revealed in the proof:

| Input | Type | Description |
| - | - | - |
| `concerns` | `Vec<Concern>` | The array of concerns from the analysis LLM, each with `category`, `severity`, and `description`. |
| `card` | `CardSummary` | Compressed Alignment Card: values, bounded actions, forbidden actions, escalation triggers. |
| `model` | `String` | Analysis model identifier. |

### Computation

The guest program executes the following deterministic steps:

1. **Validate input structure**: Verify concerns array and card summary are well-formed.
2. **Apply category mapping**: Map each concern's category to the protocol-defined severity rules.
3. **Compute verdict**: Determine the final verdict from the highest severity across all concerns.
4. **Compute concerns hash**: SHA-256 hash of the serialized concerns array (deterministic key ordering).
5. **Write journal**: Output the verdict, concerns hash, and card ID as public commitments.

### Outputs (public journal)

These are committed to the proof's public journal and are visible to verifiers:

| Output | Type | Description |
| - | - | - |
| `verdict` | `String` | The computed verdict (`clear`, `review_needed`, `boundary_violation`). |
| `concerns_hash` | `String` | SHA-256 of the serialized concerns array. Verifiers compare this to the certificate's concerns to confirm consistency. |
| `card_id` | `String` | The Alignment Card ID, proving the verdict was derived against a specific card. |

***

## STARK proof format

The proof is an SP1 STARK (Scalable Transparent ARgument of Knowledge), stored and transmitted in three components:

| Component | Description |
| - | - |
| `receipt` | The STARK proof receipt. Contains the cryptographic proof data. Base64-encoded. |
| `journal` | The public journal output. Contains the committed public values (verdict, concerns\_hash, card\_id). Base64-encoded. |
| `image_id` | The hash of the guest program binary. Uniquely identifies which program was executed. Used to verify the proof was generated by the correct verdict derivation program. |

The `image_id` is a content hash of the compiled guest program. It acts as a commitment to the exact computation that was proven. A verifier checks:

1. The `image_id` matches the expected verdict derivation program (published in the Mnemom GitHub repository).
2. The `receipt` is a valid STARK proof for the given `image_id` and `journal`.
3. The `journal` outputs are consistent with the certificate's claims.

***

## Tiered proving strategy

Not every checkpoint receives a ZK proof. Proving has non-trivial cost and latency, so AIP uses a **tiered proving strategy** that distinguishes real violations from configuration errors:

Every verdict type is proven stochastically -- there is no verdict that is proven unconditionally by default. `boundary_violation` gets its own, separately configurable sampling rate rather than a higher-priority guarantee:

| Verdict | Proving Policy | Rationale |
| - | - | - |
| `boundary_violation` (no card gaps) | **Stochastic sampling** at `proof_boundary_cap` (default 5%) -- set this to 100 to prove every confirmed violation | Confirmed behavioral violations are the highest-value target for independent verification, but proving every one by default is a cost/latency tradeoff each deployment makes explicitly. |
| `boundary_violation` (with card gaps) | **Deferred** | Card-gap violations are configuration errors (stale alignment cards), not agent misbehavior. Proving is deferred until [DDR reconciliation](/concepts/card-lifecycle#violation-reclassification) classifies the violation. |
| `review_needed` | **Stochastic sampling** at `proof_rate` (default 10%) | Sampled proofs provide statistical confidence without proving every checkpoint. |
| `clear` | **Stochastic sampling** at `proof_rate` (default 10%) | Sampling clear verdicts catches systematic errors in the derivation logic. |

### Deferred proofs

When the [policy engine](/concepts/policy-engine) detects `card_gaps` alongside a `boundary_violation`, the gateway creates a proof record with `status: deferred` instead of immediately dispatching it to the prover. This keeps a tamper-evident audit trail without generating a proof yet.

The observer's DDR ([Disagreement-Driven Review](/protocols/aip/specification#15-disagreement-driven-review)) reconciliation later classifies the violation:

| DDR Outcome | Proof Resolution |
| - | - |
| `card_gap` | `skipped_card_gap` -- proof permanently skipped |
| `noise` | `skipped_noise` -- proof permanently skipped |
| `aip_miss` | Upgraded to `pending` and dispatched to the prover |

This means only **confirmed real violations** are proven, which matters most during rapid development, when alignment cards drift often.

### On-demand proving

Any checkpoint can be proven on demand via [`POST /v1/checkpoints/{checkpoint_id}/prove`](/api-reference/endpoint/post-checkpoints-id-prove). This is useful when:

* An operator wants to verify a specific `review_needed` verdict
* An auditor requires proof for a specific checkpoint
* The stochastic sample missed a checkpoint of interest
* A deferred proof needs to be forced without waiting for DDR

The request needs an API key belonging to a member of the org that governs the checkpoint's agent; any other caller gets the same `404` as a missing checkpoint. If a proof already exists, the endpoint returns its current status. Each newly queued proof is charged in μ to that org's billing account (rate card class `analyze.zk-proof`, see [Pricing](/pricing/overview)); returning an existing proof is not charged. If the billing account cannot be resolved, no proof is queued and the request fails with `503` (or `409` when the org has no billing account).

***

## Cost and latency

| Metric | Value | Notes |
| - | - | - |
| **Price** | 100 μ per proof | Rate card class `analyze.zk-proof`, charged once per checkpoint. See [Pricing](/pricing/overview). |
| **Proving latency** | \~700ms | Typical time from proof request to receipt. Each completed proof reports its own `proving_duration_ms`. |
| **Verification latency** | \<100 milliseconds | STARK verification is fast. Local verification avoids network round-trips. |
| **Proof size** | \~2.7 MB | SP1 compressed STARK receipt. |

Proving is asynchronous. The checkpoint verdict is returned immediately; the proof is generated in the background and attached to the certificate when complete. In most conversations it completes before the next integrity checkpoint.

### Proof status and failures

Poll [`GET /v1/checkpoints/{checkpoint_id}/proof`](/api-reference/endpoint/get-checkpoints-id-proof) to follow a proof. Its `status` is one of `pending`, `proving`, `completed`, or `failed`.

The proving inputs are stored with the proof request, so a transient failure (for example, the prover restarting) does not lose the request: it is retried automatically, with no action needed from you. A proof that still cannot be generated ends in `failed`. The checkpoint itself remains valid; it just lacks the additional computational integrity guarantee.

***

## Verification

### Server-side

Submit the full certificate to [`POST /v1/verify`](/api-reference/endpoint/post-verify). The API delegates STARK verification to the prover service and returns a structured result including `checks.verdict_derivation.valid`.

### Local verification

Use the SP1 verifier SDK to verify locally:

```rust theme={null}
use sp1_sdk::{ProverClient, SP1ProofWithPublicValues};

let client = ProverClient::from_env();
let (_, vk) = client.setup(ELF);

let proof: SP1ProofWithPublicValues = deserialize(&cert.proofs.verdict_derivation.receipt);

// Verify the proof
client.verify(&proof, &vk).expect("Proof verification failed");

// Read the public values
let journal = proof.public_values.read::<VerdictJournal>();
assert_eq!(journal.verdict, cert.claims.verdict);
```

### Checking proof status

For checkpoints where proving is in progress, query the status (authenticated):

```bash theme={null}
curl https://api.mnemom.ai/v1/checkpoints/{checkpoint_id}/proof \
  -H "Authorization: Bearer $MNEMOM_TOKEN"
```

The `status` field progresses through: `pending` -> `proving` -> `completed` (or `failed`).

For deferred proofs, the progression is: `deferred` -> `skipped_card_gap` | `skipped_noise` | `pending` (then normal flow).

***

## See also

* [Integrity Certificates](/protocols/aip/certificates) -- The certificate format that contains ZK proofs
* [AIP Specification -- Verification](/protocols/aip/specification#17-verification) -- Protocol-level verification specification
* [SP1 Documentation](https://docs.succinct.xyz/) -- The zkVM platform used for proving


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.