> ## Documentation Index
> Fetch the complete documentation index at: https://docs.mnemom.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Provider Support

> Honest per-provider coverage of every Mnemom feature — what works, what's partial, and what's not supported across Anthropic, OpenAI, and Gemini

Mnemom sits in front of three upstream model providers: **Anthropic**, **OpenAI**, and **Gemini**. Safe House, AIP integrity checkpoints, CLPI policy enforcement, and DLP all route through a single gateway, but the **quality of each feature varies by provider** — because the underlying APIs differ in what they expose.

This page is the honest accounting: what Mnemom guarantees on each provider, what it partially guarantees, and what it cannot guarantee.

<Note>
  Where coverage is partial — particularly **OpenAI's integrity-checkpoint coverage** — Mnemom does not claim parity. The v1 commitment is honest per-provider differentiation, not uniform coverage. Marketing materials and the trust center reflect this.
</Note>

## Supported models

Models the v1 promise applies to. Listed in the gateway's `/models.json` registry; routed end-to-end through Safe House, AIP, CLPI, and DLP per the matrix below.

<CardGroup cols={3}>
  <Card title="Anthropic" icon="brain">
    Claude Fable 5.1
    Claude Opus 5.5
    Claude Opus 5
    Claude Sonnet 5.5
    Claude Sonnet 5
    Claude Opus 4.8
    Claude Fable 5
    Claude Opus 4.7
    Claude Sonnet 4.6
    Claude Haiku 4.5
  </Card>

  <Card title="OpenAI" icon="microchip">
    GPT-5
    GPT-5 Codex
    o3
    o3-mini
    GPT-5.6 Sol
    GPT-5.6 Terra
    GPT-5.6 Luna
    GPT-6 Astra
    GPT-6.1 Sol
    GPT-6 Sol
    GPT-6 Luna
    GPT-5.5
  </Card>

  <Card title="Gemini" icon="gem">
    Gemini 2.5 Pro
    Gemini 2.5 Flash
    Gemini 3.8 Flash
    Gemini 3.5 Flash-Lite
  </Card>
</CardGroup>

**Canonical model IDs** (what you pass in `model:` on a Mnemom API call):

```yaml theme={null}
# Supported tier. Mnemom's gateway exposes this set via
# /models.json with `supported: true` flag; passthrough models route
# but carry no v1 promise.
supported_models:
  anthropic:
    - claude-fable-5-1
    - claude-opus-5-5
    - claude-opus-5
    - claude-sonnet-5-5
    - claude-sonnet-5
    - claude-opus-4-8
    - claude-fable-5
    - claude-opus-4-7
    - claude-sonnet-4-6
    - claude-haiku-4-5-20251001
  openai:
    - gpt-5
    - gpt-5-codex
    - o3
    - o3-mini
    - gpt-5.6-sol
    - gpt-5.6-terra
    - gpt-5.6-luna
    - gpt-6-astra
    - gpt-6.1-sol
    - gpt-6-sol
    - gpt-6-luna
    - gpt-5.5
  gemini:
    - gemini-2.5-pro
    - gemini-2.5-flash
    - gemini-3.8-flash
    - gemini-3.5-flash-lite
```

Additional models may route through the gateway in **passthrough mode** — they work for inference but do not carry the v1 feature-correctness promise. See [Deprecation policy](#deprecation-policy) below.

## Feature coverage matrix

Each cell describes the v1 commitment level. Symbols:

* **✓** Fully supported. Tested in CI; Safe House features work identically to the Anthropic baseline.
* **⚠ Partial** — Provider exposes the feature, but with a documented limitation. See the supporting note.
* **N/A** — Provider does not expose this capability. Not a Mnemom limitation.

| Capability | Anthropic | OpenAI | Gemini |
| - | - | - | - |
| Streaming (SSE) | ✓ | ✓ | ✓ |
| Tool use / function calling | ✓ | ✓ | ✓ |
| **Thinking-trace inspection** (AIP) | ✓ <sup>\[1]</sup> | Not available <sup>\[2]</sup> | ✓ <sup>\[3]</sup> |
| Multimodal (image inputs) | ✓ | ✓ | ✓ |
| Prompt caching | ✓ <sup>\[4]</sup> | ⚠ Partial <sup>\[5]</sup> | ⚠ Partial <sup>\[6]</sup> |
| Batch API | ✓ | ✓ | ✓ |

### \[1] Anthropic — full extended thinking

Anthropic models expose full **extended thinking blocks** through the response API. AIP reads completed thinking blocks post-response (before delivery in `enforce`, which adds latency; post-delivery in `observe`/`nudge`), and the verifier (Claude Haiku 4.5) has full chain-of-thought visibility. This is the most complete AIP coverage Mnemom offers — and the baseline against which other providers are compared.

If the client's thinking settings return thinking blocks without readable text, AIP grades the turn's visible reply and tool calls instead of skipping the check. See [Integrity Checkpoints](/concepts/integrity-checkpoints#provider-support).

### \[2] OpenAI — no thinking-trace inspection through the gateway today

OpenAI's reasoning summaries are a **Responses API** feature (`reasoning.summary`), not something the Chat Completions API returns. The gateway's default `/openai` door proxies Chat Completions — the shape OpenAI clients send by default — and that shape never carries reasoning content back, for any OpenAI model, reasoning or not. AIP's thinking-extraction step returns empty by construction for every `/openai/*` request today, including through the `/openai/v1/responses` passthrough path. There is no per-model split here: it is uniform across `o3`, `o3-mini`, `gpt-5`, `gpt-5-codex`, `gpt-5.6-*`, `gpt-6-astra`, `gpt-6.1-sol`, `gpt-6-sol`, `gpt-6-luna`, and `gpt-5.5`.

**What this means in practice:**

* Boundary violations that surface in the model's final response are still caught equally well on OpenAI — Safe House and CLPI don't depend on the thinking trace.
* Boundary violations that would only have surfaced in the model's hidden reasoning are **not caught** on OpenAI — AIP has nothing to analyze, so a checkpoint below the minimum-evidence threshold is emitted as a synthetic `clear` without an analysis-LLM call, not a partial-confidence verdict.
* If your application relies on AIP catching boundary violations in the model's *reasoning* specifically, use Anthropic or Gemini for that workload. OpenAI models remain fully supported for inference, Safe House, and CLPI — just not for thinking-trace inspection.

This is the load-bearing honesty disclosure: Mnemom does not claim AIP thinking-trace coverage on OpenAI, on any model, through the gateway as it exists today.

### \[3] Gemini — full `thoughts` exposure

Gemini models expose a `thoughts` field on response candidates. Coverage is uniform across 2.5 Pro, 2.5 Flash, 3.8 Flash, and 3.5 Flash-Lite — all four ship Google's `thinking` capability (3.8 Flash defaults to medium effort, 3.5 Flash-Lite defaults to minimal, both adjustable) — and AIP reads `thoughts` through the gateway's response normalizer and treats it as equivalent to Anthropic extended thinking.

### \[4] Anthropic — explicit `cache_control`

Anthropic supports explicit `cache_control` block markers — customers control which prompt segments are cached. Mnemom passes `cache_control` through transparently. **Safe House still evaluates the full request** (it does not assume cached prefixes are safe just because they were previously seen). Cache hits do not bypass any checkpoint.

### \[5] OpenAI — automatic caching, no customer control

OpenAI's prompt caching is **automatic** — the API decides what to cache based on request shape. Customers cannot reason about cache hit rates the way they can on Anthropic. Mnemom passes requests through unchanged; cache decisions are OpenAI's. **Safe House dispatch remains idempotent across cache hits and misses** — the same prompt routed through Mnemom twice produces the same verdict regardless of whether OpenAI cached it.

### \[6] Gemini — separate `CachedContent` API

Gemini exposes prompt caching as a separate `CachedContent` resource (explicit cache lifecycle, named caches with TTL). The gateway today does not surface or use this API; requests are sent without referencing cached content. Customers using Gemini's cache outside Mnemom will see lower latency than they see through Mnemom.

## AIP coverage by provider — the headline commitment

Of all per-provider gaps, **thinking-trace coverage** is the load-bearing one:

| Provider | Thinking-trace coverage |
| - | - |
| Anthropic (all supported models) | Full — native extended-thinking blocks, read directly |
| Gemini (2.5 Pro, 2.5 Flash, 3.8 Flash, 3.5 Flash-Lite) | Full — native `thoughts` content parts, read directly |
| OpenAI (all supported models) | None through the gateway today — see \[2] above |

**If your application relies on AIP catching boundary violations in the model's reasoning** — particularly violations that would not surface in the final response — choose Anthropic or Gemini. OpenAI models are fully supported for inference, Safe House screening, and CLPI tool-use governance; they are not the right choice when AIP thinking-trace analysis is the load-bearing safety layer.

## Latency expectations

In aggregate:

* **Safe House dispatch** adds \~15 ms P50 / \~60 ms P95 across all providers (the dispatch path is provider-agnostic).
* **AIP analysis** runs post-response (before delivery in `enforce`, which adds the latency below; post-delivery in `observe`/`nudge`), only when a thinking trace was extracted. Cost varies by upstream-response token volume — Anthropic Opus with full extended thinking emits the largest traces and therefore the longest AIP analysis tails (P95 up to 2.5 seconds). Gemini 2.5 Flash emits thinner traces and completes AIP analysis in P50 ≤ 800 ms. OpenAI requests skip the analysis-LLM call entirely (no trace to analyze) and add no AIP latency.

The gateway does not interrupt a response mid-stream. In `enforce`, AIP runs after the upstream response completes but before the gateway delivers it to the customer, so a violating response is gated before delivery (adding latency); this holds on every provider. In `observe` and `nudge`, the response is delivered first and the verdict is recorded post-hoc.

## How we test against each provider

Mnemom's gateway adapter — the code that parses each provider's response format, extracts thinking blocks, and routes tool calls — is tested in **two layers**:

| Layer | Cadence | Catches |
| - | - | - |
| **Static-shape** | Every PR to the gateway | "Did our parser break?" Asserts the adapter handles captured response fixtures correctly. |
| **Live** | Nightly | "Did the upstream provider ship a breaking change?" Real upstream call → real response → real parse. |

All three providers (Anthropic, OpenAI, Gemini) run both layers. Live tests fire nightly against real upstream APIs and detect breaking changes within 24 hours.

<Note>
  **OpenAI reasoning content, streaming or not:** OpenAI's Chat Completions API — streaming or non-streaming — does not return reasoning content for any model. Server-side reasoning token usage is reported in `usage.completion_tokens_details.reasoning_tokens`, but the reasoning text itself never reaches the response body Mnemom's gateway proxies. This is upstream API behavior, not a Mnemom limitation, and it is why OpenAI has no thinking-trace coverage today regardless of streaming mode — see \[2] above.
</Note>

## Deprecation policy

Models the gateway routes are classified into two tiers:

<CardGroup cols={2}>
  <Card title="Supported" icon="check">
    Listed in the [supported models](#supported-models) section above. v1 promise applies. Safe House, AIP, CLPI, DLP all work to their per-provider commitment. Tested in CI. Deprecation requires **90 days' notice** via this page and the changelog.
  </Card>

  <Card title="Passthrough" icon="forward">
    Routed by the gateway but not in the supported tier. Inference works; Safe House features are best-effort. Not tested in CI. No deprecation notice — model availability tracks the provider's lifecycle.
  </Card>
</CardGroup>

**Today's deprecation schedule:**

| Model | Status | Sunset date | Migration target |
| - | - | - | - |
| `claude-3-opus-20240229` | Passthrough | Tracks Anthropic's deprecation | `claude-opus-4-8` |
| `claude-3-5-sonnet-20241022` | Passthrough | Tracks Anthropic's deprecation | `claude-sonnet-4-6` |
| `claude-sonnet-4-20250514` | Passthrough | 2026-Q3 (recommend migrating now) | `claude-sonnet-4-6` |
| `gpt-4o` | Passthrough | Tracks OpenAI's deprecation | `gpt-5` |
| `gemini-3-pro`, `gemini-3-flash` | Preview | When Google releases stable | Stay on `gemini-3-*` once supported |

Passthrough-tier models that are removed from the upstream provider also disappear from the gateway's `/models.json` registry; we do not maintain shims.

## Out of scope

Provider expansion beyond Anthropic + OpenAI + Gemini is **not in scope for v1**. Cohere, Mistral, Together, Groq, and other providers are not supported. Adding a provider is a multi-quarter effort (Safe House dispatch, AIP adapter, CLPI policy schema mapping, test coverage, docs) — tracked separately.

BYOK (bring-your-own-key) for upstream providers is also out of v1. v1 ships with Mnemom-only key custody.

## Related

* [Integrity Checkpoints](/concepts/integrity-checkpoints) — the AIP analysis machinery
* [Safe House](/concepts/safe-house) — screening layer for everything reaching the agent: the inbound message at request arrival, plus each tool result turn-internally, inline, before the request is forwarded
* [CLPI](/concepts/clpi) — Continuous Local Policy Interpretation for tool use
* [Webhook contract](/concepts/webhook-contract) — event delivery for operator surfaces (provider-agnostic)


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.