Skip to main content
Mnemom has no plan tiers. Every feature (the gateway, integrity checks, alignment cards, reputation, the dashboard) is available to every account from the start. You pay for work as it happens, in μ.
For estimates at your own volume, use the pricing calculator.

The unit

1 μ = $0.01 USD. The peg is fixed and will not be re-denominated, so a μ balance reads directly as a dollar amount. Charges are recorded in thousandths of a μ and rounded up, so a small model call can cost a fraction of a μ.

Getting μ

1

Create an account

Create an account to explore the product. Nothing is billed until you buy μ.
2

Add a payment method

Save a card to buy μ in one step. Nothing is billed until you buy.
3

Buy more, or turn on auto top-up

Purchase any amount of μ in one click, or set a threshold and target so your balance tops itself up automatically when it runs low. There’s no minimum purchase.

What you pay for

When the gateway calls a model to do work for you, that call is charged at its measured model cost: the tokens it used, at the model provider’s published prices, with no markup. A step that makes no model call costs nothing. In a Mnemom agent session (a conversation the gateway runs with Mnemom’s agent features on, such as context management or goal alignment), each successful request also has an inference share. The share is 4% of that request’s forwarded inference spend, priced at the provider’s published rates. For each request you pay the larger of two amounts, never both:
  • the measured cost of the gateway’s own model calls for that request: agent work (context folds, the goal-alignment judge, and implicit-goal calls) and, when governance is on, governance (the integrity analysis and Safe House level 2 screens);
  • the 4% share.
When the share is larger, your statement shows the measured lines plus an “Inference share” line that tops them up to it. Some of the gateway’s own calls for a request can finish after its share has already been charged, for example a long integrity analysis. From 2026-10-04 00:00:01 UTC (rate card v2026-10-share-3), such a call still counts toward that request’s comparison. It is covered by the share first, and only the part above the share is charged. Before that instant, it is charged at measured cost on its own line. Observer trace analysis runs later, from your logs. From 2026-10-04 00:00 UTC (rate card v2026-10-share-2), when it analyzes a request, its measured cost counts toward that request’s comparison: the request is charged the larger of the 4% share and all of its measured costs, observer included. If the share was already charged, the observer line is covered by it first and only the part above the share is charged. Observer analysis that is not tied to a request is charged at measured cost on its own line. Before that date, observer analysis is always charged on its own line, on top of the comparison. There is no minimum charge per request. A request that fails upstream is charged nothing. Prices on this page match Mnemom rate card v2026-10-share.

Agent work in a Mnemom agent session

Governance

The inference share

API products with a fixed price

Work that needs no model call

What is not charged

  • Your own model inference. When you send requests through the gateway with your own provider key (BYOK), the model call you asked for is passed through and your provider bills you for it. Mnemom does not resell it. Outside a Mnemom agent session there is no inference share. Inside one, the share replaces the measured cost of the gateway’s own model calls only when the share is larger.
  • Work that needs no model call, as listed in the last table.
  • Everything else on this site (agent registration, alignment and protection cards, the dashboard, reputation lookups) has no per-call charge today.

A worked example

A request in a Mnemom agent session forwards a model call that your provider prices at 2.50(250μ).4tokenscost2.50 (250 μ). 4% of that is 10 μ. The same request triggers a context fold whose model tokens cost 0.03 (3 μ). At 1× that is 3 μ. Governance is on for this session. The integrity analysis uses 0.002(0.2μ)ofmodeltokens.At1×thatis0.2μ.OnesurfaceescalatestotheSafeHouselevel2model,andthatcalluses0.002 (0.2 μ) of model tokens. At 1× that is 0.2 μ. One surface escalates to the Safe House level 2 model, and that call uses 0.0004 (0.04 μ). At 1× that is 0.04 μ. The request’s charge is the larger of the two, 10 μ: the 3.24 μ of measured lines plus a 6.76 μ inference-share line. The request costs 10 μ in total. Had the measured lines come to more than the share, the request would have been charged their own cost and there would be no inference-share line.

When your balance runs out

There is no overdraft and no credit line: your balance never goes below 0 μ. A charge your balance cannot cover is refused, and the gateway rejects requests with 402 Payment Required (see Errors) until you add μ. To avoid a gap, turn on auto top-up before you need it.

When μ expires

Each purchase or grant of μ is kept as a separate lot, and each lot has a 12-month expiry clock. The clock starts when you buy or receive the lot. Every charge except a settled reservation (below) restarts the clock on each lot it draws from. A lot that goes 12 months without such a charge expires, and its unspent μ is removed. The exception: a ResearchGrade run reserves μ when it starts and settles the reservation when it completes. Settling draws from your lots but does not restart their clocks. If ResearchGrade runs are all you spend μ on, a lot still expires 12 months after you bought it, or 12 months after the last other charge drawn from it. Granted μ is spent before purchased μ, and purchased μ is spent oldest first, so in normal use your oldest lots are the ones being drawn down.

Checking your balance and usage

  • Dashboard: www.mnemom.ai/settings/billing shows your current μ balance, and /settings/billing/usage breaks down consumption over time.
  • API: GET /billing/mu returns your org’s current μ balance, GET /billing/mu/ledger returns every ledger entry (grants, holds, settlements, refunds, expirations) behind that balance, and GET /billing/mu/statement returns an itemized monthly statement for a given YYYY-MM period. See the Mu balance, Mu ledger, and Mu statement endpoint references.
  • Payment methods and receipts: GET /billing/payment-methods lists your saved payment instruments (tokenized, so full card numbers are never returned) and GET /billing/receipts lists your completed μ purchases. See the Payment methods and Receipts references.
  • CLI: mnemom usage --org <org_id> shows per-person token and request consumption for an org, filterable by --days, --person, --provider, and --model. See the CLI reference.

Enterprise

An annual μ commitment unlocks a volume discount off the list peg. Retail purchases always stay at list price, with no discount tiers: Enterprise also adds self-hosting (run the trust plane in your own infrastructure, same μ accounting and peg) and dedicated cells (an isolated deployment in the region you choose). Contact sales for a quote.

FAQ

μ is the unit you pay in. One μ is one US cent, and the peg is fixed.
The model calls the gateway makes for you, each at its measured model cost. In a Mnemom agent session you also pay the inference share (4% of the request’s forwarded inference spend) when it is larger than the agent work’s cost. Fixed-price API products are listed above. Work that needs no model call is free.
Level 1 screening, which runs rules and makes no model call, is free. When a surface escalates to level 2, that model call is charged at its measured cost.
Create an account, add a payment method, and buy μ. Nothing is charged until you buy.
Your balance never goes below 0 μ. A charge it cannot cover is refused, and the gateway rejects requests with 402 Payment Required until you add μ. There is no overdraft. Buy more μ in one click at any time, or turn on auto top-up so your balance refills once it drops below a threshold you set.
Yes. Each lot of μ expires after 12 months without a charge drawn from it. The clock starts at purchase, and every charge except a settled ResearchGrade reservation restarts it. See “When μ expires” above.
In a Mnemom agent session, each successful request is charged the larger of two amounts: the measured cost of the gateway’s own model calls for it (context folds, the goal judge, implicit-goal calls and, when governance is on, the integrity analysis and Safe House level 2 screens), or 4% of the request’s forwarded inference spend at the provider’s published rates. When the 4% is larger, your statement shows an “Inference share” line for the difference. A gateway call that finishes after the share was charged counts toward the same comparison from 2026-10-04 00:00:01 UTC. Observer trace analysis runs later from your logs. From 2026-10-04 00:00 UTC, its cost for a request counts toward that same comparison, so you pay the larger of the 4% and all of the request’s measured costs. Before that date, and for analysis not tied to a request, it is charged at measured cost on top.
No. Every feature is available to every account regardless of spend. Buying more μ buys more work done, not more product.
Your dashboard’s billing pages show your balance and usage; GET /billing/mu, GET /billing/mu/ledger, and GET /billing/mu/statement expose the same data by API, and mnemom usage shows per-person consumption from the CLI.