Caching

Caching is the single exception to Zero data retention: to serve a repeat fast, Merki must keep the cached entry.

Quick path

  1. Send the same prompt prefix again. A hit returns faster and bills half input on matched tokens.
  2. Entries live 24 hours or until superseded.
  3. Entries are scoped to your account and never shared.

Details

TopicDecision
What matchesExact prompt-prefix match on the same model and quantization. A hit returns the cached completion path without re-running the full prefix.
Lifetime24 hours from write, or until superseded by a newer entry. There is no manual purge API; expiry is automatic.
ScopePer account, per model. Other customers never see your entries.
BillingMatched input tokens at 50% of the input rate; output tokens at the full output rate. See Pricing and Usage.
RetentionCache entries are the only retained request data. Everything else follows ZDR. See the retention matrix.
Opt-outCaching is on by default because it is the speed path. Per-request opt-out: send X-Merki-No-Cache: 1; the request bypasses cache read and write.

Checklist

  • [ ] Repeated system prompts and few-shot prefixes are structured for prefix hits.
  • [ ] Sensitive prompts that must never persist use short-lived flows until opt-out ships.
  • [ ] Cost models assume the 50% hit rate on repeated prefixes.

Next step

Price the work: Pricing.