Caching
Caching is the single exception to Zero data retention: to serve a repeat fast, Merki must keep the cached entry.
Quick path
- Send the same prompt prefix again. A hit returns faster and bills half input on matched tokens.
- Entries live 24 hours or until superseded.
- Entries are scoped to your account and never shared.
Details
| Topic | Decision |
|---|---|
| What matches | Exact prompt-prefix match on the same model and quantization. A hit returns the cached completion path without re-running the full prefix. |
| Lifetime | 24 hours from write, or until superseded by a newer entry. There is no manual purge API; expiry is automatic. |
| Scope | Per account, per model. Other customers never see your entries. |
| Billing | Matched input tokens at 50% of the input rate; output tokens at the full output rate. See Pricing and Usage. |
| Retention | Cache entries are the only retained request data. Everything else follows ZDR. See the retention matrix. |
| Opt-out | Caching is on by default because it is the speed path. Per-request opt-out: send X-Merki-No-Cache: 1; the request bypasses cache read and write. |
Checklist
- [ ] Repeated system prompts and few-shot prefixes are structured for prefix hits.
- [ ] Sensitive prompts that must never persist use short-lived flows until opt-out ships.
- [ ] Cost models assume the 50% hit rate on repeated prefixes.
Next step
Price the work: Pricing.