Merki Inference Pte. Ltd.
Production inference,
measured.
Merki runs your models on managed infrastructure and reports latency, throughput, and cost for every request. Hosted models or bring your own key, on one endpoint.
Summarize the quarterly report in one sentence.
Revenue grew 12% on stronger renewals, with churn flat.
Hosted models
Proprietary serving, priced below list.
Our own inference stack is why every price sits 25 to 35 percent below the cheapest published list price for the same model. No request fees, no minimums.
| Model | Quant | Context | Input $/M | Output $/M | Tok/s | TTFT |
|---|---|---|---|---|---|---|
| DeepSeek-V3 (0324) | FP8 | 128K | $0.200 | $0.800 | 68.4 | 0.92 s |
| DeepSeek-R1 | FP8 | 128K | $0.400 | $1.600 | 41.2 | 1.45 s |
| GLM-4.5 | BF16 | 128K | $0.300 | $1.200 | 72.8 | 0.88 s |
| GLM-4.5 | FP8 | 128K | $0.300 | $1.200 | 88.5 | 0.71 s |
| Kimi-K2 | BF16 | 256K | $0.400 | $1.600 | 52.3 | 1.10 s |
| Kimi-K2 | FP8 | 256K | $0.400 | $1.600 | 64.7 | 0.95 s |
| MiniMax-M1 | BF16 | 1M | $0.450 | $1.800 | 48.9 | 1.25 s |
| MiniMax-M1 | FP8 | 1M | $0.450 | $1.800 | 59.1 | 1.02 s |
| Mistral Large 2 | BF16 | 128K | $1.500 | $4.500 | 55.6 | 0.98 s |
| Mistral Small 3.1 | BF16 | 128K | $0.300 | $0.900 | 96.3 | 0.62 s |
| Qwen3-235B-A22B | BF16 | 128K | $0.250 | $1.000 | 44.8 | 1.38 s |
| Qwen3-235B-A22B | FP8 | 128K | $0.250 | $1.000 | 58.2 | 1.05 s |
| Qwen3-235B-A22B | INT4 | 128K | $0.250 | $1.000 | 71.5 | 0.89 s |
| Qwen3-32B | BF16 | 128K | $0.120 | $0.480 | 82.4 | 0.74 s |
| Qwen3-32B | INT4 | 128K | $0.120 | $0.480 | 104.6 | 0.58 s |
| Qwen3-32B | GGUF | 128K | $0.120 | $0.480 | 97.8 | 0.63 s |
| Llama-4-Scout | BF16 | 10M | $0.152 | $0.544 | 94.0 | 0.84 s |
| Llama-4-Scout | GGUF | 10M | $0.152 | $0.544 | 94.0 | 0.84 s |
| Gemma-3-27B | BF16 | 128K | $0.100 | $0.400 | 89.2 | 0.66 s |
| Gemma-3-27B | GGUF | 128K | $0.100 | $0.400 | 101.5 | 0.55 s |
| Phi-4 | BF16 | 128K | $0.100 | $0.400 | 105.3 | 0.51 s |
| Phi-4 | FP8 | 128K | $0.100 | $0.400 | 118.7 | 0.44 s |
Endpoints
Point your existing SDK at Merki.
Change the base URL and the key. Harnesses, IDE integrations, and agent frameworks keep working, and BYOK routes work through the same endpoints.
Start with the quickstartAccess
Three tiers, sized to the risk.
Age is gated by content and identity is verified for accountability, so each tier only asks for what it needs. The gate is what separates them.
No verification
Open to anyone, at any age. Hosted models and BYOK.
Regular
18 and over
Age assurance through Sumsub. Serves adult content.
Roleplay
Identity plus domain
Identity verified, and you prove control of a domain to point Merki at live or remote servers.
Cybersecurity
Billing
Credits, and nothing else.
No subscriptions, and no limit on how many accounts you create. Bonus caps are per account.
See full pricingCredits, not subscriptions
You buy credits and every request draws them down. There is no plan to game by opening another account.
5% more above 100 USD
A single top-up of 100 USD or more earns 5% in bonus credits, capped at 300 USD per account. Bonus credits are spent first.
Cache hits are half price
Repeated prefixes hit the cache and bill at 50% of the listed input rate.
Data
Nothing kept, except the cache.
Prompts and completions are not retained. Inference runs on Merki infrastructure and is never routed to a third-party provider.
Zero data retention
Prompts and completions are not stored after a request is served.
Caching is the exception
Cache entries live 24 hours, are scoped to your account, and are never shared.
Keys revoke themselves
A key committed to git, or found on the public internet, is revoked automatically.
See what your inference actually costs.
Tell us what you are running. We will tell you what it takes.