Models

Merki hosts a catalog of open and licensed models, and lets you bring your own key for providers you already use. Both routes are served by the same proprietary Merki inference stack, which is why throughput is high and prices are low.

Hosted models

Any model in the catalog can be called directly. You pay in credits at the listed per-million-token rates. Calls go through endpoints that are compatible with OpenAI and Anthropic clients. See Endpoints and compatibility.

Bring your own key

You can also route through your own provider key. See Bring your own key.

How to read the catalog

The catalog is a table. One row per model and quantization. Columns:

ColumnMeaning
ModelThe model name, as published by its author.
QuantThe quantization Merki serves, for example FP4, FP8, BF16, INT4, GGUF.
ContextMaximum context window, in tokens.
Input $/MPrice per million input tokens, in US dollars.
Output $/MPrice per million output tokens, in US dollars.
ModalitiesWhat the model accepts, for example text, image, video.
Tok/sOutput tokens per second, measured by Merki.
TTFTTime to first token.
UptimeMeasured availability over the reporting period.
Model cardLinks to the upstream model card and an independent speed benchmark.

Prices reflect the proprietary Merki serving stack: every hosted price sits 25–35% below the cheapest published list price for the same model. Cache hits are billed at 50% of the listed input rate. See Pricing and Caching. Where a cell is not available it reads n/a.

Snapshots

The canonical table is the model catalog. Dated snapshots are frozen monthly, for example the August 2025 snapshot. The current page always holds the latest lineup.