Usage and metering

Every request draws down credits at the model catalog rates.

Quick path

  1. Tokens are counted per request: prompt plus completion.
  2. Cost = input tokens × input rate + output tokens × output rate. Cache hits halve the input rate.
  3. Reconcile in usage: per-request rows with model, tokens, and cost.

Details

TopicDecision
CountingTokenizer of the serving model. Prompt tokens include messages, tools, and media tokens where the model accepts them.
Cache hitsMatched input tokens bill at 50% of the input rate. See Caching.
Partial deliveryOnly tokens actually delivered are billed — including revoked-key cutoffs and mid-stream failures.
RefusalsA refused request bills nothing. See Content safety.
Validation errorsA 400 before inference bills nothing.
CurrencyUSD. Balances and rows are in dollars to four decimals. See Pricing.

Checklist

  • [ ] You can map every charge to a request row.
  • [ ] Cache-hit savings are visible in the rows.
  • [ ] Failed-stream retries are not double-billed for undelivered tokens.

Next step

Control repeat-work cost: Caching.