Rate limits

Rate limits protect shared capacity. They are per key and per account, per minute.

Quick path

  1. Stay under the limits for your tier.
  2. Watch RateLimit-* headers and honor Retry-After on 429.
  3. Need more: top up history and Enterprise raises the ceiling. See Enterprise.

Details

TopicDecision
DimensionsRequests per minute (RPM) and tokens per minute (TPM), enforced per key and per account. The lower of the two ceilings applies.
Default ceilingsRegular: 60 RPM / 200K TPM. Roleplay and Cybersecurity: same defaults; higher on request with usage history. Exact ceilings are returned in response headers.
HeadersRateLimit-Limit, RateLimit-Remaining, RateLimit-Reset. On 429, Retry-After in seconds.
Over limit429 rate_limited. Back off; do not retry in a tight loop. See Errors.
SLARate-limited requests are customer errors and are excluded from availability measurement. See the Service level agreement.

Checklist

  • [ ] Client reads Retry-After and backs off with jitter.
  • [ ] Bursts are smoothed to stay under TPM.
  • [ ] Sustained need is raised via Enterprise, not via extra accounts.

Next step

Handle failures cleanly: Errors.