Z.ai Rate Limits by Tier

What are Z.ai's API rate limits?

Z.ai's entry tier (Default) allows an unpublished number of requests/min and an unpublished number of tokens/min for documented default/model family. Limits scale up through Default as cumulative spend and account age increase — see the full table below, verified 2026-08-14.

Verified 2026-08-14 — source

Limits by tier

TierQualificationModel classRPMTPMRPDConcurrent
DefaultAccount quota shown in the Z.ai consoledocumented default/model family————

— means not documented by Z.ai, never a guess.

What this means for your workload

Classification at volume: 116 calls/min and 60,320 tokens/min at the production profile.

This provider publishes no numeric cap for this workload; check the account console before launch.

Response headers

retry-afterSeconds to wait before retrying, when supplied with a 429
rate-limit response headersProvider-specific remaining and reset counters when documented

When you exceed the limit

Z.ai returns HTTP 429.

Back off and check the model-specific account quota.

FAQ

What happens when I exceed Z.ai's rate limit?

Z.ai returns HTTP 429. Back off and check the model-specific account quota.

How do I request a rate limit increase on Z.ai?

Request one from the account dashboard: https://bigmodel.cn/console

Z.ai provider hubGet a Z.ai API key

Decision and evidence guide. Verified 2026-08-14. These are dated, route-local references; unavailable values are not inferred.

Z.ai observed-limit snapshot ledger

Frozen fixture board. Formula / decision rule: effective value = exact account + key + endpoint + model + plan + window observation Boundary: A sibling key, alias, or plan cannot supply an undocumented value.

Frozen fixtureJoined inputs and observationCalculated resultState
general API key · exact GLM aliasaccount=acct-zai-12; key=fp-zai-g-01; endpoint=api.z.ai; model=GLM-5; plan=general; dimension=RPM; window=minute; console=observed 60
The console value is joined to the account, key class, endpoint, exact alias, and minute window.
RPM=60; sharing edge=account-level only where console states itPASS WITH SCOPE — observed limit.
second key under one account · async jobaccount=acct-zai-12; key=fp-zai-g-02; endpoint=api.z.ai; job=async-14; dimension=concurrency; value=Unavailable; source=response headers
The second key joins the account, but no async concurrency ceiling is returned.
concurrency = Unavailable; no copy from general RPMUNAVAILABLE — live account limit required.
coding-compatible endpoint · enterprise overrideaccount=acct-zai-99; endpoint=code.z.ai; model=GLM-Coder; plan=enterprise; override=undocumented; observed-at=2026-08-14
An operator reports an override, but no console or response evidence joins it.
enterprise value = Unavailable; report only the observation provenanceFAIL CLOSED — undocumented override.

Provenance: zai module 1 first-party evidence and surface verification date 2026-08-14. Z.ai developer limits documentation. Missing joins fail closed.

Z.ai 429 evidence-and-replay classifier

Frozen fixture board. Formula / decision rule: replay = provider evidence ∧ debit visibility ∧ idempotency/effect key Boundary: Backoff changes timing; it cannot make a non-idempotent replay safe.

Frozen fixtureJoined inputs and observationCalculated resultState
request-frequency / token-budget 429HTTP=429; code=rate_limit; request=req-zai-21; tool=none; reset=5s; debit=visible; idempotency=key-z-1
Returned timing and debit evidence identify a transient request/token class.
retry after 5s with bounded jitter; replay=eligibleRETRY ELIGIBLE — identity joined.
partial stream · tool side effectHTTP=429; request=req-zai-22; stream=partial; tool=charge-card; tool_id=tool-88; effect checkpoint=missing
A side-effecting tool call lacks a checkpoint, so replay could duplicate the effect.
retry=blocked regardless of exponential backoffFAIL CLOSED — operator reconciliation.
opaque 429 · insufficient resource or overloadHTTP=429; code=opaque; request=req-zai-23; reset=missing; debit=missing; concurrency=joined
The response cannot distinguish resource entitlement from transient overload.
class={resource, overload}; next action=collect console/support evidenceUNRESOLVED — no automatic replay.

Provenance: zai module 2 first-party evidence and surface verification date 2026-08-14. Z.ai developer limits documentation. Missing joins fail closed.

GLM workload admission board

Frozen fixture board. Formula / decision rule: capacity verdict requires a live account limit for every applicable request/token/concurrency bucket Boundary: Unknown capacity is neither zero nor unlimited.

Frozen fixtureJoined inputs and observationCalculated resultState
100 short callsaccount=acct-zai-12; key=fp-zai-g-01; request reserve=100; token reserve=20000; RPM=60; TPM=Unavailable
Request bucket alone admits only the first 60 calls; token bucket is unknown.
admitted=60; deferred=40; unknown token impactDEFER — token ceiling missing.
ten 32K prompts · four long outputs · 20 parallel toolsprompt=320000; output=64000; tools=20; concurrency=Unavailable; model=GLM-5
Exact model joins, but concurrency and token reserves do not.
admitted=0; deferred=34; binding bucket=UnavailableUNAVAILABLE — live limits required.
asynchronous batch · retry wavejob=async-19; retry_count=3; shared_pool=account; idempotency=partial; reset=Unavailable
The account pool is known, but reset timing and all retry identities are not.
next-safe schedule = Unavailable; do not enqueue retry waveFAIL CLOSED — incomplete replay joins.

Provenance: zai module 3 first-party evidence and surface verification date 2026-08-14. Z.ai developer limits documentation. Missing joins fail closed.

Run this scenario →