Rate limits
The per-minute budget every token has, the headers that report it, and the quotas some methods add.
On this page
Two kinds of limit apply to the API. Every token has a budget of calls per minute, and some methods add a quota of their own on the work they do.
The per-minute budget
Every token may make 240 method calls a minute. The budget belongs to the token, not to the API: an AI client using the same token over MCP spends from the same budget, and two scripts sharing one token share it too.
The minute starts with the first call after the previous minute ran out. Every call to a method counts, including one refused because of a missing scope or a bad argument. These do not count:
GET /v1/token,GET /v1/toolsandGET /v1/openapi.json.- A call to a method name that does not exist, or with a body that is not a JSON object.
- A repeated call answered from an earlier result through its
Idempotency-Key.
Over the budget, a call is refused with a 429 and runs nothing:
{
"success": false,
"message": "Too many tool calls. Wait 27 seconds and try again!",
"data": null,
"code": 429,
"version": "0.0.1740"
}The headers
Every call to a method reports where the token stands:
| Header | What it holds |
|---|---|
X-RateLimit-Limit |
The calls the token may make in a minute |
X-RateLimit-Remaining |
The calls left in the current minute |
X-RateLimit-Reset |
Seconds until the current minute ends |
Retry-After |
On a 429 from the budget, the seconds to wait before calling again |
GET /v1/token reports the same three numbers in data.rateLimit without spending a call.
Stay inside it
- Wait for
Retry-After. After a429from the budget, sleep for the seconds it gives, then carry on. Retrying sooner only gets more refusals. - Pace long loops. When
X-RateLimit-Remaininggets low, slow down rather than running into the wall. - Poll gently. Methods that report progress, such as
get_jobandget_publish_status, only need a call every few seconds. See Long-running work. - Give separate jobs separate tokens. Each token has its own budget, and a token per job also makes it clear which one to revoke.
Quotas on the work itself
Some methods limit the work they do, whatever token calls them. These answer with their own message, and waiting for the work to finish is the fix rather than retrying:
| Quota | Limit | Answer when reached |
|---|---|---|
| Messages waiting in a site's chat queue | 20 per site, and 8 from any one person | 429 from send_message, naming the cap |
| Analytics lookups | 1,200 per site per hour, shared by every token and every AI client reading that site | 429 from the analytics methods |
Calls to one deploy hook, including through trigger_deploy_hook |
10 a minute per hook | The call is refused until the minute passes |
Other methods have size limits on what they take or return, such as the longest prompt or the largest file. Each method's page states its own.
Next
- Errors for every status code.
- Idempotency and retries for retrying writes safely.