Modulify

Rate limits

The per-minute budget every token has, the headers that report it, and the quotas some methods add.

On this page

Two kinds of limit apply to the API. Every token has a budget of calls per minute, and some methods add a quota of their own on the work they do.

The per-minute budget

Every token may make 240 method calls a minute. The budget belongs to the token, not to the API: an AI client using the same token over MCP spends from the same budget, and two scripts sharing one token share it too.

The minute starts with the first call after the previous minute ran out. Every call to a method counts, including one refused because of a missing scope or a bad argument. These do not count:

  • GET /v1/token, GET /v1/tools and GET /v1/openapi.json.
  • A call to a method name that does not exist, or with a body that is not a JSON object.
  • A repeated call answered from an earlier result through its Idempotency-Key.

Over the budget, a call is refused with a 429 and runs nothing:

{
  "success": false,
  "message": "Too many tool calls. Wait 27 seconds and try again!",
  "data": null,
  "code": 429,
  "version": "0.0.1740"
}

The headers

Every call to a method reports where the token stands:

Header What it holds
X-RateLimit-Limit The calls the token may make in a minute
X-RateLimit-Remaining The calls left in the current minute
X-RateLimit-Reset Seconds until the current minute ends
Retry-After On a 429 from the budget, the seconds to wait before calling again

GET /v1/token reports the same three numbers in data.rateLimit without spending a call.

Stay inside it

  • Wait for Retry-After. After a 429 from the budget, sleep for the seconds it gives, then carry on. Retrying sooner only gets more refusals.
  • Pace long loops. When X-RateLimit-Remaining gets low, slow down rather than running into the wall.
  • Poll gently. Methods that report progress, such as get_job and get_publish_status, only need a call every few seconds. See Long-running work.
  • Give separate jobs separate tokens. Each token has its own budget, and a token per job also makes it clear which one to revoke.

Quotas on the work itself

Some methods limit the work they do, whatever token calls them. These answer with their own message, and waiting for the work to finish is the fix rather than retrying:

Quota Limit Answer when reached
Messages waiting in a site's chat queue 20 per site, and 8 from any one person 429 from send_message, naming the cap
Analytics lookups 1,200 per site per hour, shared by every token and every AI client reading that site 429 from the analytics methods
Calls to one deploy hook, including through trigger_deploy_hook 10 a minute per hook The call is refused until the minute passes

Other methods have size limits on what they take or return, such as the longest prompt or the largest file. Each method's page states its own.

Next