# Rate limits

Source: https://modulify.ai/docs/api/rate-limits

The per-minute budget every token has, the headers that report it, and the quotas some methods add.

Two kinds of limit apply to the API. Every token has a budget of calls per minute, and some methods add a quota of their own on the work they do.

## The per-minute budget

Every token may make 240 method calls a minute. The budget belongs to the token, not to the API: an AI client using the same token over MCP spends from the same budget, and two scripts sharing one token share it too.

The minute starts with the first call after the previous minute ran out. Every call to a method counts, including one refused because of a missing scope or a bad argument. These do not count:

- `GET /v1/token`, `GET /v1/tools` and `GET /v1/openapi.json`.
- A call to a method name that does not exist, or with a body that is not a JSON object.
- A repeated call answered from an earlier result through its `Idempotency-Key`.

Over the budget, a call is refused with a `429` and runs nothing:

```json
{
  "success": false,
  "message": "Too many tool calls. Wait 27 seconds and try again!",
  "data": null,
  "code": 429,
  "version": "0.0.1740"
}
```

## The headers

Every call to a method reports where the token stands:

| Header | What it holds |
| --- | --- |
| `X-RateLimit-Limit` | The calls the token may make in a minute |
| `X-RateLimit-Remaining` | The calls left in the current minute |
| `X-RateLimit-Reset` | Seconds until the current minute ends |
| `Retry-After` | On a `429` from the budget, the seconds to wait before calling again |

`GET /v1/token` reports the same three numbers in `data.rateLimit` without spending a call.

## Stay inside it

- **Wait for `Retry-After`.** After a `429` from the budget, sleep for the seconds it gives, then carry on. Retrying sooner only gets more refusals.
- **Pace long loops.** When `X-RateLimit-Remaining` gets low, slow down rather than running into the wall.
- **Poll gently.** Methods that report progress, such as `get_job` and `get_publish_status`, only need a call every few seconds. See [Long-running work](https://modulify.ai/docs/api/long-running-work).
- **Give separate jobs separate tokens.** Each token has its own budget, and a token per job also makes it clear which one to revoke.

## Quotas on the work itself

Some methods limit the work they do, whatever token calls them. These answer with their own message, and waiting for the work to finish is the fix rather than retrying:

| Quota | Limit | Answer when reached |
| --- | --- | --- |
| Messages waiting in a site's chat queue | 20 per site, and 8 from any one person | `429` from `send_message`, naming the cap |
| Analytics lookups | 1,200 per site per hour, shared by every token and every AI client reading that site | `429` from the analytics methods |
| Calls to one deploy hook, including through `trigger_deploy_hook` | 10 a minute per hook | The call is refused until the minute passes |

Other methods have size limits on what they take or return, such as the longest prompt or the largest file. Each method's page states its own.

## Next

- [Errors](https://modulify.ai/docs/api/errors) for every status code.
- [Idempotency and retries](https://modulify.ai/docs/api/idempotency) for retrying writes safely.