Skip to content
أكاكوسAcacus
تواصل مع المبيعاتContact sales
Docs menu

API reference

Rate limits and other limits

Each account can send a set number of requests a minute, and have a set number running at the same time. Requests also have limits on size and time. This page lists them, and what to do when you reach one.

Requests per minute

Each account can send 30 requests a minute to POST /v1/chat/completions. The count starts again at the start of each clock minute: it is a fixed window, not a rolling one.

  • Every request whose key is accepted counts, even one the API then refuses, such as bad JSON, an unknown model or a 402.
  • All the keys of an account share one count. The account's use of the Acacus chat counts too.
  • Not counted: GET /v1/models, GET /v1/models/{id}, GET /v1/usage, and requests with a missing or invalid key.

Every reply from the endpoint, once the key is accepted, says where you stand:

HeaderWhat it says
X-RateLimit-LimitRequests your account may send per minute.
X-RateLimit-RemainingRequests left in the current minute.
X-RateLimit-ResetWhen the current minute ends, in Unix seconds. The count starts again then.
Retry-AfterOnly on a 429 for this limit: seconds until the minute ends.

To see them, send any request and print the headers:

curl -sS -D - -o /dev/null https://acacus.ly/v1/chat/completions \
  -H "Authorization: Bearer $ACACUS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek-v4-flash",
    "messages": [{"role": "user", "content": "Say OK."}],
    "max_tokens": 5,
    "thinking": {"type": "disabled"}
  }' | grep -i '^x-ratelimit'
Terminal
x-ratelimit-limit: 30
x-ratelimit-remaining: 23
x-ratelimit-reset: 1790380500

Read X-RateLimit-Limit instead of writing the number into your code: the limit can change.

Requests at the same time

Each account can have 3 requests running at once. A request takes its place just before it goes to DeepSeek and gives it back once it is billed, just after the reply ends. A request sent while 3 requests are running gets 429 concurrency_limit, with no Retry-After.

  • All the keys of an account, and its Acacus chat, share these places.
  • A streamed reply you stop reading keeps its place until DeepSeek finishes it: the API still reads the whole reply and bills it.
  • A request the API refuses before it calls DeepSeek, such as a 400, 401 or 402, never takes a place.
  • If a place is ever not given back, for example when the server restarts during a request, it clears 130 seconds after the last request of yours that got a place. Requests refused with concurrency_limit in the meantime don't delay it.

When you get a 429

Both limits answer 429 with type rate_limit_error. Tell them apart by code.

Over the per-minute limit, code is rate_limit_error and Retry-After says how many seconds to wait. This is the 31st request in one minute:

Headers
retry-after: 55
x-ratelimit-limit: 30
x-ratelimit-remaining: 0
x-ratelimit-reset: 1790380440
Body
{
  "error": {
    "message": "Rate limit exceeded: at most 30 requests a minute per account. Try again in 55 seconds.",
    "type": "rate_limit_error",
    "code": "rate_limit_error",
    "param": null,
    "limit": 30
  }
}

Too many at the same time, code is concurrency_limit. This is what the fourth and fifth of five requests sent at once got:

Body
{
  "error": {
    "message": "Too many requests in flight for this account: wait for a previous reply to finish.",
    "type": "rate_limit_error",
    "code": "concurrency_limit",
    "param": null
  }
}
  • rate_limit_error: wait Retry-After seconds, then send the request again.
  • concurrency_limit: send it again when one of your running requests has finished.
  • Better still, keep under the limits: send at most 3 requests at once from all your code together, and slow down when X-RateLimit-Remaining gets low.

The OpenAI libraries retry a 429 on their own, 2 times by default, and wait the Retry-After time when there is one (checked in openai-python 3.19.2 and openai-node 7.23.0). Each retry counts as a request. See Errors.

Other limits

LimitValueNotes
Request body2 MBA larger body gets 413 with an HTML page from nginx, not JSON. Images in base64 count toward it: see Images.
Images per request32 imagesMore: 400 TOO_MANY_IMAGES. Only on deepseek-v4-flash.
Prompt tokens per imageup to 1,024 tokensAs DeepSeek counts them.
Reply length (max_tokens)8,192 tokens by default, 32,768 tokens at mostA larger value is lowered to 32,768 tokens, with no error. Reasoning counts inside it.
stop strings16 stringsDeepSeek refuses more (400).
Choices (n)1 choiceDeepSeek refuses other values (400).
Context window1,000,000 tokensDeepSeek's figure for both models. It was not tested near that size; for most requests the 2 MB body limit comes first.
Free requests2,048 reply tokens; 80,000 prompt tokens in all, and about 65,000 tokens in the new message (25,000 tokens in DeepSeek's peak hours); 320,000 bytes of JSON, image data not countedThe new message's limit comes from the daily allowance: see the largest new message.
Time with nothing sent back120 secondsThen 504, an HTML page. See Timeouts.

Timeouts

nginx, the server in front of the API, closes a request when nothing has come back for 120 seconds, and answers 504 with an HTML page. Cloudflare, in front of nginx, waits up to 125 seconds by its own documentation, so the 120 seconds are the limit a request meets.

A reply that isn't streamed sends nothing until it is complete. A long reply, a large max_tokens or thinking can take longer than 120 seconds. The request then fails with that 504, but DeepSeek still writes the reply and the API still bills it. The OpenAI libraries retry a 504 on their own, and each try is billed. Use "stream": true for long replies: a stream sends each piece as it is written, so the limit only applies to the time between two pieces. See Streaming.

Higher limits

The limits are the same for every account, and they can't be raised for one account.