Help
Troubleshooting and FAQ
Common problems, what causes them and how to fix them. If you are still stuck, write to us with the request's x-request-id (see Contact support).
Keys and access
401 MISSING_API_KEY
The API found no key in the request. Send it as Authorization: Bearer <key>. Other headers, such as x-api-key, are not read. An empty variable sends Bearer with nothing after it, which gets this error too. See Authentication and API keys.
401 INVALID_API_KEY
The key is unknown or was revoked. Check that your code gets the whole key, with no quotes or spaces around it:
echo "${ACACUS_API_KEY:0:8}" # prints sk-shfr-
echo -n "$ACACUS_API_KEY" | wc -c # prints 48If the key looks right and still fails, it was probably revoked: create a new one on the API keys page.
403 with the text error code: 1010
This comes from Cloudflare, in front of the API, not from the API itself. It refuses the default User-Agent of some HTTP clients, such as Python's urllib and Perl's LWP. Send your own User-Agent header, or use curl, requests, httpx, Node.js fetch or the OpenAI libraries. See Libraries and SDKs.
403 ACCOUNT_SUSPENDED
The account is suspended, and so are all its keys. The suspension object in the error says until when. Write to us on WhatsApp about it.
Money and the free allowance
402 although my API balance has money
Before a paid request runs, your balance must cover the most it could cost: every prompt token at the uncached price, plus a reply of the full max_tokens. If it doesn't, max_tokens is lowered to what the balance covers, down to 8,192 tokens. If even that is too much, the request is tried on the free allowance. A refusal from there has a FREE_TIER_ code and a requiredLyd field: the balance that would run the request as paid.
Lower max_tokens, send less input, or top up. See Billing and the free tier.
A paid request stopped at 2,048 tokens
Your balance didn't cover the request (see above), so it ran on the free allowance: deepseek-v4-flash only, replies of at most 2,048 tokens, and no thinking. On the Usage page such requests are marked “Free allowance”. Top up to run them as paid.
402 FREE_TIER_DAILY_LIMIT
The free allowance is per account and per UTC day, and the Acacus chat's daily free allowance draws on it too. The reason field says what happened:
- No reason: today's allowance is used up. It starts again at 00:00 UTC.
allowance: what is left is not enough for this request. The check counts the new message at the uncached price and the reply at its fullmax_tokens(up to 2,048 tokens on the free tier), so a lowermax_tokensor a shorter new message may still fit. This can happen at the start of the day too, to a long new message: at today's settings, one of more than about 65,000 tokens (about 25,000 tokens in DeepSeek's peak hours). See the largest new message.request_too_large: the request is over the free size limit of 80,000 prompt tokens, or 320,000 bytes of JSON not counting images. Send less input.
During DeepSeek's peak hours (01:00 to 04:00 and 06:00 to 10:00 UTC, Monday to Friday), free requests use up the allowance twice as fast. Paid requests cost the same at every hour.
402 FREE_TIER_EXHAUSTED
The free tier as a whole is used up for today, or it is paused for now (reason paused). Top up to continue, or try again later.
402 FREE_TIER_MODEL or FREE_TIER_THINKING
A free request asked for deepseek-v4-pro, or for thinking: a thinking field with a type other than disabled, or, with no thinking field, a reasoning_effort other than none. Use deepseek-v4-flash without thinking, or top up. See Reasoning.
Where can I see what a request cost?
The reply counts tokens, but it has no price. The Usage page lists each request with its cost: set Channel to API to see only API requests. GET /v1/usage returns your token totals.
I stopped reading a stream. Was it billed?
Yes. The reply is still generated to the end, and it is billed in full. To limit what one reply can cost, set max_tokens. See Streaming.
Limits and timeouts
429 rate_limit_error
The account sent more than 30 requests in one clock minute. All its keys and its use of the Acacus chat count together. Wait the number of seconds in the Retry-After header. The X-RateLimit-Remaining header of each chat completion reply says how many requests are left this minute. See Rate limits and other limits.
429 concurrency_limit
The account already has 3 requests in flight. Wait for one to finish, then send the next. This 429 has no Retry-After header.
413 Request Entity Too Large, as an HTML page
The request body is over 2 MB. The web server in front of the API answers before the API sees it, so the reply is HTML, not JSON. Images sent as data URLs are part of the body, and base64 makes a file about a third larger: send smaller or fewer images. See Images.
504 on a long reply
A reply that isn't streamed sends nothing until it is complete, and nginx, the server in front of the API, closes a connection when no data comes back for 120 seconds. It answers 504 with an HTML page. The reply is still generated and billed. Stream long replies, large max_tokens values and requests with thinking. The OpenAI libraries retry a 504 by themselves, and each retry is a new request that is billed too. See Streaming.
503 UPSTREAM_DOWN
DeepSeek is unavailable to us right now. Wait a minute and try again. You are not charged for it.
Replies
The model field says deepseek-flash
That is DeepSeek's own name for the model that serves deepseek-v4-flash. You are billed for the id you sent. See Models.
400 MODEL_NOT_FOUND
Model ids are exact: deepseek-v4-flash or deepseek-v4-pro. DeepSeek's own names, such as deepseek-chat, don't work here. The error message says where the ids your key can use are listed: GET /v1/models, sent with your key.
The answer is empty or cut off, and finish_reason is length
The reply reached max_tokens. Reasoning counts inside max_tokens, so a request with thinking on can use all of it before it answers. Free requests stop at 2,048 tokens. Raise max_tokens, or turn thinking off with "thinking": {"type": "disabled"}. See Reasoning.
400 “This response_format type is unavailable now”
The API doesn't support json_schema response formats, which the SDKs' parse() helpers and LangChain's default with_structured_output send. Use {"type": "json_object"}, or a tool with a JSON schema. See JSON output.
An image was refused
Images work with deepseek-v4-flash only (else 400 MODEL_NO_VISION), in user messages only (else IMAGE_NOT_IN_USER_MESSAGE), and up to 32 images in one request (else TOO_MANY_IMAGES). A 400 “Failed to download image” means DeepSeek couldn't fetch an https image: send the image as a data URL instead. See Images.
Connecting
Can I call the API from a web page or a mobile app?
Not directly. The API sends no CORS headers, so browsers block the call, and a key inside a page or an app can be read by anyone who uses it. Call the API from your server, and pass the answer on.
Does the Acacus chat share my allowance and limits?
The limits, yes: the chat and the API share them per account. The free daily allowance too, while you have no chat plan: the chat's daily free allowance comes out of it. The chat is never charged to your API balance.
Which endpoints are there?
POST /v1/chat/completions, GET /v1/models, GET /v1/models/{id} and GET /v1/usage. Every other path gets 404, including embeddings and OpenAI's Responses API. See the Overview.
Contact support
Write to [email protected] or on WhatsApp (+218 94 380 1609). Say what you sent and what you got back, and give the request's x-request-id: every reply from the API has this header. When an error comes from DeepSeek, its message also ends with DeepSeek's request_id; send that too. A reply from the servers in front of the API (a 413, a 403 with error code: 1010, or a 5xx HTML page) has no x-request-id: give its cf-ray header and the time instead.
curl -i https://acacus.ly/v1/chat/completions \
-H "Authorization: Bearer $ACACUS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek-v4-flash",
"messages": [{"role": "user", "content": "Say hello."}],
"thinking": {"type": "disabled"}
}'With curl, look for the x-request-id line among the headers. In the OpenAI libraries, errors carry it too: e.request_id in Python and err.requestID in Node.js. You can also send your own x-request-id header, and the reply carries it back unchanged.