API reference
Chat completions
Send a conversation and get the model's reply. This page lists every field of the request and of the reply, and says which fields the API checks, changes, passes on to DeepSeek or refuses.
Endpoint
Send a JSON object in the body, with your API key in the Authorization header.
| Header | Value |
|---|---|
Authorization | Bearer sk-shfr-.... Required. See Authentication and API keys. |
Content-Type | application/json. The body is read as JSON in any case. |
x-request-id | Optional. Your own id for the request. The reply carries the same value in its x-request-id header. |
A request with a system message, a user message and a few options. It turns thinking off, so it also runs on the free allowance:
curl https://acacus.ly/v1/chat/completions \
-H "Authorization: Bearer $ACACUS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek-v4-flash",
"messages": [
{"role": "system", "content": "You are a helpful assistant. Answer in one short sentence."},
{"role": "user", "content": "Name three cities in Libya."}
],
"max_tokens": 100,
"temperature": 0.7,
"thinking": {"type": "disabled"}
}'The reply the curl example got:
{
"id": "555f7067-3b95-4bd2-a686-3b7a50755d7d",
"object": "chat.completion",
"created": 1790380150,
"model": "deepseek-flash",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "Tripoli, Benghazi, and Misrata are three cities in Libya."
},
"logprobs": null,
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 23,
"completion_tokens": 16,
"total_tokens": 39,
"prompt_tokens_details": {
"cached_tokens": 0
},
"prompt_cache_hit_tokens": 0,
"prompt_cache_miss_tokens": 23
},
"system_fingerprint": "aeb56401ca74e127821c4f9126dcb669"
}Request body
Only messages is required. The gateway checks model, max_tokens, max_completion_tokens, stream and images itself, and answers with its own error codes. It passes the other fields to DeepSeek, which checks them; DeepSeek's errors come back as DeepSeek sends them. See Errors.
deepseek-v4-flash or deepseek-v4-pro. See Models. Ids are exact: any other id, such as deepseek-chat or deepseek-reasoner, gets 400 MODEL_NOT_FOUND. A value that isn't a string gets 400 INVALID_PARAMETER.Default: deepseek-v4-flash400) and a request without it (422).finish_reason "length". It must be a whole number of at least 1, else 400 INVALID_PARAMETER. A value above 32,768 tokens is lowered to 32,768 tokens, with no error. Free requests get at most 2,048 tokens. When your balance can't cover a paid request, the API may lower it to what the balance covers, but not below 8,192 tokens: see Billing and the free tier.Default: 8,192 tokensmax_tokens, with the same rules. If you send both, max_tokens is used.true sends the reply in pieces as it is written, as Server-Sent Events. See Streaming. Send true, false or null; anything else gets 400 INVALID_PARAMETER.Default: falseinclude_usage is always set to true, whatever you send, so the last chunk always carries usage.{"type": "enabled"}, {"type": "adaptive"} or {"type": "disabled"}. A paid request without it thinks, unless reasoning_effort is none, and the reasoning is billed as output. On the free tier, anything but disabled gets 402 FREE_TIER_THINKING. See Reasoning.Default: on for paid requests, off for free oneslow, high or max; none turns thinking off. When thinking is also set, thinking decides. On the free tier, a value other than none gets 402 FREE_TIER_THINKING.400).finish_reason is "stop". DeepSeek refuses 17 or more (400).{"type": "json_object"} makes the reply a JSON object; {"type": "text"} is plain text. DeepSeek refuses {"type": "json_schema"} (400). See JSON output.{"type": "function", "function": {"name", "description", "parameters"}}, with parameters as a JSON Schema. "strict": true is accepted, but whether DeepSeek enforces it was not tested. See Tool calling.auto, none, required, or {"type": "function", "function": {"name": "..."}} to make the model call that function.true adds the log probability of each reply token, in choices[].logprobs.Default: falselogprobs: how many of the most likely tokens to list at each position, from 0 to 20. DeepSeek refuses larger values (400).400).Default: 1400 INVALID_PARAMETER. They are the old form of tools and tool_choice; use those. An empty functions list is accepted.Other fields go to DeepSeek unchanged. This page doesn't describe what DeepSeek does with them.
Messages
Each item of messages is one turn of the conversation. The API keeps no history between requests: send the whole conversation every time.
system, user, assistant or tool. DeepSeek refuses a role it doesn't know (422).tool_calls. In a user message it can also be a list of parts: {"type": "text", "text": "..."} and, for deepseek-v4-flash only, {"type": "image_url", "image_url": {"url": "data:image/png;base64,..."}}. See Images.id of the tool call it answers.Images: at most 32 images a request, in user messages only, and not to deepseek-v4-pro. Each of these rules has its own error code (below).
Changed, ignored and refused
The differences from what OpenAI's API does with the same request, in one place:
| Field | What happens |
|---|---|
max_tokens above 32,768 tokens | Lowered to 32,768 tokens. |
max_tokens on the free tier | Lowered to at most 2,048 tokens. |
stream_options.include_usage | Always true. |
user_id | Replaced with an id made from your account. |
thinking.budget_tokens | Ignored by DeepSeek. |
n other than 1 | 400 from DeepSeek. |
response_format json_schema | 400 from DeepSeek. |
stop with 17 or more strings | 400 from DeepSeek. |
functions, function_call | 400 INVALID_PARAMETER. |
Images to deepseek-v4-pro | 400 MODEL_NO_VISION. |
| Images outside user messages | 400 IMAGE_NOT_IN_USER_MESSAGE. |
| More than 32 images | 400 TOO_MANY_IMAGES. |
Thinking, or a model other than deepseek-v4-flash, on the free tier | 402 FREE_TIER_THINKING or FREE_TIER_MODEL. |
The reply
A reply that isn't streamed is one JSON object: DeepSeek's reply, passed on as DeepSeek sends it.
chatcmpl-....chat.completiondeepseek-flash for deepseek-v4-flash. You are billed for the model you asked for.n is always 1.0assistantmax_tokens.index, id, type ("function"), function.name and function.arguments, the arguments as a JSON string.null unless the request set "logprobs": true. Then content lists each reply token with token, logprob, bytes and top_logprobs.stop (it was complete, or reached a stop string), length (it reached max_tokens) or tool_calls (it calls a tool). Treat any other value as an incomplete reply.prompt_cache_hit_tokens, where OpenAI's format puts it.The reply doesn't include its price. The Usage page lists the cost of each request.
Log probabilities
With "logprobs": true and "top_logprobs": 2, a one-token reply carried this:
"logprobs": {
"content": [
{
"token": "Yes",
"logprob": -0.000051020274,
"bytes": [89, 101, 115],
"top_logprobs": [
{ "token": "Yes", "logprob": -0.000051020274, "bytes": [89, 101, 115] },
{ "token": "yes", "logprob": -9.89327, "bytes": [121, 101, 115] }
]
}
]
}Tool calls
With a get_weather function in tools, the question “What is the weather in Tripoli?” got this choice. Your code runs the function and sends the result back in a tool message: see Tool calling.
{
"index": 0,
"message": {
"role": "assistant",
"content": "I'll check the current weather in Tripoli for you.",
"tool_calls": [
{
"index": 0,
"id": "call_00_l0FD1WqIMX4lwllRfZaa5995",
"type": "function",
"function": {
"name": "get_weather",
"arguments": "{\"city\": \"Tripoli\"}"
}
}
]
},
"logprobs": null,
"finish_reason": "tool_calls"
}Streamed replies
With "stream": true, the reply comes as Server-Sent Events: lines that start with data: , each holding one JSON chunk and followed by a blank line. The last line is data: [DONE]. This is a whole stream:
data: {"id":"2fdb0a67-522b-44f2-b79f-2101790a929a","object":"chat.completion.chunk","created":1790380870,"model":"deepseek-flash","system_fingerprint":"aeb56401ca74e127821c4f9126dcb669","choices":[{"index":0,"delta":{"role":"assistant","content":""},"logprobs":null,"finish_reason":null}],"usage":null}
data: {"id":"2fdb0a67-522b-44f2-b79f-2101790a929a","object":"chat.completion.chunk","created":1790380870,"model":"deepseek-flash","system_fingerprint":"aeb56401ca74e127821c4f9126dcb669","choices":[{"index":0,"delta":{"content":"Good"},"logprobs":null,"finish_reason":null}],"usage":null}
data: {"id":"2fdb0a67-522b-44f2-b79f-2101790a929a","object":"chat.completion.chunk","created":1790380870,"model":"deepseek-flash","system_fingerprint":"aeb56401ca74e127821c4f9126dcb669","choices":[{"index":0,"delta":{"content":" morning"},"logprobs":null,"finish_reason":null}],"usage":null}
data: {"id":"2fdb0a67-522b-44f2-b79f-2101790a929a","object":"chat.completion.chunk","created":1790380870,"model":"deepseek-flash","system_fingerprint":"aeb56401ca74e127821c4f9126dcb669","choices":[{"index":0,"delta":{"content":""},"logprobs":null,"finish_reason":"stop"}],"usage":{"prompt_tokens":12,"completion_tokens":2,"total_tokens":14,"prompt_tokens_details":{"cached_tokens":0},"prompt_cache_hit_tokens":0,"prompt_cache_miss_tokens":12}}
data: [DONE]chat.completion.chunkassistant, in the first chunk.index, id, type and function.name; later pieces add to function.arguments. Join the pieces that share an index.null until the last chunk.null in every chunk but the last, which carries the usage of the whole request.id, created, model and system_fingerprint repeat in every chunk. Code samples, and what happens when a stream stops early, are in Streaming.
Response headers
| Header | What it says |
|---|---|
Content-Type | application/json, or text/event-stream for a stream. |
x-request-id | On every reply from the API: the id you sent, or one the API made. Quote it when you contact support. Replies from the servers in front of the API, such as a 413, don't have it: see Errors. |
X-RateLimit-Limit | Requests your account may send per minute. |
X-RateLimit-Remaining | Requests left in the current minute. |
X-RateLimit-Reset | When the current minute ends, in Unix seconds. |
Retry-After | On a 429 for the per-minute limit: seconds to wait. |
The X-RateLimit headers come on every reply from this endpoint once the key is accepted. See Rate limits and other limits.
Errors
The gateway's own checks, made before it calls DeepSeek:
| Status and code | Cause |
|---|---|
400 INVALID_JSON | The body isn't a JSON object. |
400 INVALID_PARAMETER | model, max_tokens, max_completion_tokens or stream has a wrong value, or the request sends functions or function_call. param names the field. |
400 MODEL_NOT_FOUND | The model id isn't one this API serves. |
400 MODEL_NO_VISION, IMAGE_NOT_IN_USER_MESSAGE, TOO_MANY_IMAGES | An image the API won't send on. |
401 MISSING_API_KEY, INVALID_API_KEY | No key, or a key that doesn't work. |
402 FREE_TIER_... | Your balance or the free allowance doesn't cover the request. |
429 rate_limit_error, concurrency_limit | Too many requests in the minute, or at the same time. |
A value DeepSeek refuses comes back with DeepSeek's own error. Every error, with what to do, is in Errors.