Skip to content
أكاكوسAcacus
تواصل مع المبيعاتContact sales
Docs menu

API reference

Chat completions

Send a conversation and get the model's reply. This page lists every field of the request and of the reply, and says which fields the API checks, changes, passes on to DeepSeek or refuses.

Endpoint

POSThttps://acacus.ly/v1/chat/completions

Send a JSON object in the body, with your API key in the Authorization header.

HeaderValue
AuthorizationBearer sk-shfr-.... Required. See Authentication and API keys.
Content-Typeapplication/json. The body is read as JSON in any case.
x-request-idOptional. Your own id for the request. The reply carries the same value in its x-request-id header.

A request with a system message, a user message and a few options. It turns thinking off, so it also runs on the free allowance:

curl https://acacus.ly/v1/chat/completions \
  -H "Authorization: Bearer $ACACUS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek-v4-flash",
    "messages": [
      {"role": "system", "content": "You are a helpful assistant. Answer in one short sentence."},
      {"role": "user", "content": "Name three cities in Libya."}
    ],
    "max_tokens": 100,
    "temperature": 0.7,
    "thinking": {"type": "disabled"}
  }'

The reply the curl example got:

Response
{
  "id": "555f7067-3b95-4bd2-a686-3b7a50755d7d",
  "object": "chat.completion",
  "created": 1790380150,
  "model": "deepseek-flash",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "Tripoli, Benghazi, and Misrata are three cities in Libya."
      },
      "logprobs": null,
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 23,
    "completion_tokens": 16,
    "total_tokens": 39,
    "prompt_tokens_details": {
      "cached_tokens": 0
    },
    "prompt_cache_hit_tokens": 0,
    "prompt_cache_miss_tokens": 23
  },
  "system_fingerprint": "aeb56401ca74e127821c4f9126dcb669"
}

Request body

Only messages is required. The gateway checks model, max_tokens, max_completion_tokens, stream and images itself, and answers with its own error codes. It passes the other fields to DeepSeek, which checks them; DeepSeek's errors come back as DeepSeek sends them. See Errors.

Name
Type
Description
modelOptional
string
The model: deepseek-v4-flash or deepseek-v4-pro. See Models. Ids are exact: any other id, such as deepseek-chat or deepseek-reasoner, gets 400 MODEL_NOT_FOUND. A value that isn't a string gets 400 INVALID_PARAMETER.Default: deepseek-v4-flash
messagesRequired
array
The conversation so far, oldest first. See Messages below. DeepSeek refuses an empty list (400) and a request without it (422).
max_tokensOptional
integer
The most tokens the reply may have, reasoning included. A reply that reaches it ends with finish_reason "length". It must be a whole number of at least 1, else 400 INVALID_PARAMETER. A value above 32,768 tokens is lowered to 32,768 tokens, with no error. Free requests get at most 2,048 tokens. When your balance can't cover a paid request, the API may lower it to what the balance covers, but not below 8,192 tokens: see Billing and the free tier.Default: 8,192 tokens
max_completion_tokensOptional
integer
OpenAI's newer name for max_tokens, with the same rules. If you send both, max_tokens is used.
streamOptional
boolean
true sends the reply in pieces as it is written, as Server-Sent Events. See Streaming. Send true, false or null; anything else gets 400 INVALID_PARAMETER.Default: false
stream_optionsOptional
object
For a streamed request, include_usage is always set to true, whatever you send, so the last chunk always carries usage.
thinkingOptional
object
DeepSeek's switch for its thinking step: {"type": "enabled"}, {"type": "adaptive"} or {"type": "disabled"}. A paid request without it thinks, unless reasoning_effort is none, and the reasoning is billed as output. On the free tier, anything but disabled gets 402 FREE_TIER_THINKING. See Reasoning.Default: on for paid requests, off for free ones
reasoning_effortOptional
string
low, high or max; none turns thinking off. When thinking is also set, thinking decides. On the free tier, a value other than none gets 402 FREE_TIER_THINKING.
temperatureOptional
number
From 0 to 2. DeepSeek refuses other values (400).
top_pOptional
number
Passed to DeepSeek as sent.
stopOptional
string or array
Up to 16 strings. The reply ends just before the first of them it would write, and finish_reason is "stop". DeepSeek refuses 17 or more (400).
response_formatOptional
object
{"type": "json_object"} makes the reply a JSON object; {"type": "text"} is plain text. DeepSeek refuses {"type": "json_schema"} (400). See JSON output.
toolsOptional
array
Functions the model may call, each {"type": "function", "function": {"name", "description", "parameters"}}, with parameters as a JSON Schema. "strict": true is accepted, but whether DeepSeek enforces it was not tested. See Tool calling.
tool_choiceOptional
string or object
auto, none, required, or {"type": "function", "function": {"name": "..."}} to make the model call that function.
logprobsOptional
boolean
true adds the log probability of each reply token, in choices[].logprobs.Default: false
top_logprobsOptional
integer
With logprobs: how many of the most likely tokens to list at each position, from 0 to 20. DeepSeek refuses larger values (400).
presence_penaltyOptional
number
Passed to DeepSeek as sent.
frequency_penaltyOptional
number
Passed to DeepSeek as sent.
seedOptional
integer
Passed to DeepSeek as sent.
userOptional
string
Passed to DeepSeek as sent.
nOptional
integer
Only 1. DeepSeek refuses other values (400).Default: 1
user_idOptional
string
Replaced. The gateway sets it to an id made from your account, which keeps DeepSeek's prompt cache private to you. What you send is dropped.
functions, function_callOptional
array, string or object
Not supported: 400 INVALID_PARAMETER. They are the old form of tools and tool_choice; use those. An empty functions list is accepted.

Other fields go to DeepSeek unchanged. This page doesn't describe what DeepSeek does with them.

Messages

Each item of messages is one turn of the conversation. The API keeps no history between requests: send the whole conversation every time.

Name
Type
Description
roleRequired
string
system, user, assistant or tool. DeepSeek refuses a role it doesn't know (422).
contentRequired
string or array
The text of the message. It can be empty in an assistant message that has tool_calls. In a user message it can also be a list of parts: {"type": "text", "text": "..."} and, for deepseek-v4-flash only, {"type": "image_url", "image_url": {"url": "data:image/png;base64,..."}}. See Images.
nameOptional
string
A name for the participant. Passed to DeepSeek as sent.
tool_callsOptional
array
In an assistant message: the tool calls of an earlier reply, sent back as the reply had them.
tool_call_idOptional
string
In a tool message: the id of the tool call it answers.
reasoning_contentOptional
string
In an assistant message: the reasoning of an earlier reply. DeepSeek's documentation asks for it back in a tool loop with thinking on. See Reasoning.

Images: at most 32 images a request, in user messages only, and not to deepseek-v4-pro. Each of these rules has its own error code (below).

Changed, ignored and refused

The differences from what OpenAI's API does with the same request, in one place:

FieldWhat happens
max_tokens above 32,768 tokensLowered to 32,768 tokens.
max_tokens on the free tierLowered to at most 2,048 tokens.
stream_options.include_usageAlways true.
user_idReplaced with an id made from your account.
thinking.budget_tokensIgnored by DeepSeek.
n other than 1400 from DeepSeek.
response_format json_schema400 from DeepSeek.
stop with 17 or more strings400 from DeepSeek.
functions, function_call400 INVALID_PARAMETER.
Images to deepseek-v4-pro400 MODEL_NO_VISION.
Images outside user messages400 IMAGE_NOT_IN_USER_MESSAGE.
More than 32 images400 TOO_MANY_IMAGES.
Thinking, or a model other than deepseek-v4-flash, on the free tier402 FREE_TIER_THINKING or FREE_TIER_MODEL.

The reply

A reply that isn't streamed is one JSON object: DeepSeek's reply, passed on as DeepSeek sends it.

Name
Type
Description
id
string
Identifies the reply. It is a UUID, not chatcmpl-....
object
string
chat.completion
created
integer
When the reply was made, in Unix seconds.
model
string
DeepSeek's name for the model that answered: deepseek-flash for deepseek-v4-flash. You are billed for the model you asked for.
choices
array
One item, since n is always 1.
choices[].index
integer
0
choices[].message.role
string
assistant
choices[].message.content
string
The answer. In a reply that calls a tool it can be text or empty. It is empty when reasoning used up max_tokens.
choices[].message.reasoning_content
string
The model's reasoning, on requests with thinking on. See Reasoning.
choices[].message.tool_calls
array
The functions the model wants to call. Each has index, id, type ("function"), function.name and function.arguments, the arguments as a JSON string.
choices[].logprobs
object or null
null unless the request set "logprobs": true. Then content lists each reply token with token, logprob, bytes and top_logprobs.
choices[].finish_reason
string
Why the reply ended: stop (it was complete, or reached a stop string), length (it reached max_tokens) or tool_calls (it calls a tool). Treat any other value as an incomplete reply.
usage.prompt_tokens
integer
Tokens in the prompt, tool definitions and images included.
usage.completion_tokens
integer
Tokens in the reply, reasoning included.
usage.total_tokens
integer
The two added up.
usage.prompt_cache_hit_tokens
integer
Prompt tokens read from DeepSeek's cache, billed at the cached-input price. See Prompt caching and cost.
usage.prompt_cache_miss_tokens
integer
The other prompt tokens, billed at the input price.
usage.prompt_tokens_details.cached_tokens
integer
The same count as prompt_cache_hit_tokens, where OpenAI's format puts it.
usage.completion_tokens_details.reasoning_tokens
integer
With thinking: how many of the completion tokens are reasoning.
system_fingerprint
string
Passed on from DeepSeek.

The reply doesn't include its price. The Usage page lists the cost of each request.

Log probabilities

With "logprobs": true and "top_logprobs": 2, a one-token reply carried this:

Part of the response
"logprobs": {
  "content": [
    {
      "token": "Yes",
      "logprob": -0.000051020274,
      "bytes": [89, 101, 115],
      "top_logprobs": [
        { "token": "Yes", "logprob": -0.000051020274, "bytes": [89, 101, 115] },
        { "token": "yes", "logprob": -9.89327, "bytes": [121, 101, 115] }
      ]
    }
  ]
}

Tool calls

With a get_weather function in tools, the question “What is the weather in Tripoli?” got this choice. Your code runs the function and sends the result back in a tool message: see Tool calling.

choices[0]
{
  "index": 0,
  "message": {
    "role": "assistant",
    "content": "I'll check the current weather in Tripoli for you.",
    "tool_calls": [
      {
        "index": 0,
        "id": "call_00_l0FD1WqIMX4lwllRfZaa5995",
        "type": "function",
        "function": {
          "name": "get_weather",
          "arguments": "{\"city\": \"Tripoli\"}"
        }
      }
    ]
  },
  "logprobs": null,
  "finish_reason": "tool_calls"
}

Streamed replies

With "stream": true, the reply comes as Server-Sent Events: lines that start with data: , each holding one JSON chunk and followed by a blank line. The last line is data: [DONE]. This is a whole stream:

Stream
data: {"id":"2fdb0a67-522b-44f2-b79f-2101790a929a","object":"chat.completion.chunk","created":1790380870,"model":"deepseek-flash","system_fingerprint":"aeb56401ca74e127821c4f9126dcb669","choices":[{"index":0,"delta":{"role":"assistant","content":""},"logprobs":null,"finish_reason":null}],"usage":null}

data: {"id":"2fdb0a67-522b-44f2-b79f-2101790a929a","object":"chat.completion.chunk","created":1790380870,"model":"deepseek-flash","system_fingerprint":"aeb56401ca74e127821c4f9126dcb669","choices":[{"index":0,"delta":{"content":"Good"},"logprobs":null,"finish_reason":null}],"usage":null}

data: {"id":"2fdb0a67-522b-44f2-b79f-2101790a929a","object":"chat.completion.chunk","created":1790380870,"model":"deepseek-flash","system_fingerprint":"aeb56401ca74e127821c4f9126dcb669","choices":[{"index":0,"delta":{"content":" morning"},"logprobs":null,"finish_reason":null}],"usage":null}

data: {"id":"2fdb0a67-522b-44f2-b79f-2101790a929a","object":"chat.completion.chunk","created":1790380870,"model":"deepseek-flash","system_fingerprint":"aeb56401ca74e127821c4f9126dcb669","choices":[{"index":0,"delta":{"content":""},"logprobs":null,"finish_reason":"stop"}],"usage":{"prompt_tokens":12,"completion_tokens":2,"total_tokens":14,"prompt_tokens_details":{"cached_tokens":0},"prompt_cache_hit_tokens":0,"prompt_cache_miss_tokens":12}}

data: [DONE]
Name
Type
Description
object
string
chat.completion.chunk
choices[].delta.role
string
assistant, in the first chunk.
choices[].delta.content
string
The next piece of the answer.
choices[].delta.reasoning_content
string
The next piece of the reasoning, when the model thinks.
choices[].delta.tool_calls
array
Pieces of tool calls. The first piece of a call has index, id, type and function.name; later pieces add to function.arguments. Join the pieces that share an index.
choices[].finish_reason
string or null
null until the last chunk.
usage
object or null
null in every chunk but the last, which carries the usage of the whole request.

id, created, model and system_fingerprint repeat in every chunk. Code samples, and what happens when a stream stops early, are in Streaming.

Response headers

HeaderWhat it says
Content-Typeapplication/json, or text/event-stream for a stream.
x-request-idOn every reply from the API: the id you sent, or one the API made. Quote it when you contact support. Replies from the servers in front of the API, such as a 413, don't have it: see Errors.
X-RateLimit-LimitRequests your account may send per minute.
X-RateLimit-RemainingRequests left in the current minute.
X-RateLimit-ResetWhen the current minute ends, in Unix seconds.
Retry-AfterOn a 429 for the per-minute limit: seconds to wait.

The X-RateLimit headers come on every reply from this endpoint once the key is accepted. See Rate limits and other limits.

Errors

The gateway's own checks, made before it calls DeepSeek:

Status and codeCause
400 INVALID_JSONThe body isn't a JSON object.
400 INVALID_PARAMETERmodel, max_tokens, max_completion_tokens or stream has a wrong value, or the request sends functions or function_call. param names the field.
400 MODEL_NOT_FOUNDThe model id isn't one this API serves.
400 MODEL_NO_VISION, IMAGE_NOT_IN_USER_MESSAGE, TOO_MANY_IMAGESAn image the API won't send on.
401 MISSING_API_KEY, INVALID_API_KEYNo key, or a key that doesn't work.
402 FREE_TIER_...Your balance or the free allowance doesn't cover the request.
429 rate_limit_error, concurrency_limitToo many requests in the minute, or at the same time.

A value DeepSeek refuses comes back with DeepSeek's own error. Every error, with what to do, is in Errors.