API reference
Models
The Acacus API serves two models, both running on DeepSeek V4 models (V4.1 Flash and V4 Pro): deepseek-v4-flash, the default, and deepseek-v4-pro. You call both with your Acacus key and pay for both in Libyan dinars from your API balance.
The models
| Model | Runs on | Reads images | Context | Default reply | Longest reply |
|---|---|---|---|---|---|
deepseek-v4-flash (default) | DeepSeek V4.1 Flash | Yes | 1,000,000 tokens | 8,192 tokens | 32,768 tokens |
deepseek-v4-pro | DeepSeek V4 Pro | No | 1,000,000 tokens | 8,192 tokens | 32,768 tokens |
“Default reply” is where a reply stops when the request sets no max_tokens. The Acacus gateway checks your key, applies the limits and bills your API balance in Libyan dinars. The content of each request is processed by the models' provider, DeepSeek (see the Privacy policy).
deepseek-v4-flashis used when a request has nomodel. It reads images, and it costs less per token thandeepseek-v4-pro. The free daily allowance covers it.deepseek-v4-proreads text only: a request with an image gets400MODEL_NO_VISION. It needs a balance: on the free allowance it gets402FREE_TIER_MODEL.
Your account may have access to additional models; GET /v1/models lists what your key can use.
We don't publish comparisons of the two models' quality or speed. Try your own prompts on both, and compare the answers and what each request cost on the Usage page. Requests to either model can use DeepSeek's thinking fields: see Reasoning.
Prices
| Model | Input | Cached input | Output |
|---|---|---|---|
deepseek-v4-flash | 4.50 LYD / 1M tokens | 0.09 LYD / 1M tokens | 18.00 LYD / 1M tokens |
deepseek-v4-pro | 19.80 LYD / 1M tokens | 0.66 LYD / 1M tokens | 59.40 LYD / 1M tokens |
Libyan dinars (LYD) per 1 million tokens, in effect since 25 September 2026. Output includes reasoning tokens. The same price applies at every hour. Each request is billed to the dirham (0.001 LYD), rounded up, with no minimum.
The free daily allowance covers deepseek-v4-flash only. See Billing and the free tier.
Model names
Send the ids exactly as listed. Any other name gets 400 MODEL_NOT_FOUND, including DeepSeek's own names, such as deepseek-chat, deepseek-reasoner and deepseek-flash, and OpenAI's model names. The error says where the ids your key can use are listed:
{
"error": {
"message": "Unknown model: deepseek-chat. Model ids are exact. GET /v1/models, sent with this key, lists the models it can use: deepseek-v4-flash, deepseek-v4-pro and any others your account has.",
"type": "invalid_request_error",
"code": "MODEL_NOT_FOUND",
"param": "model"
}
}The model field of a reply is DeepSeek's name for the model that answered. For deepseek-v4-flash it is deepseek-flash, DeepSeek's name for DeepSeek V4.1 Flash. You are billed for the id you sent.
List the models
No key is needed. The list holds the models that POST /v1/chat/completions serves.
curl https://acacus.ly/v1/models{
"object": "list",
"data": [
{
"id": "deepseek-v4-flash",
"object": "model",
"created": 1758000000,
"owned_by": "shafra-agents"
},
{
"id": "deepseek-v4-pro",
"object": "model",
"created": 1758000000,
"owned_by": "shafra-agents"
}
]
}list.model.model.shafra-agents.Get one model
No key is needed. It returns the same object as the list, for one id. The SDKs call it with models.retrieve().
curl https://acacus.ly/v1/models/deepseek-v4-flash{
"id": "deepseek-v4-flash",
"object": "model",
"created": 1758000000,
"owned_by": "shafra-agents"
}Any other id gets 404 MODEL_NOT_FOUND, and the SDKs raise NotFoundError. Without a key:
{
"error": {
"message": "Unknown model: deepseek-chat. Without an API key this API lists deepseek-v4-flash, deepseek-v4-pro; send Authorization: Bearer <your key> to see the other models your key may use.",
"type": "invalid_request_error",
"code": "MODEL_NOT_FOUND",
"param": "model"
}
}Context and reply length
- Context: 1,000,000 tokens for each model. This is DeepSeek's figure, and we haven't tested requests near that size. Other limits come first: a request body can be at most 2 MB, and a free request at most 80,000 prompt tokens in all, with a smaller limit on its new message (see the largest new message).
- Reply length: without
max_tokens, a reply stops at 8,192 tokens. You can setmax_tokensup to 32,768 tokens. A larger value is lowered to 32,768 tokens, without an error. Free requests are lowered to 2,048 tokens. An additional model your key may use has its own limits: withoutmax_tokensits reply stops at 16,384 tokens, you can ask up to that model's maximum, and a long prompt leaves a shorter reply, since the prompt and the reply share the context. - Reasoning tokens count inside
max_tokens, and they are billed as output. See Reasoning.
The other limits are in Rate limits and other limits.