Guides
Billing and the free tier
The Acacus API is billed in Libyan dinars (LYD), per request, from your prepaid API balance. You pay for the tokens each request uses, and the Acacus chat is never charged to the balance. Top up from 10 LYD with a local bank card, a bank transfer or LYPay, no foreign card needed, and the credit does not expire. Until you top up, your requests use a small free daily allowance.
Prices
| Model | Input | Cached input | Output |
|---|---|---|---|
deepseek-v4-flash | 4.50 LYD / 1M tokens | 0.09 LYD / 1M tokens | 18.00 LYD / 1M tokens |
deepseek-v4-pro | 19.80 LYD / 1M tokens | 0.66 LYD / 1M tokens | 59.40 LYD / 1M tokens |
Libyan dinars (LYD) per 1 million tokens, in effect since 25 September 2026. Output includes reasoning tokens. The same price applies at every hour. Each request is billed to the dirham (0.001 LYD), rounded up, with no minimum.
The prices are the same at every hour. The same prices are on the Pricing page.
How a request is charged
- The API reads the token counts in the reply's
usage: prompt tokens read from the cache, the other prompt tokens, and reply tokens. Reply tokens include reasoning. - Cost = (prompt tokens not from the cache × the input price + prompt tokens from the cache × the cached-input price + reply tokens × the output price) ÷ 1,000,000 tokens.
- Each request is billed to the dirham (0.001 LYD), rounded up, with no minimum. The rule in force is shown under the prices.
- The price used is the one in effect at the time. A new price applies within a minute.
Real requests from these docs, priced at today's rates for deepseek-v4-flash:
| Request | Prompt, not cached | Prompt, cached | Reply | Before rounding | Charged |
|---|---|---|---|---|---|
| The Quickstart's question | 16 tokens | 0 tokens | 8 tokens | 0.000216 LYD | 0.001 LYD |
| A 3,021-token prompt, the first time | 3,021 tokens | 0 tokens | 16 tokens | 0.013883 LYD | 0.014 LYD |
| The same prompt again | 205 tokens | 2,816 tokens | 16 tokens | 0.001464 LYD | 0.002 LYD |
There is no minimum charge: a short request costs what its tokens cost, rounded up to the dirham (0.001 LYD). Long prompts that repeat cost less with the cache: see Prompt caching and cost.
The API reply doesn't include the cost. The Usage page shows what each request was charged.
Your balance and max_tokens
Before a paid request goes to DeepSeek, your balance must cover the most it could cost: every prompt token at the input price, counted with DeepSeek's tokenizer plus 10% and 1,024 tokens for each image, plus a reply as long as its max_tokens. Nothing is held: this is a check, and your API balance is charged only after the reply.
For example, a deepseek-v4-flash request with a 1,000-token prompt is checked as 1,100 prompt tokens. At today's prices it needs a balance of 0.153 LYD without max_tokens (the default, 8,192 tokens), and 0.014 LYD with "max_tokens": 500.
- If your balance can't cover it and
max_tokensis above 8,192 tokens, the API lowersmax_tokensto the largest value the balance covers, but not below 8,192 tokens. - If the balance still can't cover it, the API tries the request as a free request. The free limits then apply:
deepseek-v4-flashonly, no thinking, at most 2,048 reply tokens. The Usage page marks it “Free allowance”. - If it doesn't fit the free tier either, you get
402with a free-tier code andrequiredLyd: the smallest balance that would run the request as a paid one.
So set max_tokens to what you need: a lower value needs less balance, and it doesn't change what a reply costs, since you pay for the tokens the reply has.
When you are charged
- After the reply is complete. Nothing is taken from your API balance while a request runs.
- A streamed reply you stop reading is charged in full: the API reads DeepSeek's reply to the end. To limit it, set
max_tokens. - If DeepSeek's reply breaks off, you are charged what it reported, or else an estimate: the prompt, and the part of the reply you were sent. A request that isn't streamed then gets
502UPSTREAM_INTERRUPTED, and is charged the prompt plus an estimate of the reply for the time the request ran, at 20 tokens per second, up tomax_tokens(a request is cut off after 300 seconds). That is because the model writes the whole reply before it sends any of it. If nothing of the reply had arrived but the provider's waiting signal, nothing is charged. - Rarely, a complete reply comes without token counts. It is then charged its largest possible cost, the same amount the balance check used.
Not charged:
- Requests the API refuses before it calls DeepSeek: every error from the gateway's own checks (
400,401,402,403,404,405,429), and503UPSTREAM_DOWN. - Requests DeepSeek refuses (
400,422), and DeepSeek's server errors. - Free requests.
The free allowance
Your requests use a small free allowance when your API balance is empty, or when your balance can't cover a request (see above). Each account has one allowance a day, shared by all its keys and, when the account has no chat plan, by its free messages in the Acacus chat.
| Free requests | |
|---|---|
| Model | deepseek-v4-flash only. Other models get 402 FREE_TIER_MODEL. |
| Thinking | Not included: asking for it gets 402 FREE_TIER_THINKING. A request that sends neither thinking nor reasoning_effort runs with thinking off. |
| Reply length | max_tokens is lowered to at most 2,048 tokens. |
| Request size | Up to 80,000 prompt tokens for the whole request (DeepSeek's count plus 10%, and 1,024 tokens for each image) and 320,000 bytes of JSON, image data not counted. Larger: 402 FREE_TIER_DAILY_LIMIT with reason request_too_large. The new message has a smaller limit: see below. |
| Rate limits | The same as for paid requests: see Rate limits. |
| Resets | Every day at 00:00 UTC. |
- The allowance is an amount of money, not a number of requests: long prompts and long replies use it up faster.
- Before a free request runs, what is left of the allowance must cover an estimate of it: the new part of the prompt at the input price, the earlier part at the cached-input price, and a reply as long as its
max_tokens. The new part is the last user message, or everything from the last assistant message on, such as tool results. If the estimate doesn't fit, the request gets402with reasonallowance, and the same request with a lowermax_tokensor a shorter new part may still run. - Free requests use the allowance twice as fast during DeepSeek's peak hours (01:00 to 04:00 and 06:00 to 10:00 UTC, Monday to Friday). Paid prices are the same at every hour.
- The free tier as a whole also has a daily limit, and it can be paused. Then free requests get
402FREE_TIER_EXHAUSTED. - Free requests cost nothing. The Usage page marks them “Free allowance”.
The largest new message
Because of this check, the new part of a free request has a limit of its own, even at the start of the day. At today's settings, with max_tokens at 2,048 tokens (what a free request gets when it sets none), the new part can have about 65,000 tokens, or 69 images with a short question. During DeepSeek's peak hours it is about 25,000 tokens, or 26 images. A larger new part gets 402 with reason allowance. A lower max_tokens leaves more room.
The messages before the last assistant message count at the lower cached-input price (see the prices above), so a long conversation usually fits, as long as each new part stays under this limit. In a conversation, the new part includes the last reply you send back.
When there isn't enough credit
Every 402 has your balance, and some have a reason and requiredLyd:
| Code | What it means | What to do |
|---|---|---|
FREE_TIER_DAILY_LIMIT | You used today's free allowance. | Top up, or wait until 00:00 UTC. |
FREE_TIER_DAILY_LIMIT, reason allowance | What is left today is too small for this request. | Lower max_tokens, send less input, or top up. |
FREE_TIER_DAILY_LIMIT, reason request_too_large | The request is over the free size limits. | Send less input, or top up. |
FREE_TIER_EXHAUSTED | The free tier as a whole is used up, switched off or paused (reason paused). | Top up, or try again later. |
FREE_TIER_MODEL | A free request for a model other than deepseek-v4-flash. | Use deepseek-v4-flash, or top up. |
FREE_TIER_THINKING | A free request that asks for thinking. | Send "thinking": {"type": "disabled"}, or top up. |
With money in your API balance, a 402 means the balance is too small for this request: requiredLyd says how much it needs. The messages are in Errors. To handle a 402 in code:
# pip install openai
import os
import openai
from openai import OpenAI
client = OpenAI(
api_key=os.environ["ACACUS_API_KEY"],
base_url="https://acacus.ly/v1",
)
try:
response = client.chat.completions.create(
model="deepseek-v4-flash",
messages=[{"role": "user", "content": "Say OK."}],
max_tokens=20,
extra_body={"thinking": {"type": "disabled"}},
)
print(response.choices[0].message.content)
except openai.APIStatusError as e:
if e.status_code != 402:
raise
error = e.body if isinstance(e.body, dict) else {}
print("Refused:", e.code)
if "reason" in error:
print("Reason:", error["reason"])
print(error.get("message"))
print("Balance:", error.get("balance"), "LYD")
if "requiredLyd" in error:
print("A balance of", error["requiredLyd"], "LYD would run it as a paid request.")When the request runs, it prints the reply. The same code with model="deepseek-v4-pro", on an account without a balance, printed:
Refused: FREE_TIER_MODEL
The free tier only includes deepseek-v4-flash. Top up to use deepseek-v4-pro, or switch to deepseek-v4-flash.
Balance: 0 LYDTop up
Top up on the Top up page, from 10 LYD to 1,000 LYD at a time. You choose how to pay on the next page: a local bank card through Moamalat, credited right away, or a bank transfer or LYPay transfer, credited once it arrives. The API balance pays for API requests only: the Acacus chat is never charged to it.
Top-ups are not refundable, and the credit in your API balance does not expire. See the Terms. Past top-ups are on Top-up history.
See what you spend
The Usage page lists each request with its time, model, channel, tokens, the amount charged and its status. Choose API under Channel to see only API requests, and Download results (CSV) to save them.
From code, GET /v1/usage returns your tokens:
curl https://acacus.ly/v1/usage \
-H "Authorization: Bearer $ACACUS_API_KEY"{"todayTokens":19890,"totalTokens":19890}todayTokens: tokens since 00:00 UTC today.totalTokens: all your tokens. Each is prompt tokens, from the cache or not, plus reply tokens.- It counts all your use, the Acacus chat included, and has no amounts in LYD: those are on the Usage page.
- It needs your API key. It doesn't count toward the rate limit.