Get started
The Acacus API
The Acacus API lets your own code send messages to an AI model and get answers back, with one Acacus account and one API key. You pay in Libyan dinars from your prepaid API balance, with no foreign card, and a small free daily allowance lets you try it before you top up. The API follows OpenAI's Chat Completions format, so the OpenAI libraries work with it once you change the base URL.
What the API gives you
- One account for the API and the Acacus chat. Create and revoke keys in the console, on the API keys page.
- Payment in Libyan dinars: top up from 10 LYD with a local bank card (Moamalat), a bank transfer or LYPay. The credit does not expire, and the price is the same at every hour.
- A small free daily allowance on
deepseek-v4-flash, so you can build and test before you pay. - OpenAI's Chat Completions format, with the libraries you already use: streaming, tool calling, JSON output, images (on
deepseek-v4-flash) and reasoning, with DeepSeek's controls for it. - What each request cost, on the Usage page.
- Support from our team in Libya, by email or WhatsApp (+218 94 380 1609).
Base URL
https://acacus.ly/v1The earlier address, https://ai.shafra.ly/v1, is the same API and keeps working, so code that already uses it needs no change.
Send your API key with every request, in the Authorization header: Authorization: Bearer sk-shfr-.... See Authentication and API keys.
| Endpoint | What it does | API key |
|---|---|---|
POST /v1/chat/completions | Send messages and get a reply, whole or streamed. | Required |
GET /v1/models | List the models. | Not needed |
GET /v1/models/{id} | Get one model. | Not needed |
GET /v1/usage | Your tokens today (from 00:00 UTC) and in total. | Required |
Other OpenAI endpoints are not available: embeddings, legacy completions, the Responses API, images, audio, files, batch, fine-tuning and assistants. Their paths answer 404.
Call the API from your server. It sends no CORS headers, so a web page can't call it, and a key inside a web page or a mobile app would be exposed.
Models
The API serves two models, both running on DeepSeek V4 models (V4.1 Flash and V4 Pro). The Acacus gateway checks your key, applies the limits and bills your API balance. The content of each request is processed by the models' provider, DeepSeek, as the Privacy policy says.
| Model | Runs on | Reads images | Context | Longest reply |
|---|---|---|---|---|
deepseek-v4-flash (default) | DeepSeek V4.1 Flash | Yes | 1,000,000 tokens | 32,768 tokens |
deepseek-v4-pro | DeepSeek V4 Pro | No | 1,000,000 tokens | 32,768 tokens |
A request without model uses deepseek-v4-flash. The model field of a reply shows DeepSeek's own name for the model that answered, such as deepseek-flash. See Models.
Your account may have access to additional models; GET /v1/models lists what your key can use.
Compatibility with OpenAI
Use any OpenAI client that lets you set the base URL, with an Acacus key. Most Chat Completions code works once you change the model name. These are the differences to know:
nmust be 1.response_formatacceptsjson_objectbut notjson_schema. See JSON output.stoptakes up to 16 strings.- Without
max_tokens, a reply stops at 8,192 tokens. Amax_tokensabove 32,768 tokens is lowered to 32,768 tokens, without an error. - A streamed reply always carries
usagein its last chunk. See Streaming. - DeepSeek models can think before they answer. A paid request thinks unless it sends
"thinking": {"type": "disabled"}or"reasoning_effort": "none", and the reasoning is billed as output. See Reasoning. - Errors use OpenAI's format with our own codes. Problems with your balance or the free allowance are status
402. See Errors. - Each account can send 30 requests a minute, 3 requests at a time, of up to 2 MB each. See Rate limits and other limits.
How billing works
- You pay per token, in Libyan dinars (LYD), from your prepaid API balance. The Acacus chat is never charged to it.
- Prompt tokens, prompt tokens read from the cache, and output tokens each have a price per 1 million tokens.
- Each request is billed to the dirham (0.001 LYD), rounded up, with no minimum. The price is the same at every hour.
- Before a paid request runs, your balance must cover its largest possible cost: the whole prompt, plus output up to its
max_tokens. - With no balance, requests use a small free daily allowance on
deepseek-v4-flash.
| Model | Input | Cached input | Output |
|---|---|---|---|
deepseek-v4-flash | 4.50 LYD / 1M tokens | 0.09 LYD / 1M tokens | 18.00 LYD / 1M tokens |
deepseek-v4-pro | 19.80 LYD / 1M tokens | 0.66 LYD / 1M tokens | 59.40 LYD / 1M tokens |
Libyan dinars (LYD) per 1 million tokens, in effect since 25 September 2026. Output includes reasoning tokens. The same price applies at every hour. Each request is billed to the dirham (0.001 LYD), rounded up, with no minimum.
Top up from 10 LYD on the Top up page, by bank card (Moamalat), bank transfer or LYPay. The Usage page shows what each request cost. The details are in Billing and the free tier.
Where to go next
Questions about the API: write to [email protected] or on WhatsApp (+218 94 380 1609).