Skip to content
أكاكوسAcacus
تواصل مع المبيعاتContact sales
Docs menu

Get started

Quickstart

Send your first request to the Acacus API in a few minutes. You need an Acacus account and an API key from the console. New accounts get a small free daily allowance on deepseek-v4-flash, so you can try the API before you top up. Most code written for OpenAI's Chat Completions needs only the Acacus base URL, an Acacus key and a model id; see Compatibility with OpenAI for the differences.

1. Create an account

Sign up and confirm your email address. If you already use the Acacus chat, sign in with that account instead. The chat is never charged to your API balance.

2. Create an API key

Open API keys in the console and click Create key. The key is shown once, so copy it before you close the window. Keys start with sk-shfr-.

Save the key in an environment variable. The examples in these docs read it from ACACUS_API_KEY:

Terminal
export ACACUS_API_KEY="sk-shfr-..."
Keep the key secret. Anyone who has it can spend your balance. Use it on your server only: not in a web page, not in a mobile app and not in a git repository. If a key leaks, revoke it on the API keys page. It stops working at once.

3. Send a request

The API follows OpenAI's Chat Completions format. Use curl, or the official OpenAI library for Python or Node.js with the base URL set to https://acacus.ly/v1. The earlier address, https://ai.shafra.ly/v1, keeps working too.

curl https://acacus.ly/v1/chat/completions \
  -H "Authorization: Bearer $ACACUS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek-v4-flash",
    "messages": [
      {"role": "user", "content": "What is the capital of Libya? Answer in one sentence."}
    ],
    "thinking": {"type": "disabled"}
  }'

thinking is DeepSeek's switch for its thinking step. A paid request thinks before it answers unless it sends "thinking": {"type": "disabled"} or "reasoning_effort": "none", and you pay for the reasoning as output tokens. Free requests never think. See Reasoning.

4. Read the reply

This is the reply the curl example got:

Response
{
  "id": "682a5a6c-cea1-4f49-ba63-d138d6b39537",
  "object": "chat.completion",
  "created": 1790378330,
  "model": "deepseek-flash",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "The capital of Libya is Tripoli."
      },
      "logprobs": null,
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 16,
    "completion_tokens": 8,
    "total_tokens": 24,
    "prompt_tokens_details": {
      "cached_tokens": 0
    },
    "prompt_cache_hit_tokens": 0,
    "prompt_cache_miss_tokens": 16
  },
  "system_fingerprint": "aeb56401ca74e127821c4f9126dcb669"
}
  • choices[0].message.content is the answer.
  • model is DeepSeek's name for the model that answered. deepseek-flash is DeepSeek V4.1 Flash, which serves deepseek-v4-flash.
  • id identifies this reply.
  • usage counts the tokens. prompt_cache_hit_tokens are prompt tokens read from DeepSeek's cache, which cost less. See Prompt caching and cost.

The reply does not include the price. The Usage page lists each request with its cost. Requests the free allowance paid for are marked “Free allowance”.

5. Top up when you need more

Until you top up, your requests use a small free daily allowance, shared with the Acacus chat's daily free allowance. It covers:

  • Model: deepseek-v4-flash only. A request for deepseek-v4-pro gets error 402 FREE_TIER_MODEL.
  • No thinking. A request that asks for it gets 402 FREE_TIER_THINKING.
  • Replies of up to 2,048 tokens. A larger max_tokens is lowered to 2,048 tokens.
  • Requests of up to 80,000 prompt tokens in all. The new message has a smaller limit: about 65,000 tokens at today's settings, and about 25,000 tokens during DeepSeek's peak hours. See the largest new message.
  • The allowance resets at 00:00 UTC. When it runs out, requests get 402 FREE_TIER_DAILY_LIMIT.

For more, top up from 10 LYD on the Top up page, by bank card (Moamalat), bank transfer or LYPay. With a balance you can use both models, thinking, and replies of up to 32,768 tokens. See Billing and the free tier.

Next steps