Skip to content
أكاكوسAcacus
تواصل مع المبيعاتContact sales
Docs menu

Get started

Libraries and SDKs

The API follows OpenAI's Chat Completions format, so OpenAI's own libraries, and tools built on them, work with it. Point them at the Acacus base URL and give them your Acacus key.

Whatever you use, set these three things:

  • Base URL: https://acacus.ly/v1
  • API key: your Acacus key, which starts with sk-shfr-. See Authentication and API keys.
  • Model: deepseek-v4-flash or deepseek-v4-pro. See Models.

The samples on this page were run as written against https://acacus.ly/v1 with openai 3.19.2 (Python), openai 7.23.0 (Node.js), langchain-openai 1.6.6 and llama-index-llms-openai-like 0.8.0. Each prints its answer, or the output shown under it.

Python

Use OpenAI's openai package. It needs Python 3.10 or later.

Python
# pip install openai
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["ACACUS_API_KEY"],
    base_url="https://acacus.ly/v1",
)

response = client.chat.completions.create(
    model="deepseek-v4-flash",
    messages=[
        {"role": "system", "content": "Answer in one sentence."},
        {"role": "user", "content": "What is the capital of Libya?"},
    ],
    # DeepSeek's fields are not OpenAI parameters: send them in extra_body.
    extra_body={"thinking": {"type": "disabled"}},
)
print(response.choices[0].message.content)
  • DeepSeek's own fields, such as thinking, go in extra_body. The library adds them to the JSON body as they are.
  • reasoning_effort is also an OpenAI parameter, so pass it as a normal argument. See Reasoning.
  • stream=True and the client.chat.completions.stream() helper both work. See Streaming.

Environment variables

When you don't pass api_key and base_url, the library reads OPENAI_API_KEY and OPENAI_BASE_URL. This helps with tools that create their own OpenAI client:

Terminal
export OPENAI_API_KEY="$ACACUS_API_KEY"
export OPENAI_BASE_URL="https://acacus.ly/v1"
Python
from openai import OpenAI

client = OpenAI()  # reads OPENAI_API_KEY and OPENAI_BASE_URL

response = client.chat.completions.create(
    model="deepseek-v4-flash",
    messages=[{"role": "user", "content": "Say hello."}],
    extra_body={"thinking": {"type": "disabled"}},
)
print(response.choices[0].message.content)

Errors

An error reply raises openai.APIStatusError, or one of its subclasses: BadRequestError (400), AuthenticationError (401), PermissionDeniedError (403), NotFoundError (404), UnprocessableEntityError (422), RateLimitError (429) and InternalServerError (5xx). A 402 is a plain APIStatusError. This request names a model that doesn't exist:

Python
import os
import openai
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["ACACUS_API_KEY"],
    base_url="https://acacus.ly/v1",
)

try:
    client.chat.completions.create(
        model="deepseek-chat",  # DeepSeek's own name: not a model id here
        messages=[{"role": "user", "content": "Say hello."}],
    )
except openai.APIStatusError as e:
    print(e.status_code, e.code, e.param)
    print(e.body["message"])  # e.body is the "error" object of the reply
    print(e.request_id)  # quote it if you write to support
Output
400 MODEL_NOT_FOUND model
Unknown model: deepseek-chat. Model ids are exact. GET /v1/models, sent with this key, lists the models it can use: deepseek-v4-flash, deepseek-v4-pro and any others your account has.
2931faea-c596-4c25-8310-5fc23b7d1fbb

e.body is the error object of the reply, so the details are there: on a 402, e.body["balance"] is your balance. The codes are listed in Errors.

Node.js

Use OpenAI's openai package. The current version needs Node.js 22 or later. The samples are ES modules: save them as .mjs files.

Node.js
// npm install openai
// Save as client.mjs, then run: node client.mjs
import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.ACACUS_API_KEY,
  baseURL: "https://acacus.ly/v1",
});

const response = await client.chat.completions.create({
  model: "deepseek-v4-flash",
  messages: [
    { role: "system", content: "Answer in one sentence." },
    { role: "user", content: "What is the capital of Libya?" },
  ],
  // DeepSeek's fields are not OpenAI parameters: add them to the object.
  thinking: { type: "disabled" },
});
console.log(response.choices[0].message.content);
  • DeepSeek's own fields, such as thinking, go straight in the request object. The library sends them as they are.
  • stream: true and the client.chat.completions.stream() helper both work. See Streaming.

TypeScript

TypeScript refuses fields that are not in the SDK's types (error TS2353 for thinking). Add them to the request type:

TypeScript
// npm install openai
// Save as client.mts, then run: node client.mts
import OpenAI from "openai";

// DeepSeek's thinking field is not in the SDK's types: add it to the request type.
type AcacusChatParams = OpenAI.ChatCompletionCreateParamsNonStreaming & {
  thinking?: { type: "enabled" | "adaptive" | "disabled" };
};

const client = new OpenAI({
  apiKey: process.env.ACACUS_API_KEY,
  baseURL: "https://acacus.ly/v1",
});

const params: AcacusChatParams = {
  model: "deepseek-v4-flash",
  messages: [{ role: "user", content: "What is the capital of Libya?" }],
  thinking: { type: "disabled" },
};

const response = await client.chat.completions.create(params);
console.log(response.choices[0].message.content);

We checked this file with tsc --strict and ran it with Node.js 22.23, which runs TypeScript files directly. With an older Node.js, compile it with tsc first, or run it with a tool such as tsx.

Errors

An error reply throws OpenAI.APIError, or a subclass with the same names as in Python. A 402 is a plain APIError.

Node.js
import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.ACACUS_API_KEY,
  baseURL: "https://acacus.ly/v1",
});

try {
  await client.chat.completions.create({
    model: "deepseek-chat", // DeepSeek's own name: not a model id here
    messages: [{ role: "user", content: "Say hello." }],
  });
} catch (err) {
  if (!(err instanceof OpenAI.APIError)) throw err;
  console.log(err.status, err.code, err.param);
  console.log(err.error.message); // err.error is the "error" object of the reply
  console.log(err.requestID); // quote it if you write to support
}
Output
400 MODEL_NOT_FOUND model
Unknown model: deepseek-chat. Model ids are exact. GET /v1/models, sent with this key, lists the models it can use: deepseek-v4-flash, deepseek-v4-pro and any others your account has.
59acaf58-3d6b-4aff-a70c-8ed6f252eb85

err.error is the error object of the reply: on a 402, err.error.balance is your balance.

Retries and timeouts

  • Both OpenAI libraries retry a request that gets status 408, 409, 429 or 5xx, times out or can't connect: 2 more tries by default, after a short wait, or after the Retry-After time when the reply has one.
  • Each retry is a new request. It counts toward your rate limit, and it is billed like any other request. A 502 UPSTREAM_INTERRUPTED reply is billed for its input and an estimate of its reply (20 tokens per second of the time it ran, up to max_tokens), so its retry is billed again. To decide for yourself, set max_retries=0 (Python) or maxRetries: 0 (Node.js).
  • Both libraries wait up to 10 minutes for a reply, but nginx, in front of the API, closes a connection when no data comes back for 120 seconds and answers 504. So stream long replies. See Streaming and Rate limits and other limits.

LangChain

Use ChatOpenAI from the langchain-openai package, with the Acacus base URL:

Python
# pip install langchain-openai
import os
from langchain_openai import ChatOpenAI
from pydantic import BaseModel

llm = ChatOpenAI(
    model="deepseek-v4-flash",
    api_key=os.environ["ACACUS_API_KEY"],
    base_url="https://acacus.ly/v1",
    extra_body={"thinking": {"type": "disabled"}},  # DeepSeek's field
)

print(llm.invoke("What is the capital of Libya? Answer in one sentence.").content)


class City(BaseModel):
    name: str
    country: str


# The default method (json_schema) is refused: use function calling.
structured = llm.with_structured_output(City, method="function_calling")
print(structured.invoke("Tripoli is the capital of which country?"))
Output
The capital of Libya is Tripoli.
name='Tripoli' country='Libya'
  • invoke, stream, bind_tools and with_structured_output with method="function_calling" or method="json_mode" work.
  • with_structured_output without a method sends a json_schema response format, which the API refuses with 400 (“This response_format type is unavailable now”).
  • DeepSeek's thinking field goes in extra_body. reasoning_effort= works as a normal argument.
  • Don't set reasoning= or use_responses_api=True. They switch ChatOpenAI to OpenAI's Responses API, which this API doesn't have (404).

LlamaIndex

Use OpenAILike from the llama-index-llms-openai-like package. LlamaIndex's OpenAI class accepts only OpenAI's model names: with deepseek-v4-flash it raises “Unknown model”.

Python
# pip install llama-index-llms-openai-like
import os
from llama_index.llms.openai_like import OpenAILike

llm = OpenAILike(
    model="deepseek-v4-flash",
    api_base="https://acacus.ly/v1",
    api_key=os.environ["ACACUS_API_KEY"],
    is_chat_model=True,  # send requests to /v1/chat/completions
    is_function_calling_model=True,  # the models can call tools
    additional_kwargs={"extra_body": {"thinking": {"type": "disabled"}}},  # DeepSeek's field
)

print(llm.complete("What is the capital of Libya? Answer in one sentence."))
Output
The capital of Libya is Tripoli.
  • is_chat_model=True is required. Without it, LlamaIndex calls /v1/completions, which this API doesn't have (404).
  • With is_function_calling_model=True, tools (for example predict_and_call) and structured_predict work.
  • DeepSeek's own fields go in additional_kwargs, under extra_body.
  • By default OpenAILike waits 60 seconds for a reply and retries 3 times, and each retry is a new request. Change that with timeout and max_retries.
  • context_window (3,900 tokens by default) tells LlamaIndex how much text to put in one prompt. You can raise it, but keep each request under 2 MB. On the free tier, LlamaIndex puts the retrieved text in the new user message, which has a smaller limit than the 80,000 tokens of a whole request: about 65,000 tokens at today's settings, and about 25,000 tokens during DeepSeek's peak hours. See the largest new message.

Plain HTTP

Any HTTP client works. Send a JSON body with the Authorization header. The reply is JSON, or Server-Sent Events when you stream (see Streaming).

curl https://acacus.ly/v1/chat/completions \
  -H "Authorization: Bearer $ACACUS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek-v4-flash",
    "messages": [
      {"role": "system", "content": "Answer in one sentence."},
      {"role": "user", "content": "What is the capital of Libya?"}
    ],
    "thinking": {"type": "disabled"}
  }'

403 “error code: 1010”

Cloudflare, in front of the API, refuses the default User-Agent of Python's urllib and of Perl's LWP (libwww-perl) with 403 and the text error code: 1010. Send your own User-Agent header, or use a client that works as it is: curl, requests, httpx, Node.js fetch and the OpenAI libraries.

What works and what doesn't

FeatureWorksNotes
Chat completions, whole or streamedYesSee Chat completions and Streaming.
Stream helpers: chat.completions.stream() in Python and Node.jsYes
Tool calling: tools, tool_choiceYesSee Tool calling.
JSON mode: response_format json_objectYesSee JSON output.
Structured outputs: json_schema, chat.completions.parse(), Pydantic or zod formatsNo (400)Use json_object and check the result, or a tool with a JSON schema.
Images in user messagesdeepseek-v4-flash onlySee Images.
DeepSeek's thinking fieldYesextra_body in Python, in the request object in Node.js. See Reasoning.
models.list(), models.retrieve()YesNo key needed.
n above 1No (400)One reply per request.
Responses API: responses.create()No (404)Use chat completions.
Embeddings, legacy completions, images, audio, files, batch, fine-tuning, assistantsNo (404)The API serves chat completions and models only.

Next steps