Get started
Libraries and SDKs
The API follows OpenAI's Chat Completions format, so OpenAI's own libraries, and tools built on them, work with it. Point them at the Acacus base URL and give them your Acacus key.
Whatever you use, set these three things:
- Base URL:
https://acacus.ly/v1 - API key: your Acacus key, which starts with
sk-shfr-. See Authentication and API keys. - Model:
deepseek-v4-flashordeepseek-v4-pro. See Models.
The samples on this page were run as written against https://acacus.ly/v1 with openai 3.19.2 (Python), openai 7.23.0 (Node.js), langchain-openai 1.6.6 and llama-index-llms-openai-like 0.8.0. Each prints its answer, or the output shown under it.
Python
Use OpenAI's openai package. It needs Python 3.10 or later.
# pip install openai
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["ACACUS_API_KEY"],
base_url="https://acacus.ly/v1",
)
response = client.chat.completions.create(
model="deepseek-v4-flash",
messages=[
{"role": "system", "content": "Answer in one sentence."},
{"role": "user", "content": "What is the capital of Libya?"},
],
# DeepSeek's fields are not OpenAI parameters: send them in extra_body.
extra_body={"thinking": {"type": "disabled"}},
)
print(response.choices[0].message.content)- DeepSeek's own fields, such as
thinking, go inextra_body. The library adds them to the JSON body as they are. reasoning_effortis also an OpenAI parameter, so pass it as a normal argument. See Reasoning.stream=Trueand theclient.chat.completions.stream()helper both work. See Streaming.
Environment variables
When you don't pass api_key and base_url, the library reads OPENAI_API_KEY and OPENAI_BASE_URL. This helps with tools that create their own OpenAI client:
export OPENAI_API_KEY="$ACACUS_API_KEY"
export OPENAI_BASE_URL="https://acacus.ly/v1"from openai import OpenAI
client = OpenAI() # reads OPENAI_API_KEY and OPENAI_BASE_URL
response = client.chat.completions.create(
model="deepseek-v4-flash",
messages=[{"role": "user", "content": "Say hello."}],
extra_body={"thinking": {"type": "disabled"}},
)
print(response.choices[0].message.content)Errors
An error reply raises openai.APIStatusError, or one of its subclasses: BadRequestError (400), AuthenticationError (401), PermissionDeniedError (403), NotFoundError (404), UnprocessableEntityError (422), RateLimitError (429) and InternalServerError (5xx). A 402 is a plain APIStatusError. This request names a model that doesn't exist:
import os
import openai
from openai import OpenAI
client = OpenAI(
api_key=os.environ["ACACUS_API_KEY"],
base_url="https://acacus.ly/v1",
)
try:
client.chat.completions.create(
model="deepseek-chat", # DeepSeek's own name: not a model id here
messages=[{"role": "user", "content": "Say hello."}],
)
except openai.APIStatusError as e:
print(e.status_code, e.code, e.param)
print(e.body["message"]) # e.body is the "error" object of the reply
print(e.request_id) # quote it if you write to support400 MODEL_NOT_FOUND model
Unknown model: deepseek-chat. Model ids are exact. GET /v1/models, sent with this key, lists the models it can use: deepseek-v4-flash, deepseek-v4-pro and any others your account has.
2931faea-c596-4c25-8310-5fc23b7d1fbbe.body is the error object of the reply, so the details are there: on a 402, e.body["balance"] is your balance. The codes are listed in Errors.
Node.js
Use OpenAI's openai package. The current version needs Node.js 22 or later. The samples are ES modules: save them as .mjs files.
// npm install openai
// Save as client.mjs, then run: node client.mjs
import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.ACACUS_API_KEY,
baseURL: "https://acacus.ly/v1",
});
const response = await client.chat.completions.create({
model: "deepseek-v4-flash",
messages: [
{ role: "system", content: "Answer in one sentence." },
{ role: "user", content: "What is the capital of Libya?" },
],
// DeepSeek's fields are not OpenAI parameters: add them to the object.
thinking: { type: "disabled" },
});
console.log(response.choices[0].message.content);- DeepSeek's own fields, such as
thinking, go straight in the request object. The library sends them as they are. stream: trueand theclient.chat.completions.stream()helper both work. See Streaming.
TypeScript
TypeScript refuses fields that are not in the SDK's types (error TS2353 for thinking). Add them to the request type:
// npm install openai
// Save as client.mts, then run: node client.mts
import OpenAI from "openai";
// DeepSeek's thinking field is not in the SDK's types: add it to the request type.
type AcacusChatParams = OpenAI.ChatCompletionCreateParamsNonStreaming & {
thinking?: { type: "enabled" | "adaptive" | "disabled" };
};
const client = new OpenAI({
apiKey: process.env.ACACUS_API_KEY,
baseURL: "https://acacus.ly/v1",
});
const params: AcacusChatParams = {
model: "deepseek-v4-flash",
messages: [{ role: "user", content: "What is the capital of Libya?" }],
thinking: { type: "disabled" },
};
const response = await client.chat.completions.create(params);
console.log(response.choices[0].message.content);We checked this file with tsc --strict and ran it with Node.js 22.23, which runs TypeScript files directly. With an older Node.js, compile it with tsc first, or run it with a tool such as tsx.
Errors
An error reply throws OpenAI.APIError, or a subclass with the same names as in Python. A 402 is a plain APIError.
import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.ACACUS_API_KEY,
baseURL: "https://acacus.ly/v1",
});
try {
await client.chat.completions.create({
model: "deepseek-chat", // DeepSeek's own name: not a model id here
messages: [{ role: "user", content: "Say hello." }],
});
} catch (err) {
if (!(err instanceof OpenAI.APIError)) throw err;
console.log(err.status, err.code, err.param);
console.log(err.error.message); // err.error is the "error" object of the reply
console.log(err.requestID); // quote it if you write to support
}400 MODEL_NOT_FOUND model
Unknown model: deepseek-chat. Model ids are exact. GET /v1/models, sent with this key, lists the models it can use: deepseek-v4-flash, deepseek-v4-pro and any others your account has.
59acaf58-3d6b-4aff-a70c-8ed6f252eb85err.error is the error object of the reply: on a 402, err.error.balance is your balance.
Retries and timeouts
- Both OpenAI libraries retry a request that gets status
408,409,429or5xx, times out or can't connect: 2 more tries by default, after a short wait, or after theRetry-Aftertime when the reply has one. - Each retry is a new request. It counts toward your rate limit, and it is billed like any other request. A
502UPSTREAM_INTERRUPTEDreply is billed for its input and an estimate of its reply (20 tokens per second of the time it ran, up tomax_tokens), so its retry is billed again. To decide for yourself, setmax_retries=0(Python) ormaxRetries: 0(Node.js). - Both libraries wait up to 10 minutes for a reply, but nginx, in front of the API, closes a connection when no data comes back for 120 seconds and answers
504. So stream long replies. See Streaming and Rate limits and other limits.
LangChain
Use ChatOpenAI from the langchain-openai package, with the Acacus base URL:
# pip install langchain-openai
import os
from langchain_openai import ChatOpenAI
from pydantic import BaseModel
llm = ChatOpenAI(
model="deepseek-v4-flash",
api_key=os.environ["ACACUS_API_KEY"],
base_url="https://acacus.ly/v1",
extra_body={"thinking": {"type": "disabled"}}, # DeepSeek's field
)
print(llm.invoke("What is the capital of Libya? Answer in one sentence.").content)
class City(BaseModel):
name: str
country: str
# The default method (json_schema) is refused: use function calling.
structured = llm.with_structured_output(City, method="function_calling")
print(structured.invoke("Tripoli is the capital of which country?"))The capital of Libya is Tripoli.
name='Tripoli' country='Libya'invoke,stream,bind_toolsandwith_structured_outputwithmethod="function_calling"ormethod="json_mode"work.with_structured_outputwithout a method sends ajson_schemaresponse format, which the API refuses with400(“This response_format type is unavailable now”).- DeepSeek's
thinkingfield goes inextra_body.reasoning_effort=works as a normal argument. - Don't set
reasoning=oruse_responses_api=True. They switchChatOpenAIto OpenAI's Responses API, which this API doesn't have (404).
LlamaIndex
Use OpenAILike from the llama-index-llms-openai-like package. LlamaIndex's OpenAI class accepts only OpenAI's model names: with deepseek-v4-flash it raises “Unknown model”.
# pip install llama-index-llms-openai-like
import os
from llama_index.llms.openai_like import OpenAILike
llm = OpenAILike(
model="deepseek-v4-flash",
api_base="https://acacus.ly/v1",
api_key=os.environ["ACACUS_API_KEY"],
is_chat_model=True, # send requests to /v1/chat/completions
is_function_calling_model=True, # the models can call tools
additional_kwargs={"extra_body": {"thinking": {"type": "disabled"}}}, # DeepSeek's field
)
print(llm.complete("What is the capital of Libya? Answer in one sentence."))The capital of Libya is Tripoli.is_chat_model=Trueis required. Without it, LlamaIndex calls/v1/completions, which this API doesn't have (404).- With
is_function_calling_model=True, tools (for examplepredict_and_call) andstructured_predictwork. - DeepSeek's own fields go in
additional_kwargs, underextra_body. - By default
OpenAILikewaits 60 seconds for a reply and retries 3 times, and each retry is a new request. Change that withtimeoutandmax_retries. context_window(3,900 tokens by default) tells LlamaIndex how much text to put in one prompt. You can raise it, but keep each request under 2 MB. On the free tier, LlamaIndex puts the retrieved text in the new user message, which has a smaller limit than the 80,000 tokens of a whole request: about 65,000 tokens at today's settings, and about 25,000 tokens during DeepSeek's peak hours. See the largest new message.
Plain HTTP
Any HTTP client works. Send a JSON body with the Authorization header. The reply is JSON, or Server-Sent Events when you stream (see Streaming).
curl https://acacus.ly/v1/chat/completions \
-H "Authorization: Bearer $ACACUS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek-v4-flash",
"messages": [
{"role": "system", "content": "Answer in one sentence."},
{"role": "user", "content": "What is the capital of Libya?"}
],
"thinking": {"type": "disabled"}
}'403 “error code: 1010”
Cloudflare, in front of the API, refuses the default User-Agent of Python's urllib and of Perl's LWP (libwww-perl) with 403 and the text error code: 1010. Send your own User-Agent header, or use a client that works as it is: curl, requests, httpx, Node.js fetch and the OpenAI libraries.
What works and what doesn't
| Feature | Works | Notes |
|---|---|---|
| Chat completions, whole or streamed | Yes | See Chat completions and Streaming. |
Stream helpers: chat.completions.stream() in Python and Node.js | Yes | |
Tool calling: tools, tool_choice | Yes | See Tool calling. |
JSON mode: response_format json_object | Yes | See JSON output. |
Structured outputs: json_schema, chat.completions.parse(), Pydantic or zod formats | No (400) | Use json_object and check the result, or a tool with a JSON schema. |
| Images in user messages | deepseek-v4-flash only | See Images. |
DeepSeek's thinking field | Yes | extra_body in Python, in the request object in Node.js. See Reasoning. |
models.list(), models.retrieve() | Yes | No key needed. |
n above 1 | No (400) | One reply per request. |
Responses API: responses.create() | No (404) | Use chat completions. |
| Embeddings, legacy completions, images, audio, files, batch, fine-tuning, assistants | No (404) | The API serves chat completions and models only. |