Get started
Use with coding tools
Every tool below asks for the same three things: a base URL, a key and a model name. For a tool that lets you type a custom OpenAI-compatible endpoint, the Acacus API takes the same three.
The three settings
| Setting | Value |
|---|---|
| Base URL (also called API base, endpoint or host) | https://acacus.ly/v1. Include /v1. Leave off the endpoint name: the tool adds /chat/completions itself. https://acacus.ly alone is the website, not the API. |
| API key | A key from API keys, starting with sk-. See Authentication. |
| Model name | deepseek-v4-flash or deepseek-v4-pro, or an open model's id from Models. The id must match exactly; the tool does not look it up. If the tool offers a model list, it reads it from GET /v1/models, which lists the models your key can use. |
To check the three at once from a terminal, ask for the model list. A list of models means the URL and the key are right:
curl https://acacus.ly/v1/models -H "Authorization: Bearer $ACACUS_API_KEY"Cursor
In Cursor's settings, open Models. Under the OpenAI API key, paste your Acacus key and turn on the override of the OpenAI base URL, then enter https://acacus.ly/v1. Add a custom model named exactly deepseek-v4-flash (or another id) and select it in the chat. Cursor sends the request from its own servers, not from your computer, so the API must be reachable from the internet, which it is. Features that run on Cursor's own models, such as its tab completion, do not use your key.
Cline
In Cline's settings, set the API Provider to OpenAI Compatible. Enter https://acacus.ly/v1 as the Base URL, your key as the API Key, and deepseek-v4-flash as the Model ID. Under the model configuration, set the context window to 1,000,000 and the maximum output to 16,384. Turn on image support only for deepseek-v4-flash, the one model that reads images.
Kilo Code and Roo Code
Both are close relatives of Cline and use the same fields. In the provider settings choose OpenAI Compatible, then enter the base URL https://acacus.ly/v1, your key and the model ID. Set the context window and the maximum output as for Cline.
Continue
Continue reads its models from config.yaml (open it from the gear icon in the Continue panel). Use the openai provider and set apiBase:
# ~/.continue/config.yaml
name: Acacus
version: 1.0.0
schema: v1
models:
- name: Acacus Flash
provider: openai
model: deepseek-v4-flash
apiBase: https://acacus.ly/v1
apiKey: <your Acacus key>
roles:
- chat
- edit
- applyContinue's tab autocomplete and its codebase indexing need an endpoint for completions and for embeddings. The Acacus API has neither for API keys (see Chat completions), so use Acacus for chat, edit and apply only.
Aider
Aider reads the standard OpenAI environment variables. The openai/ before the model name tells Aider to use the OpenAI-style endpoint:
export OPENAI_API_BASE=https://acacus.ly/v1
export OPENAI_API_KEY="$ACACUS_API_KEY"
aider --model openai/deepseek-v4-flashAider may warn that it does not know the model's context size or price. The warning is harmless.
OpenAI libraries and LangChain
Point the library's client at the base URL and use your key. The same code works in Python and JavaScript. Send thinking in extra_body (Python) or as a field (JavaScript) when you want the reasoning step off. More on the libraries is in Libraries and SDKs.
# pip install openai
import os
from openai import OpenAI
client = OpenAI(api_key=os.environ["ACACUS_API_KEY"], base_url="https://acacus.ly/v1")
reply = client.chat.completions.create(
model="deepseek-v4-flash",
messages=[{"role": "user", "content": "Write a Python function that reverses a string."}],
# `thinking` is not an OpenAI parameter: the library sends it through extra_body.
extra_body={"thinking": {"type": "disabled"}},
)
print(reply.choices[0].message.content)LangChain's ChatOpenAI takes the same three settings:
# pip install langchain-openai
import os
from langchain_openai import ChatOpenAI
llm = ChatOpenAI(
model="deepseek-v4-flash",
api_key=os.environ["ACACUS_API_KEY"],
base_url="https://acacus.ly/v1",
extra_body={"thinking": {"type": "disabled"}},
)
print(llm.invoke("Write a Python function that reverses a string.").content)What to expect from coding agents
- Limits. An API key can send 120 requests a minute, with 8 requests running at once. Agent tools that start several calls together stay inside that. The limits are separate from the Acacus web chat. See Rate limits.
- Reply length. A tool that asks for a very large
max_tokensgets at most 16,384 tokens on the DeepSeek models. The request does not fail; the number is lowered. - Long replies. Coding tools stream by default, which is what a long answer needs. A request that does not stream and runs past 110 seconds is cancelled with a
504and is not charged. Tools that make one-off calls without streaming, such as a title or a summary, should stay well inside that. - Stop. The Stop button of these tools ends the stream. The reply stops being generated and you are charged for the text you received.
- Thinking. The DeepSeek models think before they answer unless the request turns it off, and the thinking is charged as output. Most tools cannot send the
thinkingfield. In a long tool loop, the model host once refused a later turn that did not send the earlier reasoning back, and that has not been re-checked since the host changed: see Tool calling. For the open models the API adds no thinking setting of its own. - Empty balance. When the balance cannot cover a request, the API falls back to the daily free allowance: short replies, no thinking,
deepseek-v4-flashonly. Every reply carries the headerx-acacus-billing(balanceorfree-tier), andGET /v1/balancereturns the balance. See Billing.
Tools that cannot use the API
A tool can use an Acacus key only if its settings let you enter your own OpenAI-compatible base URL. A tool that works only with its own built-in models cannot send a request to Acacus, whatever key you have. Google Antigravity is one: it has no field for a custom OpenAI-compatible endpoint, so an Acacus key cannot be entered in its agent. The same goes for any tool with no base URL setting.
Two ways around it:
- Use an extension from this page, such as Cline, Kilo Code or Continue, if the editor can install it. Acacus has not tested this inside Antigravity or other editors derived from VS Code.
- Use Acacus Code, our own editor extension. It signs in with your Acacus account and runs on your plan, not on API keys or the API balance.