Guides
Tool calling
Describe functions your code can run. The model can then answer with a request to call one, with the arguments as JSON. Your code runs the function and sends the result back, and the model uses it in its answer. The model never runs anything itself.
How it works
- Send the messages with a list of
tools. - If the model wants a function, the reply has
finish_reasonset totool_calls, andmessage.tool_callsnames the function and its arguments. - Run the function in your code.
- Send the conversation again, with the assistant's message and one
toolmessage per call holding the result. - The model answers, or asks for more calls. Repeat until it answers.
The examples use deepseek-v4-flash, which the free allowance covers, with thinking off.
A complete example
This program gives the model one function, get_weather, runs the calls the model asks for, and prints the final answer. Replace the body of get_weather with your own code.
# pip install openai
import json
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["ACACUS_API_KEY"],
base_url="https://acacus.ly/v1",
)
tools = [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get the current weather in a city.",
"parameters": {
"type": "object",
"properties": {
"city": {"type": "string", "description": "The city, for example Tripoli."},
},
"required": ["city"],
},
},
}
]
def get_weather(city):
# Your own code goes here, for example a call to a weather service.
return {"city": city, "temperature_c": 29, "conditions": "clear"}
functions = {"get_weather": get_weather}
messages = [{"role": "user", "content": "What is the weather in Benghazi right now?"}]
for step in range(5): # a limit, in case the model keeps asking for calls
response = client.chat.completions.create(
model="deepseek-v4-flash",
messages=messages,
tools=tools,
# DeepSeek's switch for its thinking step (not an OpenAI parameter).
extra_body={"thinking": {"type": "disabled"}},
)
message = response.choices[0].message
messages.append(message) # the assistant's turn, sent back as received
if not message.tool_calls:
break
for call in message.tool_calls:
# The model wrote these arguments: check them before you use them.
try:
args = json.loads(call.function.arguments)
result = functions[call.function.name](**args)
except Exception as error:
result = {"error": str(error)}
messages.append({
"role": "tool",
"tool_call_id": call.id,
"content": json.dumps(result),
})
print(f"Called {call.function.name}({call.function.arguments})")
print(message.content)The Python version printed:
Called get_weather({"city": "Benghazi"})
Right now in Benghazi it's **29°C** and **clear**. ☀️Both versions append the assistant's message as they received it, tool_calls included, before the results. The next two sections show the same exchange as plain HTTP.
Define your tools
tools is a list. Each entry describes one function in OpenAI's format:
"function".get_weather. DeepSeek allows letters, digits, underscores and dashes, up to 128 characters. Names must differ.The model writes the arguments itself. They are usually valid JSON that fits your schema, but nothing guarantees it: parse them, check them, and send an error back as the result if they are wrong, as the example does.
Step 1: the model asks for a call
This is the first request of the example, sent with curl:
curl https://acacus.ly/v1/chat/completions \
-H "Authorization: Bearer $ACACUS_API_KEY" \
-H "Content-Type: application/json" \
-d @- <<'EOF'
{
"model": "deepseek-v4-flash",
"messages": [
{"role": "user", "content": "What is the weather in Benghazi right now?"}
],
"tools": [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get the current weather in a city.",
"parameters": {
"type": "object",
"properties": {
"city": {"type": "string", "description": "The city, for example Tripoli."}
},
"required": ["city"]
}
}
}
],
"thinking": {"type": "disabled"}
}
EOFThe reply:
{
"id": "ae6933d4-93c1-415d-a963-4b7737ab4eb4",
"object": "chat.completion",
"created": 1790380858,
"model": "deepseek-flash",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "I'll check the current weather in Benghazi for you.",
"tool_calls": [
{
"index": 0,
"id": "call_00_IbkYSTYQ0RLaWGnyXWp13952",
"type": "function",
"function": {
"name": "get_weather",
"arguments": "{\"city\": \"Benghazi\"}"
}
}
]
},
"logprobs": null,
"finish_reason": "tool_calls"
}
],
"usage": {
"prompt_tokens": 289,
"completion_tokens": 52,
"total_tokens": 341,
"prompt_tokens_details": {
"cached_tokens": 128
},
"prompt_cache_hit_tokens": 128,
"prompt_cache_miss_tokens": 161
},
"system_fingerprint": "aeb56401ca74e127821c4f9126dcb669"
}finish_reasonistool_calls: the model is waiting for results.message.contentcan hold a few words for the user, or be empty.- Each entry of
message.tool_callshas anid, the function'snameand itsarguments. The arguments are a string of JSON: parse it.
Step 2: send the results back
Send the whole conversation again, with the same tools, and add two things at the end: the assistant's message as you received it, tool_calls included, then one message with the role tool for each call. Its tool_call_id is the call's id, and its content is the result as a string. JSON text is fine.
curl https://acacus.ly/v1/chat/completions \
-H "Authorization: Bearer $ACACUS_API_KEY" \
-H "Content-Type: application/json" \
-d @- <<'EOF'
{
"model": "deepseek-v4-flash",
"messages": [
{"role": "user", "content": "What is the weather in Benghazi right now?"},
{
"role": "assistant",
"content": "I'll check the current weather in Benghazi for you.",
"tool_calls": [
{
"id": "call_00_IbkYSTYQ0RLaWGnyXWp13952",
"type": "function",
"function": {"name": "get_weather", "arguments": "{\"city\": \"Benghazi\"}"}
}
]
},
{
"role": "tool",
"tool_call_id": "call_00_IbkYSTYQ0RLaWGnyXWp13952",
"content": "{\"city\": \"Benghazi\", \"temperature_c\": 29, \"conditions\": \"clear\"}"
}
],
"tools": [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get the current weather in a city.",
"parameters": {
"type": "object",
"properties": {
"city": {"type": "string", "description": "The city, for example Tripoli."}
},
"required": ["city"]
}
}
}
],
"thinking": {"type": "disabled"}
}
EOFThe model answers with the result:
{
"id": "f1d53d24-3fa7-47cf-a722-d7946eefa413",
"object": "chat.completion",
"created": 1790380874,
"model": "deepseek-flash",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "Right now in Benghazi, the weather is **clear** with a temperature of **29°C** (about 84°F). Warm and sunny — a good day to be outdoors!"
},
"logprobs": null,
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 374,
"completion_tokens": 40,
"total_tokens": 414,
"prompt_tokens_details": {
"cached_tokens": 128
},
"prompt_cache_hit_tokens": 128,
"prompt_cache_miss_tokens": 246
},
"system_fingerprint": "aeb56401ca74e127821c4f9126dcb669"
}The API keeps no conversation between requests, so an id only has to match between the assistant message and its tool message in the same request.
Choose when tools are used
tool_choice tells the model whether it may, must or must not call a tool. All four forms were tested:
| tool_choice | What the model does |
|---|---|
"auto" | Decides for itself whether to call a tool or answer. This is the default when you send tools. |
"none" | Answers in text and calls no tool. |
"required" | Calls at least one tool. |
{"type": "function", "function": {"name": "get_weather"}} | Calls that function. |
A forced call happens even when the messages give the model nothing to fill the arguments with. In a test, the message “Hello” with get_weather forced gave {"city": "San Francisco"}. With thinking on, DeepSeek refuses "required" and a named function with a 400 error: turn thinking off to use them.
Several calls at once
The model can ask for more than one call in the same reply. Run each one and send one tool message per call, each with its own tool_call_id. The complete example already does this. Asked about Tripoli and Benghazi, the Python version got both calls in one reply and printed:
Called get_weather({"city": "Tripoli"})
Called get_weather({"city": "Benghazi"})
Here's the current weather in both cities:
| City | Temperature | Conditions |
|------|-------------|------------|
| Tripoli | 29 °C | Clear |
| Benghazi | 29 °C | Clear |
Both Tripoli and Benghazi are currently clear and mild, sitting at 29 °C each — nice, warm conditions in both cities right now.Streaming tool calls
With "stream": true, a call arrives in pieces in delta.tool_calls. The first piece of each call has its index, id and function.name, and the next pieces add text to function.arguments. These are the delta objects of a stream of the step 1 request, from the first piece of the call (the line ... stands for 7 pieces left out):
{"tool_calls":[{"index":0,"id":"call_00_wjYoWEFhczYxOMAjZVAN8483","type":"function","function":{"name":"get_weather","arguments":""}}]}
{"tool_calls":[{"index":0,"function":{"arguments":"{"}}]}
{"tool_calls":[{"index":0,"function":{"arguments":"\""}}]}
{"tool_calls":[{"index":0,"function":{"arguments":"city"}}]}
...
{"tool_calls":[{"index":0,"function":{"arguments":"}"}}]}
{"content":""}The last chunk has finish_reason set to tool_calls. Join the pieces that share an index:
# pip install openai
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["ACACUS_API_KEY"],
base_url="https://acacus.ly/v1",
)
tools = [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get the current weather in a city.",
"parameters": {
"type": "object",
"properties": {
"city": {"type": "string", "description": "The city, for example Tripoli."},
},
"required": ["city"],
},
},
}
]
stream = client.chat.completions.create(
model="deepseek-v4-flash",
messages=[{"role": "user", "content": "What is the weather in Benghazi right now?"}],
tools=tools,
stream=True,
extra_body={"thinking": {"type": "disabled"}},
)
calls = {} # the calls, joined from their pieces by index
for chunk in stream:
if not chunk.choices:
continue
delta = chunk.choices[0].delta
if delta.content:
print(delta.content, end="", flush=True)
for piece in delta.tool_calls or []:
call = calls.setdefault(piece.index, {"id": "", "name": "", "arguments": ""})
call["id"] = piece.id or call["id"]
if piece.function:
call["name"] += piece.function.name or ""
call["arguments"] += piece.function.arguments or ""
if chunk.choices[0].finish_reason:
print(f"\nfinish_reason: {chunk.choices[0].finish_reason}")
for call in calls.values():
print(call)The Python version printed:
I'll check the current weather in Benghazi for you.
finish_reason: tool_calls
{'id': 'call_00_af0ml5Jj9oFbxBRshrhU2864', 'name': 'get_weather', 'arguments': '{"city": "Benghazi"}'}See Streaming for the rest of the stream format.
Tool calls with thinking on
On paid requests the model can think before it calls a tool. Its message then also has reasoning_content. DeepSeek's rule: in a request that has tools, every earlier assistant message must carry its reasoning_content, including turns without a tool call, or DeepSeek answers 400.
The simplest way to follow it is the one the complete example uses: append each assistant message exactly as you received it. Python's messages.append(message) and Node's messages.push(message) keep reasoning_content. If you store conversations yourself, store it too. See Reasoning.
What it costs
Tool definitions are sent with every request and count as prompt tokens. In step 1, the one-line question and one tool came to 289 prompt tokens. Results you send back are prompt tokens too.
The tool definitions sit near the start of the prompt, before the conversation, so a request that starts the same way as an earlier one reads that part from DeepSeek's cache, at a lower price. In both steps above, 128 prompt tokens came from the cache, because requests that started the same way had been sent a few minutes before. See Prompt caching and cost.
Not supported
- The older
functionsandfunction_callfields: a request with them gets400INVALID_PARAMETER. Usetoolsandtool_choice. - Schema enforcement with
strict, as explained above. - Tool calls the model didn't make. DeepSeek's documentation says its Chat Completions API doesn't support inserting them into a conversation.