Skip to content
أكاكوسAcacus
تواصل مع المبيعاتContact sales
Docs menu

Guides

Tool calling

Describe functions your code can run. The model can then answer with a request to call one, with the arguments as JSON. Your code runs the function and sends the result back, and the model uses it in its answer. The model never runs anything itself.

How it works

  1. Send the messages with a list of tools.
  2. If the model wants a function, the reply has finish_reason set to tool_calls, and message.tool_calls names the function and its arguments.
  3. Run the function in your code.
  4. Send the conversation again, with the assistant's message and one tool message per call holding the result.
  5. The model answers, or asks for more calls. Repeat until it answers.

The examples use deepseek-v4-flash, which the free allowance covers, with thinking off.

A complete example

This program gives the model one function, get_weather, runs the calls the model asks for, and prints the final answer. Replace the body of get_weather with your own code.

# pip install openai
import json
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["ACACUS_API_KEY"],
    base_url="https://acacus.ly/v1",
)

tools = [
    {
        "type": "function",
        "function": {
            "name": "get_weather",
            "description": "Get the current weather in a city.",
            "parameters": {
                "type": "object",
                "properties": {
                    "city": {"type": "string", "description": "The city, for example Tripoli."},
                },
                "required": ["city"],
            },
        },
    }
]


def get_weather(city):
    # Your own code goes here, for example a call to a weather service.
    return {"city": city, "temperature_c": 29, "conditions": "clear"}


functions = {"get_weather": get_weather}

messages = [{"role": "user", "content": "What is the weather in Benghazi right now?"}]

for step in range(5):  # a limit, in case the model keeps asking for calls
    response = client.chat.completions.create(
        model="deepseek-v4-flash",
        messages=messages,
        tools=tools,
        # DeepSeek's switch for its thinking step (not an OpenAI parameter).
        extra_body={"thinking": {"type": "disabled"}},
    )
    message = response.choices[0].message
    messages.append(message)  # the assistant's turn, sent back as received
    if not message.tool_calls:
        break
    for call in message.tool_calls:
        # The model wrote these arguments: check them before you use them.
        try:
            args = json.loads(call.function.arguments)
            result = functions[call.function.name](**args)
        except Exception as error:
            result = {"error": str(error)}
        messages.append({
            "role": "tool",
            "tool_call_id": call.id,
            "content": json.dumps(result),
        })
        print(f"Called {call.function.name}({call.function.arguments})")

print(message.content)

The Python version printed:

Output
Called get_weather({"city": "Benghazi"})
Right now in Benghazi it's **29°C** and **clear**. ☀️

Both versions append the assistant's message as they received it, tool_calls included, before the results. The next two sections show the same exchange as plain HTTP.

Define your tools

tools is a list. Each entry describes one function in OpenAI's format:

Name
Type
Description
typeRequired
string
Always "function".
function.nameRequired
string
The name the model uses to call it, such as get_weather. DeepSeek allows letters, digits, underscores and dashes, up to 128 characters. Names must differ.
function.descriptionOptional
string
What the function does. The model reads it to decide when to call the function.
function.parametersOptional
object
The arguments, as a JSON Schema object. Leave it out for a function without arguments.
function.strictOptional
boolean
Accepted, but it doesn't make DeepSeek follow the schema: DeepSeek applies it only on its own beta address, which Acacus doesn't use. Check the arguments in your code.

The model writes the arguments itself. They are usually valid JSON that fits your schema, but nothing guarantees it: parse them, check them, and send an error back as the result if they are wrong, as the example does.

Step 1: the model asks for a call

This is the first request of the example, sent with curl:

curl https://acacus.ly/v1/chat/completions \
  -H "Authorization: Bearer $ACACUS_API_KEY" \
  -H "Content-Type: application/json" \
  -d @- <<'EOF'
{
  "model": "deepseek-v4-flash",
  "messages": [
    {"role": "user", "content": "What is the weather in Benghazi right now?"}
  ],
  "tools": [
    {
      "type": "function",
      "function": {
        "name": "get_weather",
        "description": "Get the current weather in a city.",
        "parameters": {
          "type": "object",
          "properties": {
            "city": {"type": "string", "description": "The city, for example Tripoli."}
          },
          "required": ["city"]
        }
      }
    }
  ],
  "thinking": {"type": "disabled"}
}
EOF

The reply:

Response
{
  "id": "ae6933d4-93c1-415d-a963-4b7737ab4eb4",
  "object": "chat.completion",
  "created": 1790380858,
  "model": "deepseek-flash",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "I'll check the current weather in Benghazi for you.",
        "tool_calls": [
          {
            "index": 0,
            "id": "call_00_IbkYSTYQ0RLaWGnyXWp13952",
            "type": "function",
            "function": {
              "name": "get_weather",
              "arguments": "{\"city\": \"Benghazi\"}"
            }
          }
        ]
      },
      "logprobs": null,
      "finish_reason": "tool_calls"
    }
  ],
  "usage": {
    "prompt_tokens": 289,
    "completion_tokens": 52,
    "total_tokens": 341,
    "prompt_tokens_details": {
      "cached_tokens": 128
    },
    "prompt_cache_hit_tokens": 128,
    "prompt_cache_miss_tokens": 161
  },
  "system_fingerprint": "aeb56401ca74e127821c4f9126dcb669"
}
  • finish_reason is tool_calls: the model is waiting for results.
  • message.content can hold a few words for the user, or be empty.
  • Each entry of message.tool_calls has an id, the function's name and its arguments. The arguments are a string of JSON: parse it.

Step 2: send the results back

Send the whole conversation again, with the same tools, and add two things at the end: the assistant's message as you received it, tool_calls included, then one message with the role tool for each call. Its tool_call_id is the call's id, and its content is the result as a string. JSON text is fine.

curl https://acacus.ly/v1/chat/completions \
  -H "Authorization: Bearer $ACACUS_API_KEY" \
  -H "Content-Type: application/json" \
  -d @- <<'EOF'
{
  "model": "deepseek-v4-flash",
  "messages": [
    {"role": "user", "content": "What is the weather in Benghazi right now?"},
    {
      "role": "assistant",
      "content": "I'll check the current weather in Benghazi for you.",
      "tool_calls": [
        {
          "id": "call_00_IbkYSTYQ0RLaWGnyXWp13952",
          "type": "function",
          "function": {"name": "get_weather", "arguments": "{\"city\": \"Benghazi\"}"}
        }
      ]
    },
    {
      "role": "tool",
      "tool_call_id": "call_00_IbkYSTYQ0RLaWGnyXWp13952",
      "content": "{\"city\": \"Benghazi\", \"temperature_c\": 29, \"conditions\": \"clear\"}"
    }
  ],
  "tools": [
    {
      "type": "function",
      "function": {
        "name": "get_weather",
        "description": "Get the current weather in a city.",
        "parameters": {
          "type": "object",
          "properties": {
            "city": {"type": "string", "description": "The city, for example Tripoli."}
          },
          "required": ["city"]
        }
      }
    }
  ],
  "thinking": {"type": "disabled"}
}
EOF

The model answers with the result:

Response
{
  "id": "f1d53d24-3fa7-47cf-a722-d7946eefa413",
  "object": "chat.completion",
  "created": 1790380874,
  "model": "deepseek-flash",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "Right now in Benghazi, the weather is **clear** with a temperature of **29°C** (about 84°F). Warm and sunny — a good day to be outdoors!"
      },
      "logprobs": null,
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 374,
    "completion_tokens": 40,
    "total_tokens": 414,
    "prompt_tokens_details": {
      "cached_tokens": 128
    },
    "prompt_cache_hit_tokens": 128,
    "prompt_cache_miss_tokens": 246
  },
  "system_fingerprint": "aeb56401ca74e127821c4f9126dcb669"
}

The API keeps no conversation between requests, so an id only has to match between the assistant message and its tool message in the same request.

Choose when tools are used

tool_choice tells the model whether it may, must or must not call a tool. All four forms were tested:

tool_choiceWhat the model does
"auto"Decides for itself whether to call a tool or answer. This is the default when you send tools.
"none"Answers in text and calls no tool.
"required"Calls at least one tool.
{"type": "function", "function": {"name": "get_weather"}}Calls that function.

A forced call happens even when the messages give the model nothing to fill the arguments with. In a test, the message “Hello” with get_weather forced gave {"city": "San Francisco"}. With thinking on, DeepSeek refuses "required" and a named function with a 400 error: turn thinking off to use them.

Several calls at once

The model can ask for more than one call in the same reply. Run each one and send one tool message per call, each with its own tool_call_id. The complete example already does this. Asked about Tripoli and Benghazi, the Python version got both calls in one reply and printed:

Output
Called get_weather({"city": "Tripoli"})
Called get_weather({"city": "Benghazi"})
Here's the current weather in both cities:

| City | Temperature | Conditions |
|------|-------------|------------|
| Tripoli | 29 °C | Clear |
| Benghazi | 29 °C | Clear |

Both Tripoli and Benghazi are currently clear and mild, sitting at 29 °C each — nice, warm conditions in both cities right now.

Streaming tool calls

With "stream": true, a call arrives in pieces in delta.tool_calls. The first piece of each call has its index, id and function.name, and the next pieces add text to function.arguments. These are the delta objects of a stream of the step 1 request, from the first piece of the call (the line ... stands for 7 pieces left out):

delta of each chunk
{"tool_calls":[{"index":0,"id":"call_00_wjYoWEFhczYxOMAjZVAN8483","type":"function","function":{"name":"get_weather","arguments":""}}]}
{"tool_calls":[{"index":0,"function":{"arguments":"{"}}]}
{"tool_calls":[{"index":0,"function":{"arguments":"\""}}]}
{"tool_calls":[{"index":0,"function":{"arguments":"city"}}]}
...
{"tool_calls":[{"index":0,"function":{"arguments":"}"}}]}
{"content":""}

The last chunk has finish_reason set to tool_calls. Join the pieces that share an index:

# pip install openai
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["ACACUS_API_KEY"],
    base_url="https://acacus.ly/v1",
)

tools = [
    {
        "type": "function",
        "function": {
            "name": "get_weather",
            "description": "Get the current weather in a city.",
            "parameters": {
                "type": "object",
                "properties": {
                    "city": {"type": "string", "description": "The city, for example Tripoli."},
                },
                "required": ["city"],
            },
        },
    }
]

stream = client.chat.completions.create(
    model="deepseek-v4-flash",
    messages=[{"role": "user", "content": "What is the weather in Benghazi right now?"}],
    tools=tools,
    stream=True,
    extra_body={"thinking": {"type": "disabled"}},
)

calls = {}  # the calls, joined from their pieces by index
for chunk in stream:
    if not chunk.choices:
        continue
    delta = chunk.choices[0].delta
    if delta.content:
        print(delta.content, end="", flush=True)
    for piece in delta.tool_calls or []:
        call = calls.setdefault(piece.index, {"id": "", "name": "", "arguments": ""})
        call["id"] = piece.id or call["id"]
        if piece.function:
            call["name"] += piece.function.name or ""
            call["arguments"] += piece.function.arguments or ""
    if chunk.choices[0].finish_reason:
        print(f"\nfinish_reason: {chunk.choices[0].finish_reason}")

for call in calls.values():
    print(call)

The Python version printed:

Output
I'll check the current weather in Benghazi for you.
finish_reason: tool_calls
{'id': 'call_00_af0ml5Jj9oFbxBRshrhU2864', 'name': 'get_weather', 'arguments': '{"city": "Benghazi"}'}

See Streaming for the rest of the stream format.

Tool calls with thinking on

On paid requests the model can think before it calls a tool. Its message then also has reasoning_content. DeepSeek's rule: in a request that has tools, every earlier assistant message must carry its reasoning_content, including turns without a tool call, or DeepSeek answers 400.

The simplest way to follow it is the one the complete example uses: append each assistant message exactly as you received it. Python's messages.append(message) and Node's messages.push(message) keep reasoning_content. If you store conversations yourself, store it too. See Reasoning.

What it costs

Tool definitions are sent with every request and count as prompt tokens. In step 1, the one-line question and one tool came to 289 prompt tokens. Results you send back are prompt tokens too.

The tool definitions sit near the start of the prompt, before the conversation, so a request that starts the same way as an earlier one reads that part from DeepSeek's cache, at a lower price. In both steps above, 128 prompt tokens came from the cache, because requests that started the same way had been sent a few minutes before. See Prompt caching and cost.

Not supported

  • The older functions and function_call fields: a request with them gets 400 INVALID_PARAMETER. Use tools and tool_choice.
  • Schema enforcement with strict, as explained above.
  • Tool calls the model didn't make. DeepSeek's documentation says its Chat Completions API doesn't support inserting them into a conversation.
Treat arguments as untrusted input. The model chooses the function and writes the arguments, and text in the conversation can steer it. Check both before your code acts, especially before it changes data, spends money or sends messages.

Next steps