Skip to content
أكاكوسAcacus
تواصل مع المبيعاتContact sales
Docs menu

Guides

Reasoning

DeepSeek's models can think before they answer: they write out their reasoning first, then the answer. You get the reasoning in its own field, and you pay for it as output tokens. Thinking needs a balance: the free allowance doesn't include it.

Turn thinking on or off

Two request fields control it. Neither is an OpenAI parameter; both are DeepSeek's.

Name
Type
Description
thinking.typeOptional
string
"enabled" turns thinking on, "disabled" turns it off. "adaptive" is accepted too and behaved like "enabled" in our tests. When the request has both fields, this one decides whether the model thinks.
reasoning_effortOptional
string
How much the model thinks: "low", "high" or "max". "none" turns thinking off. DeepSeek also accepts "minimal" (treated as low) and "medium" and "xhigh" (treated as high).

In the OpenAI Python library, reasoning_effort is a normal argument and thinking goes in extra_body. In the Node.js library, put both in the request object; for TypeScript, see Libraries and SDKs. A budget_tokens value has no effect: to limit reasoning, use max_tokens.

The request hasPaid requestFree request
Neither fieldThinks, at high effortDoesn't think (the gateway turns thinking off)
"thinking": {"type": "enabled"}ThinksRefused: 402 FREE_TIER_THINKING
"thinking": {"type": "disabled"}Doesn't thinkDoesn't think
reasoning_effort "low", "high" or "max", and no thinkingThinks, at that effortRefused: 402 FREE_TIER_THINKING
reasoning_effort "none", and no thinkingDoesn't thinkDoesn't think

Paid requests think unless you turn it off

Thinking is on by default. A paid request with neither field thinks, at high effort, and the reasoning is billed as output tokens. If you don't need it, send "thinking": {"type": "disabled"}, as the examples in these docs do.

Send a request with thinking

curl https://acacus.ly/v1/chat/completions \
  -H "Authorization: Bearer $ACACUS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek-v4-flash",
    "messages": [
      {"role": "user", "content": "Which is larger, 9.11 or 9.8?"}
    ],
    "thinking": {"type": "enabled"},
    "reasoning_effort": "high"
  }'

These examples need a balance. On a free account they get 402 FREE_TIER_THINKING (see Free tier below).

Read the reasoning

The reasoning comes in reasoning_content, next to content in the message. The reply has these fields besides the usual ones:

Name
Type
Description
choices[].message.reasoning_content
string
The reasoning, written before the answer. content holds the answer. Absent when the model didn't think.
choices[].delta.reasoning_content
string
In a stream: the next piece of the reasoning. The reasoning streams first, then the answer in delta.content.
usage.completion_tokens_details.reasoning_tokens
integer
How many of the output tokens were reasoning.
usage.completion_tokens
integer
All output tokens, the reasoning included. This is what you pay the output price for.

Reasoning can take a while, so it is often better to stream it:

# pip install openai
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["ACACUS_API_KEY"],
    base_url="https://acacus.ly/v1",
)

stream = client.chat.completions.create(
    model="deepseek-v4-flash",
    messages=[
        {"role": "user", "content": "Which is larger, 9.11 or 9.8?"},
    ],
    reasoning_effort="high",
    stream=True,
    extra_body={"thinking": {"type": "enabled"}},
)

answering = False
for chunk in stream:
    if not chunk.choices:
        continue
    delta = chunk.choices[0].delta
    reasoning = getattr(delta, "reasoning_content", None)
    if reasoning:
        print(reasoning, end="", flush=True)  # the reasoning comes first
    if delta.content:
        if not answering:
            print("\n\nAnswer: ", end="")
            answering = True
        print(delta.content, end="", flush=True)  # then the answer
    if chunk.usage:
        print("\n", chunk.usage)

Reasoning, max_tokens and cost

The reasoning counts toward max_tokens. When you don't set max_tokens, the gateway sends 8,192 tokens, also with thinking on. You can set up to 32,768 tokens: a larger value is lowered to 32,768 tokens.

If the reasoning uses up max_tokens, the reply ends with finish_reason set to length, with part of an answer or none. Give hard questions a larger max_tokens. Before a paid request runs, your balance must cover its full max_tokens as output: see Billing and the free tier.

A reply that isn't streamed sends nothing until it is complete, and nginx, in front of the API, closes the connection after 120 seconds without data, with a 504. Stream long reasoning. See Streaming.

Conversations and tools

In a conversation without tools, you don't need to send the reasoning back: DeepSeek ignores reasoning_content in earlier assistant messages. In a request with tools, every earlier assistant message must keep its reasoning_content, or DeepSeek answers 400. See Tool calling.

Keep one setting for the whole conversation. On deepseek-v4-flash each thinking setting has its own prompt cache, so changing it makes the next request's prompt uncached, at the full input price. See Prompt caching and cost.

Other parameters

DeepSeek documents how thinking changes some other parameters:

  • temperature, presence_penalty and frequency_penalty have no effect. They don't cause an error.
  • top_p works only with thinking on, from 0.95 to 1. A lower value counts as 0.95. With thinking off, DeepSeek ignores it.
  • tool_choice "required" or a named function gets a 400 error. Turn thinking off to force a tool call.

deepseek-v4-pro

deepseek-v4-pro takes the same fields. Whether "thinking": {"type": "disabled"} turns its thinking off has not been tested, so check usage.completion_tokens_details.reasoning_tokens in its replies.

Free tier

Free requests don't think. A free request that asks for thinking (the rows marked “Refused” in the table above) gets this error before it runs, and nothing is charged:

Response (402)
{
  "error": {
    "message": "Thinking isn't included in the free tier. Top up to use it, or send the request without thinking: no reasoning_effort (or \"none\"), and no thinking field or {\"type\": \"disabled\"}.",
    "type": "insufficient_quota",
    "code": "FREE_TIER_THINKING",
    "param": null,
    "balance": 0
  }
}

A free request with neither field is sent with thinking off, so none of its output, at most 2,048 tokens on the free tier, goes to reasoning. See Billing and the free tier.

Next steps