Guides
Reasoning
DeepSeek's models can think before they answer: they write out their reasoning first, then the answer. You get the reasoning in its own field, and you pay for it as output tokens. Thinking needs a balance: the free allowance doesn't include it.
Turn thinking on or off
Two request fields control it. Neither is an OpenAI parameter; both are DeepSeek's.
"enabled" turns thinking on, "disabled" turns it off. "adaptive" is accepted too and behaved like "enabled" in our tests. When the request has both fields, this one decides whether the model thinks."low", "high" or "max". "none" turns thinking off. DeepSeek also accepts "minimal" (treated as low) and "medium" and "xhigh" (treated as high).In the OpenAI Python library, reasoning_effort is a normal argument and thinking goes in extra_body. In the Node.js library, put both in the request object; for TypeScript, see Libraries and SDKs. A budget_tokens value has no effect: to limit reasoning, use max_tokens.
| The request has | Paid request | Free request |
|---|---|---|
| Neither field | Thinks, at high effort | Doesn't think (the gateway turns thinking off) |
"thinking": {"type": "enabled"} | Thinks | Refused: 402 FREE_TIER_THINKING |
"thinking": {"type": "disabled"} | Doesn't think | Doesn't think |
reasoning_effort "low", "high" or "max", and no thinking | Thinks, at that effort | Refused: 402 FREE_TIER_THINKING |
reasoning_effort "none", and no thinking | Doesn't think | Doesn't think |
Paid requests think unless you turn it off
"thinking": {"type": "disabled"}, as the examples in these docs do.Send a request with thinking
curl https://acacus.ly/v1/chat/completions \
-H "Authorization: Bearer $ACACUS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek-v4-flash",
"messages": [
{"role": "user", "content": "Which is larger, 9.11 or 9.8?"}
],
"thinking": {"type": "enabled"},
"reasoning_effort": "high"
}'These examples need a balance. On a free account they get 402 FREE_TIER_THINKING (see Free tier below).
Read the reasoning
The reasoning comes in reasoning_content, next to content in the message. The reply has these fields besides the usual ones:
content holds the answer. Absent when the model didn't think.delta.content.Reasoning can take a while, so it is often better to stream it:
# pip install openai
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["ACACUS_API_KEY"],
base_url="https://acacus.ly/v1",
)
stream = client.chat.completions.create(
model="deepseek-v4-flash",
messages=[
{"role": "user", "content": "Which is larger, 9.11 or 9.8?"},
],
reasoning_effort="high",
stream=True,
extra_body={"thinking": {"type": "enabled"}},
)
answering = False
for chunk in stream:
if not chunk.choices:
continue
delta = chunk.choices[0].delta
reasoning = getattr(delta, "reasoning_content", None)
if reasoning:
print(reasoning, end="", flush=True) # the reasoning comes first
if delta.content:
if not answering:
print("\n\nAnswer: ", end="")
answering = True
print(delta.content, end="", flush=True) # then the answer
if chunk.usage:
print("\n", chunk.usage)Reasoning, max_tokens and cost
The reasoning counts toward max_tokens. When you don't set max_tokens, the gateway sends 8,192 tokens, also with thinking on. You can set up to 32,768 tokens: a larger value is lowered to 32,768 tokens.
If the reasoning uses up max_tokens, the reply ends with finish_reason set to length, with part of an answer or none. Give hard questions a larger max_tokens. Before a paid request runs, your balance must cover its full max_tokens as output: see Billing and the free tier.
A reply that isn't streamed sends nothing until it is complete, and nginx, in front of the API, closes the connection after 120 seconds without data, with a 504. Stream long reasoning. See Streaming.
Conversations and tools
In a conversation without tools, you don't need to send the reasoning back: DeepSeek ignores reasoning_content in earlier assistant messages. In a request with tools, every earlier assistant message must keep its reasoning_content, or DeepSeek answers 400. See Tool calling.
Keep one setting for the whole conversation. On deepseek-v4-flash each thinking setting has its own prompt cache, so changing it makes the next request's prompt uncached, at the full input price. See Prompt caching and cost.
Other parameters
DeepSeek documents how thinking changes some other parameters:
temperature,presence_penaltyandfrequency_penaltyhave no effect. They don't cause an error.top_pworks only with thinking on, from 0.95 to 1. A lower value counts as 0.95. With thinking off, DeepSeek ignores it.tool_choice"required"or a named function gets a400error. Turn thinking off to force a tool call.
deepseek-v4-pro
deepseek-v4-pro takes the same fields. Whether "thinking": {"type": "disabled"} turns its thinking off has not been tested, so check usage.completion_tokens_details.reasoning_tokens in its replies.
Free tier
Free requests don't think. A free request that asks for thinking (the rows marked “Refused” in the table above) gets this error before it runs, and nothing is charged:
{
"error": {
"message": "Thinking isn't included in the free tier. Top up to use it, or send the request without thinking: no reasoning_effort (or \"none\"), and no thinking field or {\"type\": \"disabled\"}.",
"type": "insufficient_quota",
"code": "FREE_TIER_THINKING",
"param": null,
"balance": 0
}
}A free request with neither field is sent with thinking off, so none of its output, at most 2,048 tokens on the free tier, goes to reasoning. See Billing and the free tier.