Skip to content
أكاكوسAcacus
تواصل مع المبيعاتContact sales
Docs menu

Guides

Prompt caching and cost

When a request starts with the same tokens as earlier ones, DeepSeek can read that part from its cache, and you pay the lower cached-input price for it. It works on its own: there is nothing to turn on.

How it works

  • DeepSeek saves parts of each prompt in its cache: the prompt up to the end of the last message, the prompt with the reply, and the start of a long prompt at regular intervals.
  • A new request reads a saved part from the cache only when its prompt begins with all of it. So the part read from the cache can end before the first difference.
  • When two requests begin the same way, DeepSeek also saves the part they share. It does that after the second of them, so the second may read nothing from the cache and the third can read the shared part. A long start is also saved at regular intervals, so the second request often reads most of it already, as in the example below.
  • Tool definitions are part of the prompt, and are cached like the rest of it.
  • Your cache is private to your account. The gateway sends DeepSeek an id made from your account in user_id, so another account can't read from your cache, and you can't read from theirs. All your keys share it.
  • DeepSeek decides what it keeps, and for how long: it clears what isn't used, usually within a few hours to a few days, it says. A cache hit is not promised.

See it in the usage

These examples send two questions with the same long system message and print what usage says about each:

# A long system message that every request repeats word for word.
RULES=$(seq -f 'Rule %g: every request is billed per token in Libyan dinars.' 1 200 | paste -sd ' ' -)

for QUESTION in "What does rule 7 say?" "What does rule 150 say?"; do
  echo "$QUESTION"
  curl -sS https://acacus.ly/v1/chat/completions \
    -H "Authorization: Bearer $ACACUS_API_KEY" \
    -H "Content-Type: application/json" \
    -d @- <<EOF | grep -oE '"(prompt|completion)_tokens":[0-9]+|"prompt_cache_(hit|miss)_tokens":[0-9]+'
{
  "model": "deepseek-v4-flash",
  "messages": [
    {"role": "system", "content": "Answer questions about these rules in one sentence.\n\n$RULES"},
    {"role": "user", "content": "$QUESTION"}
  ],
  "max_tokens": 50,
  "thinking": {"type": "disabled"}
}
EOF
done

On an account whose cache was empty, the Python example printed:

Terminal
What does rule 7 say?
  prompt tokens: 3021
  read from the cache: 0
  not in the cache: 3021
  reply tokens: 16
What does rule 150 say?
  prompt tokens: 3021
  read from the cache: 2816
  not in the cache: 205
  reply tokens: 16

The first request found nothing in the cache. The second began with the same system message, so 2,816 tokens of its 3,021 tokens were read from the cache. The other 205 tokens, the end of the system message and the new question, were not. Run it again and even the first request reads from the cache.

Name
Type
Description
usage.prompt_tokens
integer
All prompt tokens: the ones read from the cache and the others.
usage.prompt_cache_hit_tokens
integer
Prompt tokens read from the cache, billed at the cached-input price.
usage.prompt_cache_miss_tokens
integer
Prompt tokens not read from the cache, billed at the input price.
usage.prompt_tokens_details.cached_tokens
integer
The same count as prompt_cache_hit_tokens, where OpenAI's format puts it. The OpenAI libraries read it from here.

In a stream the same fields are in the usage of the last chunk. See Streaming.

What it saves

ModelInputCached inputOutput
deepseek-v4-flash4.50 LYD / 1M tokens0.09 LYD / 1M tokens18.00 LYD / 1M tokens
deepseek-v4-pro19.80 LYD / 1M tokens0.66 LYD / 1M tokens59.40 LYD / 1M tokens

Libyan dinars (LYD) per 1 million tokens, in effect since 25 September 2026. Output includes reasoning tokens. The same price applies at every hour. Each request is billed to the dirham (0.001 LYD), rounded up, with no minimum.

The second request above, priced at today's rates for deepseek-v4-flash: input tokens at 4.50 LYD / 1M tokens, cached input tokens at 0.09 LYD / 1M tokens and output tokens at 18.00 LYD / 1M tokens.

Without the cacheWith the cache
Prompt tokens at the input price3,021 tokens205 tokens
Prompt tokens at the cached-input price0 tokens2,816 tokens
Reply tokens at the output price16 tokens16 tokens
Cost before rounding0.013883 LYD0.001464 LYD
Charged0.014 LYD0.002 LYD

Each request is billed to the dirham (0.001 LYD), rounded up, with no minimum, so on a short prompt the cache may save less than a dirham. It saves the most on long prompts that repeat: a long system message, a document you ask several questions about, or a long conversation.

Getting more from the cache

  • Put what stays the same first and what changes last: the system message, the tool definitions and any long document first, the new question at the end.
  • Keep the parts that stay the same exactly the same, down to spaces and the order of the tools.
  • In a conversation, send the earlier messages back unchanged. Editing an old message changes the start of every prompt after it.
  • Don't switch thinking on and off within a conversation. On deepseek-v4-flash each thinking setting has its own cache, so a switch makes one request read nothing from it. See Reasoning.
  • Send requests that share a start from the same account: accounts don't share a cache.

Other things that change the cost

  • Reply tokens have their own price, in the table above. Reasoning is billed as reply tokens, and a paid request thinks unless it sends "thinking": {"type": "disabled"} or "reasoning_effort": "none". See Reasoning.
  • max_tokens doesn't change what a reply costs: you pay for the tokens it has. It changes how long the reply may get, and how much balance a paid request needs before it runs.
  • Images are billed as prompt tokens. See Images.
  • How each request is charged, the free allowance and topping up are in Billing and the free tier.