Guides
Prompt caching and cost
When a request starts with the same tokens as earlier ones, DeepSeek can read that part from its cache, and you pay the lower cached-input price for it. It works on its own: there is nothing to turn on.
How it works
- DeepSeek saves parts of each prompt in its cache: the prompt up to the end of the last message, the prompt with the reply, and the start of a long prompt at regular intervals.
- A new request reads a saved part from the cache only when its prompt begins with all of it. So the part read from the cache can end before the first difference.
- When two requests begin the same way, DeepSeek also saves the part they share. It does that after the second of them, so the second may read nothing from the cache and the third can read the shared part. A long start is also saved at regular intervals, so the second request often reads most of it already, as in the example below.
- Tool definitions are part of the prompt, and are cached like the rest of it.
- Your cache is private to your account. The gateway sends DeepSeek an id made from your account in
user_id, so another account can't read from your cache, and you can't read from theirs. All your keys share it. - DeepSeek decides what it keeps, and for how long: it clears what isn't used, usually within a few hours to a few days, it says. A cache hit is not promised.
See it in the usage
These examples send two questions with the same long system message and print what usage says about each:
# A long system message that every request repeats word for word.
RULES=$(seq -f 'Rule %g: every request is billed per token in Libyan dinars.' 1 200 | paste -sd ' ' -)
for QUESTION in "What does rule 7 say?" "What does rule 150 say?"; do
echo "$QUESTION"
curl -sS https://acacus.ly/v1/chat/completions \
-H "Authorization: Bearer $ACACUS_API_KEY" \
-H "Content-Type: application/json" \
-d @- <<EOF | grep -oE '"(prompt|completion)_tokens":[0-9]+|"prompt_cache_(hit|miss)_tokens":[0-9]+'
{
"model": "deepseek-v4-flash",
"messages": [
{"role": "system", "content": "Answer questions about these rules in one sentence.\n\n$RULES"},
{"role": "user", "content": "$QUESTION"}
],
"max_tokens": 50,
"thinking": {"type": "disabled"}
}
EOF
doneOn an account whose cache was empty, the Python example printed:
What does rule 7 say?
prompt tokens: 3021
read from the cache: 0
not in the cache: 3021
reply tokens: 16
What does rule 150 say?
prompt tokens: 3021
read from the cache: 2816
not in the cache: 205
reply tokens: 16The first request found nothing in the cache. The second began with the same system message, so 2,816 tokens of its 3,021 tokens were read from the cache. The other 205 tokens, the end of the system message and the new question, were not. Run it again and even the first request reads from the cache.
prompt_cache_hit_tokens, where OpenAI's format puts it. The OpenAI libraries read it from here.In a stream the same fields are in the usage of the last chunk. See Streaming.
What it saves
| Model | Input | Cached input | Output |
|---|---|---|---|
deepseek-v4-flash | 4.50 LYD / 1M tokens | 0.09 LYD / 1M tokens | 18.00 LYD / 1M tokens |
deepseek-v4-pro | 19.80 LYD / 1M tokens | 0.66 LYD / 1M tokens | 59.40 LYD / 1M tokens |
Libyan dinars (LYD) per 1 million tokens, in effect since 25 September 2026. Output includes reasoning tokens. The same price applies at every hour. Each request is billed to the dirham (0.001 LYD), rounded up, with no minimum.
The second request above, priced at today's rates for deepseek-v4-flash: input tokens at 4.50 LYD / 1M tokens, cached input tokens at 0.09 LYD / 1M tokens and output tokens at 18.00 LYD / 1M tokens.
| Without the cache | With the cache | |
|---|---|---|
| Prompt tokens at the input price | 3,021 tokens | 205 tokens |
| Prompt tokens at the cached-input price | 0 tokens | 2,816 tokens |
| Reply tokens at the output price | 16 tokens | 16 tokens |
| Cost before rounding | 0.013883 LYD | 0.001464 LYD |
| Charged | 0.014 LYD | 0.002 LYD |
Each request is billed to the dirham (0.001 LYD), rounded up, with no minimum, so on a short prompt the cache may save less than a dirham. It saves the most on long prompts that repeat: a long system message, a document you ask several questions about, or a long conversation.
Getting more from the cache
- Put what stays the same first and what changes last: the system message, the tool definitions and any long document first, the new question at the end.
- Keep the parts that stay the same exactly the same, down to spaces and the order of the tools.
- In a conversation, send the earlier messages back unchanged. Editing an old message changes the start of every prompt after it.
- Don't switch thinking on and off within a conversation. On
deepseek-v4-flasheach thinking setting has its own cache, so a switch makes one request read nothing from it. See Reasoning. - Send requests that share a start from the same account: accounts don't share a cache.
Other things that change the cost
- Reply tokens have their own price, in the table above. Reasoning is billed as reply tokens, and a paid request thinks unless it sends
"thinking": {"type": "disabled"}or"reasoning_effort": "none". See Reasoning. max_tokensdoesn't change what a reply costs: you pay for the tokens it has. It changes how long the reply may get, and how much balance a paid request needs before it runs.- Images are billed as prompt tokens. See Images.
- How each request is charged, the free allowance and topping up are in Billing and the free tier.