Skip to content
أكاكوسAcacus
تواصل مع المبيعاتContact sales
Docs menu

Guides

Streaming

Set stream to true and the reply arrives in small pieces while the model writes it, as Server-Sent Events. Your app can show the answer as it grows.

Turn streaming on

Add "stream": true to the request. Everything else stays the same as a request without it.

curl -N https://acacus.ly/v1/chat/completions \
  -H "Authorization: Bearer $ACACUS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek-v4-flash",
    "messages": [
      {"role": "user", "content": "Count from 1 to 5."}
    ],
    "stream": true,
    "thinking": {"type": "disabled"}
  }'

With curl, -N prints each event as it arrives. The OpenAI libraries read the events for you and give you one chunk at a time. Each chunk holds the next piece of the answer in choices[0].delta.content.

What the stream looks like

This is what the curl example received. The line ... stands for 11 chunks left out here.

Stream
data: {"id":"3aa4ba59-9e85-4743-be89-1251c433652e","object":"chat.completion.chunk","created":1790380302,"model":"deepseek-flash","system_fingerprint":"aeb56401ca74e127821c4f9126dcb669","choices":[{"index":0,"delta":{"role":"assistant","content":""},"logprobs":null,"finish_reason":null}],"usage":null}

data: {"id":"3aa4ba59-9e85-4743-be89-1251c433652e","object":"chat.completion.chunk","created":1790380302,"model":"deepseek-flash","system_fingerprint":"aeb56401ca74e127821c4f9126dcb669","choices":[{"index":0,"delta":{"content":"1"},"logprobs":null,"finish_reason":null}],"usage":null}

data: {"id":"3aa4ba59-9e85-4743-be89-1251c433652e","object":"chat.completion.chunk","created":1790380302,"model":"deepseek-flash","system_fingerprint":"aeb56401ca74e127821c4f9126dcb669","choices":[{"index":0,"delta":{"content":","},"logprobs":null,"finish_reason":null}],"usage":null}

...

data: {"id":"3aa4ba59-9e85-4743-be89-1251c433652e","object":"chat.completion.chunk","created":1790380302,"model":"deepseek-flash","system_fingerprint":"aeb56401ca74e127821c4f9126dcb669","choices":[{"index":0,"delta":{"content":"."},"logprobs":null,"finish_reason":null}],"usage":null}

data: {"id":"3aa4ba59-9e85-4743-be89-1251c433652e","object":"chat.completion.chunk","created":1790380302,"model":"deepseek-flash","system_fingerprint":"aeb56401ca74e127821c4f9126dcb669","choices":[{"index":0,"delta":{"content":""},"logprobs":null,"finish_reason":"stop"}],"usage":{"prompt_tokens":12,"completion_tokens":14,"total_tokens":26,"prompt_tokens_details":{"cached_tokens":0},"prompt_cache_hit_tokens":0,"prompt_cache_miss_tokens":12}}

data: [DONE]
  • Each event is one line: data: and a JSON object. A blank line follows each event.
  • Every chunk has the same id and "object": "chat.completion.chunk".
  • The first chunk sets delta.role to "assistant", with empty content.
  • Each chunk after it adds a piece of the answer in delta.content. Join the pieces in order.
  • The last chunk has finish_reason and usage. Every chunk before it has "usage": null.
  • The stream ends with data: [DONE].

DeepSeek can also send lines that start with a colon, such as : keep-alive, while a request waits to start. They are comments: skip them. The OpenAI libraries do this for you.

The last chunk

The chunk that ends the reply carries the finish reason and the token counts together:

Last chunk
{
  "id": "3aa4ba59-9e85-4743-be89-1251c433652e",
  "object": "chat.completion.chunk",
  "created": 1790380302,
  "model": "deepseek-flash",
  "system_fingerprint": "aeb56401ca74e127821c4f9126dcb669",
  "choices": [
    {
      "index": 0,
      "delta": {
        "content": ""
      },
      "logprobs": null,
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 12,
    "completion_tokens": 14,
    "total_tokens": 26,
    "prompt_tokens_details": {
      "cached_tokens": 0
    },
    "prompt_cache_hit_tokens": 0,
    "prompt_cache_miss_tokens": 12
  }
}

finish_reason is stop when the model finished, length when the reply reached max_tokens, and tool_calls when the model asks your code to run a function. DeepSeek's documentation lists three more values, for a reply that was cut short: content_filter, insufficient_system_resource and aborted.

A streamed reply always carries usage. The gateway turns on stream_options.include_usage for every streamed request, because it needs the token counts to bill you, so sending include_usage: false changes nothing. The usage comes in the same chunk as finish_reason, not in an extra chunk with an empty choices list. Read it from whichever chunk has it, as the examples do.

Reasoning and tool calls

With thinking on, the reasoning streams first, in delta.reasoning_content, and the answer follows in delta.content. Thinking is for paid requests only. See Reasoning.

A tool call streams in delta.tool_calls. The first piece of each call has its index, id and function.name. The next pieces add text to function.arguments. Join the pieces that share an index. The last chunk has finish_reason set to tool_calls. See Tool calling for an example.

Errors

The request is checked before the reply starts. If something is wrong, you get one JSON error with its HTTP status instead of a stream: for example 401 for a bad key, 402 when your balance and free allowance don't cover the request, 400 for a bad parameter and 429 for too many requests. A value DeepSeek refuses, such as a temperature above 2, gets DeepSeek's JSON error with status 400 in the same way.

The OpenAI libraries raise these as errors before you read the first chunk. Without a library, check the status code before you read events. The codes are listed in Errors.

If DeepSeek's stream breaks in the middle, your stream ends without a chunk that has a finish_reason and without data: [DONE]. The OpenAI libraries then end the loop without an error, so check that a finish_reason arrived, and treat the reply as incomplete if it didn't. You are billed an estimate: the input, and the part of the reply that was sent to you.

Stopping early

Closing the connection, or breaking out of the loop, doesn't stop the reply. DeepSeek writes it to the end, and you pay for all of it. Until it ends, it also counts toward your 3 requests at a time. To limit how long a reply can be, and what it can cost, set max_tokens.

Long replies

nginx, the server in front of the API, closes a connection that has sent nothing for 120 seconds, and answers 504. Cloudflare, in front of nginx, waits longer. A reply that isn't streamed sends nothing until it is complete. So a long one, with a large max_tokens or with thinking on, can be cut off after the model has done the work, and you still pay for it. A streamed reply sends each piece as soon as it is written. Stream long replies.

Response headers

These are the headers of the curl example's reply that you may want to read:

Headers
content-type: text/event-stream
cache-control: no-cache
x-ratelimit-limit: 30
x-ratelimit-remaining: 23
x-ratelimit-reset: 1790380320
x-request-id: e9c2cd35-9fc5-4e19-aa2d-3ef37b48bfa4

The x-ratelimit-* headers show your per-minute limit, what is left of it, and when it resets (see Rate limits and other limits). Quote x-request-id when you ask support about a request.

Without a library

Any HTTP client that can read a response while it arrives can stream. Check the status code first: an error is a single JSON object. Then read the body line by line. Skip blank lines and lines that start with a colon, parse the JSON after data:, and stop at data: [DONE]. These examples use Python's requests library and the fetch built into Node.js.

# pip install requests
import json
import os
import requests

response = requests.post(
    "https://acacus.ly/v1/chat/completions",
    headers={"Authorization": f"Bearer {os.environ['ACACUS_API_KEY']}"},
    json={
        "model": "deepseek-v4-flash",
        "messages": [{"role": "user", "content": "Count from 1 to 5."}],
        "stream": True,
        "thinking": {"type": "disabled"},
    },
    stream=True,
)
if response.status_code != 200:
    # An error comes as one JSON object, not as events.
    raise SystemExit(f"{response.status_code}: {response.text}")

usage = None
for raw in response.iter_lines(chunk_size=None):  # each piece as it arrives
    line = raw.decode("utf-8")
    if not line.startswith("data:"):
        continue  # a blank line between events, or a ": comment" line
    data = line[len("data:"):].strip()
    if data == "[DONE]":
        break
    chunk = json.loads(data)
    for choice in chunk["choices"]:
        print(choice["delta"].get("content") or "", end="", flush=True)
    if chunk.get("usage"):
        usage = chunk["usage"]

print()
print(usage)

Next steps