ドキュメント · 導入

ストリーミング

Set stream: true to receive tokens as server-sent events (SSE) while they are generated. You get the first token sooner and avoid client timeouts on long outputs.

本文は現在英語版のみ提供しています。

curl https://<your-endpoint>/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-4.1-mini",
    "messages": [{"role": "user", "content": "Hello!"}],
    "stream": true
  }'

Event format

Each event is a data: {JSON} line with incremental text in choices[0].delta.content. The stream ends with data: [DONE].

data: {"id":"chatcmpl-...","choices":[{"index":0,"delta":{"role":"assistant","content":""}}]}

data: {"id":"chatcmpl-...","choices":[{"index":0,"delta":{"content":"Hel"}}]}

data: {"id":"chatcmpl-...","choices":[{"index":0,"delta":{"content":"lo!"},"finish_reason":"stop"}]}

data: [DONE]

Usage in streams

Set stream_options.include_usage and the final chunk carries the usage for the whole request.

stream = client.chat.completions.create(
    model="gpt-4.1-mini",
    messages=[{"role": "user", "content": "Write a haiku about routers."}],
    stream=True,
    stream_options={"include_usage": True},
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)
    if chunk.usage:            # final chunk: empty choices, carries usage
        print("\n", chunk.usage)

Things to know

  • If you relay streams through your own reverse proxy (e.g. Nginx), disable response buffering (proxy_buffering off), or the content arrives all at once at the end.
  • Set the read timeout for streams as the gap between chunks, not the total request duration.
  • Reasoning models may think for a while before the first token. That is expected.
  • In the Anthropic format, stream events match Anthropic's (message_start, content_block_delta, …) — see Anthropic format.