Docs · Production

Best practices

Practical advice for running in production: more reliable, faster and cheaper.

Retries

Retry 429, 500, 502, 503, 504, timeouts and connection errors. Do not retry 400, 401, 403 or 404 — those are request or permission problems that will not fix themselves.

Use exponential backoff with jitter, 3–5 attempts at most, and honor Retry-After when present. The official OpenAI / Anthropic SDKs retry for you — just set max_retries.

import random, time
import openai

RETRYABLE = (openai.RateLimitError, openai.APITimeoutError, openai.APIConnectionError, openai.InternalServerError)

def chat_with_retry(**kwargs):
    for attempt in range(5):
        try:
            return client.chat.completions.create(**kwargs)
        except RETRYABLE:
            if attempt == 4:
                raise
            time.sleep(min(30, 2 ** attempt) + random.random())   # exponential backoff + jitter
        # 400 / 401 / 403 / 404 are not retried

Timeouts

  • Long outputs and reasoning models can take minutes. Set the client's total timeout to at least 300 seconds.
  • For streams, set the read timeout as the gap between chunks (e.g. 60 seconds) rather than limiting the whole request.
  • Avoid very short timeouts with aggressive retries: the upstream may still be generating, which adds load and can lead to duplicate charges.

Concurrency

Cap in-flight requests with a semaphore or connection pool and ramp up gradually. Run batch jobs off-peak and back off on 429.

import asyncio
from openai import AsyncOpenAI

client = AsyncOpenAI(base_url="https://<your-endpoint>/v1", api_key="YOUR_API_KEY", max_retries=3)
sem = asyncio.Semaphore(8)            # max requests in flight

async def one(prompt):
    async with sem:
        r = await client.chat.completions.create(
            model="gpt-4.1-mini", messages=[{"role": "user", "content": prompt}])
        return r.choices[0].message.content

async def main(prompts):
    return await asyncio.gather(*(one(p) for p in prompts))

Prompt caching

OpenAI and other models cache repeated long prefixes automatically (typically 1,024+ tokens). Cached tokens are billed at the lower cache read price and show up in usage.prompt_tokens_details.cached_tokens:

"usage": {
  "prompt_tokens": 12840,
  "completion_tokens": 312,
  "prompt_tokens_details": { "cached_tokens": 12288 }
}

Claude needs explicit cache_control breakpoints:

msg = client.messages.create(
    model="claude-sonnet-4-5",
    max_tokens=1024,
    system=[{
        "type": "text",
        "text": LONG_STABLE_INSTRUCTIONS,           # long, stable prefix
        "cache_control": {"type": "ephemeral"},     # cache breakpoint
    }],
    messages=[{"role": "user", "content": question}],
)
print(msg.usage.cache_creation_input_tokens, msg.usage.cache_read_input_tokens)
  • Put stable content (system prompt, tool definitions, reference documents) first and the changing parts last.
  • Keep timestamps, random IDs and other per-request values out of the prefix — any byte change invalidates the cache.

Cost control

  • Match the model to the task: start small and upgrade only when quality requires it.
  • Set sensible output caps to avoid needlessly long replies.
  • Trim context: truncate or summarize long histories, and send only the document chunks you need.
  • Lower the effort on reasoning models where you can.
  • Use separate API keys per project and review each one's usage and cost in the console.
  • Estimate per-request cost on the model page before you launch.

Security

  • Keep API keys on the server; route front-end calls through your own backend.
  • Log the time, model, status code and latency of each request — invaluable when debugging.