ベストプラクティス
Practical advice for running in production: more reliable, faster and cheaper.
本文は現在英語版のみ提供しています。
Retries
Retry 429, 500, 502, 503, 504, timeouts and connection errors. Do not retry 400, 401, 403 or 404 — those are request or permission problems that will not fix themselves.
Use exponential backoff with jitter, 3–5 attempts at most, and honor Retry-After when present. The official OpenAI / Anthropic SDKs retry for you — just set max_retries.
import random, time
import openai
RETRYABLE = (openai.RateLimitError, openai.APITimeoutError, openai.APIConnectionError, openai.InternalServerError)
def chat_with_retry(**kwargs):
for attempt in range(5):
try:
return client.chat.completions.create(**kwargs)
except RETRYABLE:
if attempt == 4:
raise
time.sleep(min(30, 2 ** attempt) + random.random()) # exponential backoff + jitter
# 400 / 401 / 403 / 404 are not retriedTimeouts
- Long outputs and reasoning models can take minutes. Set the client's total timeout to at least 300 seconds.
- For streams, set the read timeout as the gap between chunks (e.g. 60 seconds) rather than limiting the whole request.
- Avoid very short timeouts with aggressive retries: the upstream may still be generating, which adds load and can lead to duplicate charges.
Concurrency
Cap in-flight requests with a semaphore or connection pool and ramp up gradually. Run batch jobs off-peak and back off on 429.
import asyncio
from openai import AsyncOpenAI
client = AsyncOpenAI(base_url="https://<your-endpoint>/v1", api_key="YOUR_API_KEY", max_retries=3)
sem = asyncio.Semaphore(8) # max requests in flight
async def one(prompt):
async with sem:
r = await client.chat.completions.create(
model="gpt-4.1-mini", messages=[{"role": "user", "content": prompt}])
return r.choices[0].message.content
async def main(prompts):
return await asyncio.gather(*(one(p) for p in prompts))Prompt caching
OpenAI and other models cache repeated long prefixes automatically (typically 1,024+ tokens). Cached tokens are billed at the lower cache read price and show up in usage.prompt_tokens_details.cached_tokens:
"usage": {
"prompt_tokens": 12840,
"completion_tokens": 312,
"prompt_tokens_details": { "cached_tokens": 12288 }
}Claude needs explicit cache_control breakpoints:
msg = client.messages.create(
model="claude-sonnet-4-5",
max_tokens=1024,
system=[{
"type": "text",
"text": LONG_STABLE_INSTRUCTIONS, # long, stable prefix
"cache_control": {"type": "ephemeral"}, # cache breakpoint
}],
messages=[{"role": "user", "content": question}],
)
print(msg.usage.cache_creation_input_tokens, msg.usage.cache_read_input_tokens)- Put stable content (system prompt, tool definitions, reference documents) first and the changing parts last.
- Keep timestamps, random IDs and other per-request values out of the prefix — any byte change invalidates the cache.
Cost control
- Match the model to the task: start small and upgrade only when quality requires it.
- Set sensible output caps to avoid needlessly long replies.
- Trim context: truncate or summarize long histories, and send only the document chunks you need.
- Lower the effort on reasoning models where you can.
- Use separate API keys per project and review each one's usage and cost in the console.
- Estimate per-request cost on the model page before you launch.
Security
- Keep API keys on the server; route front-end calls through your own backend.
- Log the time, model, status code and latency of each request — invaluable when debugging.