文档 · 生产实践
最佳实践
让生产环境更稳、更快、更省的实践建议。
重试
可重试:429、500、502、503、504、超时和连接错误。不要重试:400、401、403、404——参数或权限问题,重试也不会成功。
使用指数退避加随机抖动,最多 3–5 次;响应带 Retry-After 头时按它等待。OpenAI / Anthropic 官方 SDK 已内置重试,设置 max_retries 即可。
import random, time
import openai
RETRYABLE = (openai.RateLimitError, openai.APITimeoutError, openai.APIConnectionError, openai.InternalServerError)
def chat_with_retry(**kwargs):
for attempt in range(5):
try:
return client.chat.completions.create(**kwargs)
except RETRYABLE:
if attempt == 4:
raise
time.sleep(min(30, 2 ** attempt) + random.random()) # exponential backoff + jitter
# 400 / 401 / 403 / 404 are not retried超时
- 长输出和推理模型可能需要数分钟,客户端总超时建议不低于 300 秒。
- 流式请求按「两次数据之间的间隔」设置读超时(例如 60 秒),比限制整次时长更合理。
- 不要设置很短的超时再反复重试:上游可能仍在生成,既增加负载,也可能产生重复费用。
并发
用信号量或连接池限制同时在途的请求数,逐步加压;批量任务尽量错峰运行,并对 429 做退避。
import asyncio
from openai import AsyncOpenAI
client = AsyncOpenAI(base_url="https://<your-endpoint>/v1", api_key="YOUR_API_KEY", max_retries=3)
sem = asyncio.Semaphore(8) # max requests in flight
async def one(prompt):
async with sem:
r = await client.chat.completions.create(
model="gpt-4.1-mini", messages=[{"role": "user", "content": prompt}])
return r.choices[0].message.content
async def main(prompts):
return await asyncio.gather(*(one(p) for p in prompts))提示词缓存
OpenAI 等模型会自动缓存重复的长前缀(通常 1024 tokens 以上),命中部分按更低的「缓存读取」价格计费,可在 usage.prompt_tokens_details.cached_tokens 中看到:
"usage": {
"prompt_tokens": 12840,
"completion_tokens": 312,
"prompt_tokens_details": { "cached_tokens": 12288 }
}Claude 需要用 cache_control 显式标记缓存断点:
msg = client.messages.create(
model="claude-sonnet-4-5",
max_tokens=1024,
system=[{
"type": "text",
"text": LONG_STABLE_INSTRUCTIONS, # long, stable prefix
"cache_control": {"type": "ephemeral"}, # cache breakpoint
}],
messages=[{"role": "user", "content": question}],
)
print(msg.usage.cache_creation_input_tokens, msg.usage.cache_read_input_tokens)- 固定内容(系统提示词、工具定义、参考文档)放在最前面,每次变化的内容放在最后。
- 前缀中不要出现时间戳、随机 ID 等每次都变的内容,任何字节变化都会让缓存失效。
成本控制
- 按任务选模型:先用小模型,效果不够再升级。
- 设置合理的输出上限,避免无意义的长输出。
- 精简上下文:对长对话做截断或摘要,只传需要的文档片段。
- 推理模型按需调低 effort。
- 为不同项目使用不同 API Key,在控制台分别查看用量与费用。
- 上线前在模型详情页估算单次请求成本。
安全
- API Key 只放在服务端,前端请求通过你自己的后端转发。
- 记录每次请求的时间、模型、状态码和耗时,排查问题时会非常有用。