文档 · 模型能力

视觉理解

带「视觉理解」能力的模型可以同时理解图片和文字:识别票据、读图表、看截图、描述照片等。

import base64

# Option 1: a publicly reachable image URL
url_part = {"type": "image_url", "image_url": {"url": "https://example.com/receipt.jpg"}}

# Option 2: a local file as a base64 data URL
b64 = base64.b64encode(open("receipt.jpg", "rb").read()).decode()
data_part = {"type": "image_url", "image_url": {"url": f"data:image/jpeg;base64,{b64}"}}

resp = client.chat.completions.create(
    model="gpt-4.1-mini",
    messages=[{
        "role": "user",
        "content": [
            {"type": "text", "text": "What is the total amount on this receipt?"},
            data_part,
        ],
    }],
)
print(resp.choices[0].message.content)

图片输入方式

  • 公网 URL:图片必须能被公网访问,内网地址和需要登录的链接无法读取。
  • base64 data URL:推荐用于本地文件或内网图片。
  • 一条消息中可以包含多张图片,与文字混排。

Anthropic 格式

msg = client.messages.create(
    model="claude-sonnet-4-5",
    max_tokens=1024,
    messages=[{
        "role": "user",
        "content": [
            {"type": "image", "source": {"type": "base64", "media_type": "image/jpeg", "data": b64}},
            # or {"type": "image", "source": {"type": "url", "url": "https://example.com/receipt.jpg"}}
            {"type": "text", "text": "What is the total amount on this receipt?"},
        ],
    }],
)

计费与限制

  • 图片会折算成输入 tokens(部分模型单独列出图片输入价格),分辨率越高消耗越多。
  • 上传前把图片缩放到够用的分辨率,可以明显降低成本和延迟。
  • 常见支持格式为 JPEG、PNG、WebP、GIF;单张图片过大会返回 400。