https://api.lazu.ai/v1/responsesResponses API
Responses 适合 reasoning 模型、多模态输入和 Lazu file_id 解引用。上传文件或使用较新的 OpenAI response 特性时,优先使用这个 endpoint。
计费按所选模型与线路的价格计算 token 用量;可能包含其他计费维度。 价格与线路 ↗
何时使用
使用 Responses
上传文件、reasoning 控制、文档 workflow,以及需要服务端标准化的多模态输入。
使用 Chat completions
简单聊天、SDK 兼容、tool calling,以及已经围绕
/v1/chat/completions 构建的应用。
请求 Body
modelstringrequired来自 /api/models/catalog 的模型 ID。优先选择
supported_endpoints 包含 /v1/responses 路径 的模型。
inputstring | object[]required文本输入,或 response input messages 数组。
instructionsstringnullableresponse 的系统级 instructions。
reasoningobjectnullable支持 reasoning 的模型可使用 effort 控制,例如
{"effort":"medium"}。
toolsobject[]nullableOpenAI-compatible Responses shape 的工具定义。
streambooleannullable当模型支持时,返回 response events stream。
max_output_tokensintegernullable生成输出 token 的上限。
Content parts
input_textcontent part发送给模型的文本内容。
input_imagecontent part图片内容。Lazu 可以解引用 purpose 为 vision 的上传文件。
input_filecontent part上传文件引用。Lazu 会在服务端读取文件内容后再转发给上游。
File 解引用
先通过 Files 上传文件,再引用返回的 file_id。
如果请求中发生了解引用,Lazu 会添加 X-Lazu-File-Dereference: 1。
限制:
- 单文件仍受 purpose 对应大小限制。
- 单次 Responses 调用中解引用的文件总大小必须低于 64 MB。
- Chat completions 不会自动解引用
file_id。
无状态兼容桥接
部分 catalog 条目会通过无损 Chat 桥接提供 Responses;可查看
supported_endpoints[].mode。走这类路由时,省略
store 可以正常请求,并返回
X-Lazu-Warning: stateless_bridge;显式
store: false 正常请求且不产生 warning。显式
store: true、无法保真的状态字段或未知跨协议字段会在请求上游前返回
400 protocol_bridge_unsupported。标准 Responses body 不会增加
Lazu 私有字段。
响应
idstringResponse ID。
outputobject[]输出消息、reasoning items、tool calls 或其它 response events。
usage.input_tokensintegernullable上游返回 usage 时的输入 token 数。
usage.output_tokensintegernullable上游返回 usage 时的输出 token 数。
usage.input_tokens_details.cached_tokensintegernullable支持 response-level cache usage 的 provider 返回的 cache read tokens。
完整对账请使用同一个 API Key 调用:
GET /api/usage/requests/{request_id}相关页面
完整工具续轮示例
先从当前 Key 的目录选择支持 Responses 和函数工具的模型,再设置 LAZU_API_KEY、LAZU_MODEL。无状态续轮要保留全部 output item(包括 reasoning),再用原 call_id 追加工具结果。桥接无法保留的加密或供应商私有 reasoning 状态仍只能走原生路径。
import json
import os
from openai import OpenAI, APIStatusError
client = OpenAI(
api_key=os.environ["LAZU_API_KEY"],
base_url=os.environ.get("LAZU_BASE_URL", "https://api.lazu.ai/v1"),
max_retries=0,
)
model = os.environ["LAZU_MODEL"]
tools = [{"type": "function", "name": "add", "description": "Add two integers",
"parameters": {"type": "object", "properties": {
"a": {"type": "integer"}, "b": {"type": "integer"}},
"required": ["a", "b"], "additionalProperties": False}}]
history = [{"role": "user", "content": "Use add to calculate 2 + 3."}]
try:
first = client.responses.create(model=model, input=history, tools=tools,
tool_choice={"type": "function", "name": "add"}, store=False,
max_output_tokens=1024)
history.extend(item.model_dump(exclude_unset=True) for item in first.output)
for item in first.output:
if item.type == "function_call":
if item.name != "add":
raise ValueError("Unexpected tool")
args = json.loads(item.arguments)
history.append({"type": "function_call_output", "call_id": item.call_id,
"output": json.dumps({"result": args["a"] + args["b"]})})
final = client.responses.create(model=model, input=history, tools=tools,
tool_choice="none", store=False, max_output_tokens=1024)
print(final.output_text)
except APIStatusError as exc:
print(exc.status_code, exc.response.headers.get("x-lazu-request-id"), exc.body)
raise收到输出后不要自动重放。413 request_body_too_large 应缩小请求;503 gateway_overloaded 可稍后退避重试,它不表示供应商故障。部分用量通过请求详情对账。
curl https://api.lazu.ai/v1/responses \
-H "Authorization: Bearer $LAZU_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-6-luna",
"input": [{
"role": "user",
"content": [
{"type": "input_text", "text": "Say hi"}
]
}]
}'
{
"id": "resp_01ABCDEF",
"model": "gpt-6-luna",
"output": [
{
"type": "message",
"role": "assistant",
"content": [{ "type": "output_text", "text": "Hi!" }]
}
],
"usage": { "input_tokens": 9, "output_tokens": 2 }
}
- Request ID
- req_demo_01
- 模型 · 线路
- gpt-6-luna · stable
- Token · 输入 / 输出
- 1,200 / 320
- Key
- demo-key
仅为示例,不是报价,也不是你这次请求的结果。发送上方请求即可看到你自己的回执。
查看真实请求回执 →适合场景
- Reasoning 模型
- 已上传的 PDF 或图片
- 结构化多模态输入