https://api.lazu.ai/v1/responsesResponses API
Responses 適合 reasoning 模型、多模態輸入和 Lazu file_id 解引用。上傳檔案或使用較新的 OpenAI response 特性時,優先使用這個 endpoint。
計費依所選模型與線路的價格計算 token 用量;可能包含其他計費維度。 價格與線路 ↗
何時使用
使用 Responses
上傳檔案、reasoning 控制、文件 workflow,以及需要服務端標準化的多模態輸入。
使用 Chat completions
簡單聊天、SDK 相容、tool calling,以及已經圍繞
/v1/chat/completions 建構的應用。
請求 Body
modelstringrequired來自 /api/models/catalog 的模型 ID。優先選擇
supported_endpoints 包含 /v1/responses 路径 的模型。
inputstring | object[]required文字輸入,或 response input messages 陣列。
instructionsstringnullableresponse 的系統級 instructions。
reasoningobjectnullable支援 reasoning 的模型可使用 effort 控制,例如
{"effort":"medium"}。
toolsobject[]nullableOpenAI-compatible Responses shape 的工具定義。
streambooleannullable當模型支援時,返回 response events stream。
max_output_tokensintegernullable生成輸出 token 的上限。
File 解引用
先透過 Files 上傳檔案,再引用返回的
file_id。如果請求中發生了解引用,Lazu 會加入
X-Lazu-File-Dereference: 1。
限制:
- 單檔案仍受 purpose 對應大小限制。
- 單次 Responses 呼叫中解引用的檔案總大小必須低於 64 MB。
- Chat completions 不會自動解引用
file_id。
無狀態相容橋接
部分 catalog 項目會透過無損 Chat 橋接提供 Responses;可查看
supported_endpoints[].mode。走這類路由時,省略
store 可以正常請求,並回傳
X-Lazu-Warning: stateless_bridge;明確指定
store: false 正常請求且不產生 warning。明確指定
store: true、無法保真的狀態欄位或未知跨協議欄位,會在呼叫上游前回傳
400 protocol_bridge_unsupported。標準 Responses body 不會增加
Lazu 私有欄位。
響應
idstringResponse ID。
outputobject[]輸出訊息、reasoning items、tool calls 或其它 response events。
usage.input_tokensintegernullable上游返回 usage 時的輸入 token 數。
usage.output_tokensintegernullable上游返回 usage 時的輸出 token 數。
usage.input_tokens_details.cached_tokensintegernullable支援 response-level cache usage 的 provider 返回的 cache read tokens。
完整對帳請使用同一把 API Key 呼叫:
GET /api/usage/requests/{request_id}相關頁面
完整工具續輪範例
先從目前 Key 的目錄選擇支援 Responses 和函式工具的模型,再設定 LAZU_API_KEY、LAZU_MODEL。無狀態續輪要保留全部 output item(包括 reasoning),再用原 call_id 附加工具結果。橋接無法保留的加密或供應商私有 reasoning 狀態仍只能走原生路徑。
import json
import os
from openai import OpenAI, APIStatusError
client = OpenAI(
api_key=os.environ["LAZU_API_KEY"],
base_url=os.environ.get("LAZU_BASE_URL", "https://api.lazu.ai/v1"),
max_retries=0,
)
model = os.environ["LAZU_MODEL"]
tools = [{"type": "function", "name": "add", "description": "Add two integers",
"parameters": {"type": "object", "properties": {
"a": {"type": "integer"}, "b": {"type": "integer"}},
"required": ["a", "b"], "additionalProperties": False}}]
history = [{"role": "user", "content": "Use add to calculate 2 + 3."}]
try:
first = client.responses.create(model=model, input=history, tools=tools,
tool_choice={"type": "function", "name": "add"}, store=False,
max_output_tokens=1024)
history.extend(item.model_dump(exclude_unset=True) for item in first.output)
for item in first.output:
if item.type == "function_call":
if item.name != "add":
raise ValueError("Unexpected tool")
args = json.loads(item.arguments)
history.append({"type": "function_call_output", "call_id": item.call_id,
"output": json.dumps({"result": args["a"] + args["b"]})})
final = client.responses.create(model=model, input=history, tools=tools,
tool_choice="none", store=False, max_output_tokens=1024)
print(final.output_text)
except APIStatusError as exc:
print(exc.status_code, exc.response.headers.get("x-lazu-request-id"), exc.body)
raise收到輸出後不要自動重播。413 request_body_too_large 應縮小請求;503 gateway_overloaded 可稍後退避重試,它不表示供應商故障。部分用量透過請求詳情對帳。
curl https://api.lazu.ai/v1/responses \
-H "Authorization: Bearer $LAZU_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-6-luna",
"input": [{
"role": "user",
"content": [
{"type": "input_text", "text": "Say hi"}
]
}]
}'
{
"id": "resp_01ABCDEF",
"model": "gpt-6-luna",
"output": [
{
"type": "message",
"role": "assistant",
"content": [{ "type": "output_text", "text": "Hi!" }]
}
],
"usage": { "input_tokens": 9, "output_tokens": 2 }
}
- Request ID
- req_demo_01
- 模型 · 線路
- gpt-6-luna · stable
- Token · 輸入 / 輸出
- 1,200 / 320
- Key
- demo-key
僅為範例,不是報價,也不是你這次請求的結果。送出上方請求即可看到你自己的回執。
查看真實請求回執 →適合場景
- Reasoning 模型
- 已上傳的 PDF 或圖片
- 結構化多模態輸入