https://api.lazu.ai/v1/responsesResponses API
Use Responses for reasoning models, multimodal inputs and Lazu file_id dereferencing. This is the recommended endpoint when a request needs uploaded files or newer OpenAI response features.
BillingUsage tokens at the selected model and lane price; other billable dimensions may apply. Pricing & lanes ↗
When to use it
Use Responses
Uploaded files, reasoning controls, document workflows, and multimodal inputs that should be normalized server-side.
Use Chat completions
Simple chat, SDK compatibility, tool calling, and existing apps already
built around /v1/chat/completions.
Request body
modelstringrequiredModel ID from /api/models/catalog. Prefer models where
supported_endpoints includes the path /v1/responses.
inputstring | object[]requiredText input or an array of response input messages.
instructionsstringnullableSystem-level instructions for the response.
reasoningobjectnullableReasoning effort controls for supported models, for example
{"effort":"medium"}.
toolsobject[]nullableTool definitions using the OpenAI-compatible Responses shape.
streambooleannullableStreams response events when supported by the selected model.
max_output_tokensintegernullableUpper bound for generated output tokens.
Content parts
input_textcontent partText content sent to the model.
input_imagecontent partImage content. Lazu can dereference uploaded file_id values
with purpose vision.
input_filecontent partUploaded file reference. Lazu dereferences file content server-side before forwarding the request upstream.
File dereferencing
Upload via Files, then reference the resulting
file_id. Lazu adds X-Lazu-File-Dereference: 1 when the
request dereferenced files.
Limits:
- Single file purpose limit still applies.
- Total dereferenced files in one Responses call must stay under 64 MB.
- Chat completions does not auto-dereference
file_id.
Stateless compatibility bridges
Some catalog entries expose Responses through a lossless Chat bridge; check
supported_endpoints[].mode. On such a route, omitted
store is accepted and returns
X-Lazu-Warning: stateless_bridge, while
store: false is accepted without a warning. Explicit
store: true, stateful fields, and unknown cross-protocol fields
return 400 protocol_bridge_unsupported before an upstream call.
The standard Responses body is not extended with Lazu-only fields.
Response
idstringResponse ID.
outputobject[]Output messages, reasoning items, tool calls, or other response events.
usage.input_tokensintegernullableInput token count when the upstream reports usage.
usage.output_tokensintegernullableOutput token count when the upstream reports usage.
usage.input_tokens_details.cached_tokensintegernullableCache read tokens for providers that expose response-level cache usage.
For full reconciliation, use
GET /api/usage/requests/{request_id} with the same API key.
See also
Complete tool round trip
Select a model with Responses and function tools from your token-scoped catalog, then set LAZU_API_KEY and LAZU_MODEL. Keep every output item when continuing a stateless conversation, including reasoning items; append the tool result with the same call_id. Encrypted or provider-specific reasoning state remains native-only when a bridge cannot preserve it.
import json
import os
from openai import OpenAI, APIStatusError
client = OpenAI(
api_key=os.environ["LAZU_API_KEY"],
base_url=os.environ.get("LAZU_BASE_URL", "https://api.lazu.ai/v1"),
max_retries=0,
)
model = os.environ["LAZU_MODEL"]
tools = [{"type": "function", "name": "add", "description": "Add two integers",
"parameters": {"type": "object", "properties": {
"a": {"type": "integer"}, "b": {"type": "integer"}},
"required": ["a", "b"], "additionalProperties": False}}]
history = [{"role": "user", "content": "Use add to calculate 2 + 3."}]
try:
first = client.responses.create(model=model, input=history, tools=tools,
tool_choice={"type": "function", "name": "add"}, store=False,
max_output_tokens=1024)
history.extend(item.model_dump(exclude_unset=True) for item in first.output)
for item in first.output:
if item.type == "function_call":
if item.name != "add":
raise ValueError("Unexpected tool")
args = json.loads(item.arguments)
history.append({"type": "function_call_output", "call_id": item.call_id,
"output": json.dumps({"result": args["a"] + args["b"]})})
final = client.responses.create(model=model, input=history, tools=tools,
tool_choice="none", store=False, max_output_tokens=1024)
print(final.output_text)
except APIStatusError as exc:
print(exc.status_code, exc.response.headers.get("x-lazu-request-id"), exc.body)
raiseDo not replay automatically after output has arrived. For 413 request_body_too_large, reduce the request; for 503 gateway_overloaded, retry later with backoff. Local overload does not imply a provider failure. Use request details to reconcile partial usage.
curl https://api.lazu.ai/v1/responses \
-H "Authorization: Bearer $LAZU_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-6-luna",
"input": [{
"role": "user",
"content": [
{"type": "input_text", "text": "Say hi"}
]
}]
}'
{
"id": "resp_01ABCDEF",
"model": "gpt-6-luna",
"output": [
{
"type": "message",
"role": "assistant",
"content": [{ "type": "output_text", "text": "Hi!" }]
}
],
"usage": { "input_tokens": 9, "output_tokens": 2 }
}
- Request ID
- req_demo_01
- Model · Lane
- gpt-6-luna · stable
- Tokens · input / output
- 1,200 / 320
- Key
- demo-key
Example only — not a quote and not the result of your request. Send the request above to see your own receipt.
Read a real request receipt →Best for
- Reasoning models
- Uploaded PDFs or images
- Structured multimodal input