https://api.lazu.ai/v1/chat/completionsChat completions
Send OpenAI-compatible chat requests through Lazu. Use this page as the working reference: base URL, auth, required body fields, examples, response usage, streaming and tool calling all live here.
BillingUsage tokens at the selected model and lane price; other billable dimensions may apply. Pricing & lanes ↗
Basic configuration
Base URL
Use https://api.lazu.ai/v1 with OpenAI SDKs, or call the full
path https://api.lazu.ai/v1/chat/completions directly.
Model discovery
Read GET /api/models/catalog at runtime. Filter entries where
supported_endpoint_types contains chat.
Request body
modelstringrequiredModel ID, route alias or future route policy name. For explicit models, pass
IDs from GET /api/models/catalog. For SDK compatibility,
/v1/models remains available as a flat list.
messagesobject[]requiredOrdered conversation messages. Roles follow the OpenAI shape:
system, user, assistant, and
tool.
rolestringcontentstring | content_part[]tool_call_idstringstreambooleannullableWhen true, Lazu forwards a Server-Sent Events stream and ends
with data: [DONE].
toolsobject[]nullableFunction/tool definitions. Check parameters.tools in the model
catalog before sending tools to a model.
type"function"function.namestringfunction.parametersobjecttool_choicestring | objectnullableOpenAI-compatible tool choice control. Use auto,
none, required, or a named tool object when the
selected model supports it.
response_formatobjectnullableStructured output control. Use {"type":"json_object"}
or a JSON schema object when the model supports strict structured output.
temperaturenumbernullableSampling temperature. Most models accept 0 to 2,
but provider-specific limits can differ.
max_tokensintegernullableMaximum generated tokens. The final cap is still bounded by the selected model's context and output limits.
stream_optionsobjectnullableUse {"include_usage":true} when the upstream supports
streaming usage trailers.
Message content
rolestringrequiredMessage role.
systemuserassistanttoolcontentstring | content_part[]requiredPlain text for text-only turns, or an array of content parts for multimodal requests.
tool_callsobject[]nullableAssistant tool calls returned by the model.
tool_call_idstringnullableRequired on tool messages so the model can associate a tool
result with the earlier tool call.
Vision input
For image input, send OpenAI-compatible content parts. Use either HTTPS image URLs or data URLs:
{
"model": "gpt-6-luna",
"messages": [
{
"role": "user",
"content": [
{"type": "text", "text": "What's in this image?"},
{
"type": "image_url",
"image_url": {"url": "data:image/png;base64,..."}
}
]
}
]
}For PDFs and large documents, upload through Files and use
Responses. Chat completions does not automatically
dereference file_id.
Tools
{
"model": "gpt-6-luna",
"messages": [{"role": "user", "content": "Weather in Tokyo?"}],
"tools": [
{
"type": "function",
"function": {
"name": "get_weather",
"parameters": {
"type": "object",
"properties": {
"city": {"type": "string"}
},
"required": ["city"]
}
}
}
]
}Tool support is not universal. Prefer /api/models/catalog and look
at parameters.tools before routing agent traffic.
Response
idstringLazu or provider response ID. Use the response header
X-Lazu-Request-Id for request-level reconciliation.
choicesobject[]Assistant output choices. Streaming responses send incremental
delta objects.
usage.prompt_tokensintegerPrompt input tokens.
usage.completion_tokensintegerGenerated output tokens.
usage.prompt_tokens_details.cached_tokensintegernullableCache read tokens when the upstream reports them.
usage.prompt_tokens_details.cache_write_tokensintegernullableCache creation/write tokens when the upstream reports them.
usage.prompt_tokens_details.cache_miss_tokensintegernullableCache misses when the provider reports them separately, for example DeepSeek-compatible usage.
For complete usage, billing line items, provider raw usage fields and routing metadata, call:
curl https://api.lazu.ai/api/usage/requests/req_lazu_01ABCDEF \
-H "Authorization: Bearer $LAZU_API_KEY"Errors
Relay errors use OpenAI-compatible error envelopes where possible and include a request ID. See Errors for code meanings and retry behavior.
See also
curl https://api.lazu.ai/v1/chat/completions \
-H "Authorization: Bearer $LAZU_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-6-luna",
"messages": [
{"role": "user", "content": "Say hello from Lazu"}
]
}'
{
"id": "chatcmpl-9aZx",
"model": "gpt-6-luna",
"choices": [
{
"index": 0,
"message": { "role": "assistant", "content": "Hello!" },
"finish_reason": "stop"
}
],
"usage": { "prompt_tokens": 11, "completion_tokens": 3 }
}
- Request ID
- req_demo_01
- Model · Lane
- gpt-6-luna · stable
- Tokens · input / output
- 1,200 / 320
- Key
- demo-key
Example only — not a quote and not the result of your request. Send the request above to see your own receipt.
Read a real request receipt →