---
title: "Responses API (OpenAI-compatible)"
description: "OpenAI-compatible /v1/responses with file_id dereferencing, reasoning controls, request fields and examples."
source: https://lazu.ai/docs/endpoints/responses
updated: 2026-10-02
---

# Responses API

**POST** `/v1/responses`

Use Responses for reasoning models, multimodal inputs and Lazu file\_id dereferencing. This is the recommended endpoint when a request needs uploaded files or newer OpenAI response features.

## Example request

```bash
curl https://api.lazu.ai/v1/responses \
  -H "Authorization: Bearer $LAZU_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-6-luna",
    "input": [{
      "role": "user",
      "content": [
        {"type": "input_text", "text": "Say hi"}
      ]
    }]
  }'
```

## Example response

```json
{
  "id": "resp_01ABCDEF",
  "model": "gpt-6-luna",
  "output": [
    {
      "type": "message",
      "role": "assistant",
      "content": [{ "type": "output_text", "text": "Hi!" }]
    }
  ],
  "usage": { "input_tokens": 9, "output_tokens": 2 }
}
```

## When to use it

- **Use Responses** — Uploaded files, reasoning controls, document workflows, and multimodal
  inputs that should be normalized server-side.
- **Use Chat completions** — Simple chat, SDK compatibility, tool calling, and existing apps already
  built around `/v1/chat/completions`.

## Request body

- `model` `string` (required) — Model ID from `/api/models/catalog`. Prefer models where
  `supported_endpoints` includes the path `/v1/responses`.
- `input` `string | object[]` (required) — Text input or an array of response input messages.
- `instructions` `string` (optional) — System-level instructions for the response.
- `reasoning` `object` (optional) — Reasoning effort controls for supported models, for example
  `{"effort":"medium"}`.
- `tools` `object[]` (optional) — Tool definitions using the OpenAI-compatible Responses shape.
- `stream` `boolean` (optional) — Streams response events when supported by the selected model.
- `max_output_tokens` `integer` (optional) — Upper bound for generated output tokens.

## Content parts

- `input_text` `content part` — Text content sent to the model.
- `input_image` `content part` — Image content. Lazu can dereference uploaded `file_id` values
  with purpose `vision`.
- `input_file` `content part` — Uploaded file reference. Lazu dereferences file content server-side before
  forwarding the request upstream.

## File dereferencing

Upload via [Files](https://lazu.ai/docs/endpoints/files), then reference the resulting

`file_id`. Lazu adds `X-Lazu-File-Dereference: 1` when the
request dereferenced files.

Limits:

- Single file purpose limit still applies.
- Total dereferenced files in one Responses call must stay under 64 MB.
- Chat completions does not auto-dereference `file_id`.

## Stateless compatibility bridges

Some catalog entries expose Responses through a lossless Chat bridge; check
`supported_endpoints[].mode`. On such a route, omitted
`store` is accepted and returns
`X-Lazu-Warning: stateless_bridge`, while
`store: false` is accepted without a warning. Explicit
`store: true`, stateful fields, and unknown cross-protocol fields
return `400 protocol_bridge_unsupported` before an upstream call.
The standard Responses body is not extended with Lazu-only fields.

## Response

- `id` `string` — Response ID.
- `output` `object[]` — Output messages, reasoning items, tool calls, or other response events.
- `usage.input_tokens` `integer` (optional) — Input token count when the upstream reports usage.
- `usage.output_tokens` `integer` (optional) — Output token count when the upstream reports usage.
- `usage.input_tokens_details.cached_tokens` `integer` (optional) — Cache read tokens for providers that expose response-level cache usage.

For full reconciliation, use

`GET /api/usage/requests/{request_id}` with the same API key.

## See also

- [Files API](https://lazu.ai/docs/endpoints/files)
- [Chat completions](https://lazu.ai/docs/endpoints/chat)
- [Model catalog](https://lazu.ai/docs/models/catalog)

## Complete tool round trip

Select a model with Responses and function tools from your token-scoped catalog, then set `LAZU_API_KEY` and `LAZU_MODEL`. Keep every output item when continuing a stateless conversation, including reasoning items; append the tool result with the same `call_id`. Encrypted or provider-specific reasoning state remains native-only when a bridge cannot preserve it.

```python
import json
import os
from openai import OpenAI, APIStatusError

client = OpenAI(
    api_key=os.environ["LAZU_API_KEY"],
    base_url=os.environ.get("LAZU_BASE_URL", "https://api.lazu.ai/v1"),
    max_retries=0,
)
model = os.environ["LAZU_MODEL"]
tools = [{"type": "function", "name": "add", "description": "Add two integers",
          "parameters": {"type": "object", "properties": {
              "a": {"type": "integer"}, "b": {"type": "integer"}},
              "required": ["a", "b"], "additionalProperties": False}}]
history = [{"role": "user", "content": "Use add to calculate 2 + 3."}]
try:
    first = client.responses.create(model=model, input=history, tools=tools,
        tool_choice={"type": "function", "name": "add"}, store=False,
        max_output_tokens=1024)
    history.extend(item.model_dump(exclude_unset=True) for item in first.output)
    for item in first.output:
        if item.type == "function_call":
            if item.name != "add":
                raise ValueError("Unexpected tool")
            args = json.loads(item.arguments)
            history.append({"type": "function_call_output", "call_id": item.call_id,
                            "output": json.dumps({"result": args["a"] + args["b"]})})
    final = client.responses.create(model=model, input=history, tools=tools,
        tool_choice="none", store=False, max_output_tokens=1024)
    print(final.output_text)
except APIStatusError as exc:
    print(exc.status_code, exc.response.headers.get("x-lazu-request-id"), exc.body)
    raise
```

Do not replay automatically after output has arrived. For 413 `request_body_too_large`, reduce the request; for 503 `gateway_overloaded`, retry later with backoff. Local overload does not imply a provider failure. Use request details to reconcile partial usage.
