<!-- Lazu · LLM API pricing comparison: Claude, GPT, Gemini, DeepSeek · https://lazu.ai/blog/llm-api-price-comparison -->

# LLM API pricing comparison: Claude, GPT, Gemini, DeepSeek

2026-09-30 · Lazu team

Short answer: pick by task first, then by price. Hard problems and long agent runs go to Claude Opus or GPT-6 Sol; everyday coding and writing to Claude Sonnet; high volume and low latency to Gemini Flash or DeepSeek. Every price table below reads Lazu's live catalog, so it changes when prices do and is never stale.

## How to read the tables

- Prices are **USD per million tokens**, input and output separately.
- The first row is the maker's **official price**. Below it are Lazu's **stable lane**, which runs on the official API, and the **discount lane** where one exists.
- The discount lane is a reverse-engineered route, not the official API, and we recommend it only inside the matching agent tool: Claude in Claude Code, GPT in Codex, Gemini in Antigravity. For your own code, read the stable-lane price.

## Flagships: hard problems and long tasks


### [claude-opus-5-5](https://lazu.ai/models/anthropic/claude-opus-5-5)

USD per million tokens

| Lane | Input | Output | Cache read | Cache write 5m |
| --- | --- | --- | --- | --- |
| Anthropic official price | $4.00 | $20.00 | $0.20 | $5.00 |
| Lazu Stable · Same as official | $4.00 | $20.00 | $0.20 | $5.00 |
| Lazu Discount · −57% | $1.72 | $8.60 | $0.086 | $2.15 |

A dash means no price is recorded, not free usage. The request receipt is the final billing record.


### [gpt-6-sol](https://lazu.ai/models/openai/gpt-6-sol)

USD per million tokens

| Lane | Input | Output | Cache read | Cache write 5m |
| --- | --- | --- | --- | --- |
| OpenAI official price | $2.00 | $10.00 | $0.20 | $2.50 |
| Input over 272k | $4.00 | $15.00 | $0.40 | $5.00 |
| Lazu Stable · Same as official | $2.00 | $10.00 | $0.20 | $2.50 |
| Lazu Discount · −82% | $0.36 | $1.80 | $0.036 | $0.45 |

A request whose input (cached tokens included) passes the threshold is billed entirely at the long-context prices.

A dash means no price is recorded, not free usage. The request receipt is the final billing record.

Flagships cost the most on output. An agent task can run to hundreds of thousands of tokens, most of it context sent again and again, so count the **cache read price** too: cached input usually costs a small fraction of the full rate.

## Workhorses: everyday coding and writing


### [claude-sonnet-5-5](https://lazu.ai/models/anthropic/claude-sonnet-5-5)

USD per million tokens

| Lane | Input | Output | Cache read | Cache write 5m |
| --- | --- | --- | --- | --- |
| Anthropic official price | $2.00 | $10.00 | $0.20 | $2.50 |
| Lazu Stable · Same as official | $2.00 | $10.00 | $0.20 | $2.50 |
| Lazu Discount · −57% | $0.86 | $4.30 | $0.086 | $1.07 |

A dash means no price is recorded, not free usage. The request receipt is the final billing record.

### [claude-opus-5-5](https://lazu.ai/models/anthropic/claude-opus-5-5)

USD per million tokens

| Lane | Input | Output | Cache read | Cache write 5m |
| --- | --- | --- | --- | --- |
| Anthropic official price | $4.00 | $20.00 | $0.20 | $5.00 |
| Lazu Stable · Same as official | $4.00 | $20.00 | $0.20 | $5.00 |
| Lazu Discount · −57% | $1.72 | $8.60 | $0.086 | $2.15 |

A dash means no price is recorded, not free usage. The request receipt is the final billing record.

Sonnet is enough for a lot of everyday coding, at a noticeably lower price than Opus. The table puts Opus alongside for a direct comparison.

## Value: high volume, low latency


### [gemini-3.8-flash](https://lazu.ai/models/google/gemini-3.8-flash)

USD per million tokens

| Lane | Input | Output | Cache read | Cache write 5m |
| --- | --- | --- | --- | --- |
| Google official price | $0.75 | $3.75 | $0.075 | — |
| Lazu Stable · Same as official | $0.75 | $3.75 | $0.075 | — |
| Lazu Discount · −75% | $0.1875 | $0.9375 | $0.01875 | — |

A dash means no price is recorded, not free usage. The request receipt is the final billing record.


### [deepseek-v4-pro](https://lazu.ai/models/deepseek/deepseek-v4-pro)

USD per million tokens

| Lane | Input | Output | Cache read | Cache write 5m |
| --- | --- | --- | --- | --- |
| DeepSeek official price | $0.66 | $1.98 | $0.00363 | — |
| Lazu Stable | $0.825 | $2.48 | $0.0275 | — |

A dash means no price is recorded, not free usage. The request receipt is the final billing record.


### [deepseek-v4-flash](https://lazu.ai/models/deepseek/deepseek-v4-flash)

USD per million tokens

| Lane | Input | Output | Cache read | Cache write 5m |
| --- | --- | --- | --- | --- |
| DeepSeek official price | $0.15 | $0.60 | $0.003 | — |
| Lazu Stable | $0.1875 | $0.75 | $0.00375 | — |

A dash means no price is recorded, not free usage. The request receipt is the final billing record.

DeepSeek prices by time of day: peak hours cost double, and the model page lists the hours. Run batch jobs off-peak where you can.

## A worked example

Say you make 2,000 requests a day, each with 3,000 input tokens and 500 output tokens:

- that is 6 million input tokens and 1 million output tokens a day;
- the daily cost is 6 × the input price + 1 × the output price;
- plug in the prices from the tables above to get each model's daily cost. Output prices are usually 4–8× input prices, so long outputs widen the gap.

## Ways to spend less

1. **Use caching.** Keep system prompts and long documents at the start and unchanged; the cached part is billed at the cache price.
2. **Match the model to the task.** One key calls every model, so send simple classification and summaries to a Flash-class model.
3. **Put agent tools on the discount lane.** In Claude Code and Codex you can switch the matching models to the discount lane, and keep your own code on the stable lane.
4. **Tune from the bill.** The console's usage log breaks cost down by model; find the most expensive kind of request first.

Every model and price is in the [model catalog](/models), updated live.

