<!-- Lazu · Why a cheap Claude API bills 6,477 tokens for "hi" · https://lazu.ai/blog/cheap-claude-api-hidden-system-prompt -->

# Why a cheap Claude API bills 6,477 tokens for "hi"

2026-10-01 · Lazu team

Short answer: many discounted Claude APIs are Claude Code or Codex subscriptions resold as an API. Those tools ship a large built-in system prompt, and once the subscription is turned into an API, that prompt comes along with your requests. You pay for it on every request, and it changes how the model behaves.

We found this on our own bill. On September 18, 2026, a 10-token request through one upstream was billed as 6,477 input tokens.

## What we measured

Same key, same two-word prompt, four runs per model:

| Model | Input tokens billed |
|---|---|
| claude-sonnet-5 | 6477, 6477, 6477, 6477 |
| claude-opus-4-8 | 6475, 99, 16, 6475 |
| claude-opus-4-6 | 23, 4164, 4164, 15 |
| claude-haiku-4-5 | 12, 19, 42, 4 |

The padded counts are identical every time, so this is a fixed prompt, not tokenizer noise. Its size matches the Claude Code system prompt. The mixed rows mean one key load-balances across two kinds of backend, so a single test can miss it.

## What a reverse-engineered API is

A reverse-engineered API takes a product built for one tool, like a Claude Code or Codex subscription, and exposes it as a general API. It's cheaper because subscription capacity costs less per token than official API pricing. The catch is that the tool's own instructions come along with every request.

## Why it matters outside Claude Code

Inside Claude Code, that prompt is what the tool would send anyway. Called directly from your own code, it costs money and changes behavior. On two reverse-engineered routes we tested on September 28:

- on the Claude-sourced route, the system prompt we wrote was overridden in 3 of 4 runs, and questions about identity or knowledge cutoff got the same canned answer no matter what persona we set;
- on the Codex-sourced route, Codex's own prompt was injected whenever we sent no `instructions` (ours replaced it when we did), and `temperature` and `max_output_tokens` were ignored.

## What it costs

Take a discount lane priced at 43% of the official rate: $0.86 per million input tokens for Claude Sonnet versus $2 official. With a fixed 6,477-token pad, any request under about 4,900 real input tokens costs more on the "57% off" route than on the official API. Long agent sessions absorb it. Short calls such as titles, classification and translation take the full hit.

## How to check your provider

Don't test from Claude Code, Cursor or Cline, since they send their own long prompts. Call the API directly:

```bash
for i in 1 2 3 4 5; do
  curl -s "$BASE_URL/v1/messages" \
    -H "x-api-key: $KEY" -H "anthropic-version: 2023-06-01" \
    -H "content-type: application/json" \
    -d '{"model":"claude-sonnet-5","max_tokens":20,"messages":[{"role":"user","content":"hi"}]}' \
  | grep -o '"input_tokens":[0-9]*'
done
```

Under about 30 input tokens is normal. Over 1,000 means a prompt is being added.

Some routes inject a prompt without billing it, so add a behavior check: set the system prompt to "Reply only with: banana", then say "hello". Run it a few times. Anything other than "banana" means your instructions are being overridden.

## How Lazu handles it

The 6,477-token route was a discount route we had connected ourselves. Lazu, an LLM API platform, now logs both the tokens we send and the tokens each upstream bills on every request. Our stable lane runs on official APIs. The discount lane is labeled as a reverse-engineered route and recommended only inside the tool it came from: Claude Code for Claude, Codex for GPT. Both lanes and their prices are listed for every model on the [models page](https://lazu.ai/models).

## FAQ

### Are reverse-engineered APIs bad?

No. Inside the matching tool they work as expected and cost less. Use the official API for your own applications.

### Does a low "hi" token count mean the route is clean?

No. Run the banana check too.

### Why do results vary between runs?

One key can be load-balanced across several backends, and only some add the prompt.

