# Claude Haiku 5.5 price estimates

These are calculated estimates, not measured bills. No paid Claude requests were made.

## Source and verification
User-provided `official-pricing-source.txt`, dated 2026-10-08, transcribes and paraphrases official sources:
- https://platform.claude.com/docs/en/about-claude/pricing
- https://www.anthropic.com/claude-haiku-5-5

The session could not independently retrieve these pages. Model identity and pricing are used as supplied in that transcription; the results are conditional on its accuracy and applicability.

## Rates (USD per million tokens)
| Input length | Input | Output | 5-minute write | 1-hour write | Read |
|---|---:|---:|---:|---:|---:|
| <=100,000 | 0.10 | 0.50 | 0.125 | 0.20 | 0.01 |
| >100,000 | 0.50 | 2.50 | 0.625 | 1.00 | 0.05 |

## Method
Hold output at O=1,000 tokens per request. N denotes total input tokens. Choose the lower rate row when N<=100,000, otherwise the higher row. Apply the selected row to the entire request, including output; this is not marginal pricing on excess tokens.

Uncached: U=(N*input_rate + O*output_rate)/1,000,000.
First cache write: W=(N*write_rate + O*output_rate)/1,000,000.
Successful repeat hit: H=(N*read_rate + O*output_rate)/1,000,000.
Ten independent uncached requests: 10*U.
Ten cached requests: W+9*H, calculated separately for each write lifetime.

Use independent prompt series at each N. Each cached series begins with an empty cache, then repeats the exact same cacheable prompt nine times. Previous generated outputs are not appended to subsequent prompts. All input is assumed eligible and covered by a cache breakpoint; there is no uncached suffix, extra history, or additional input. All nine later requests are assumed to hit while the cache is valid; each hit refreshes the original lifetime. Output is generated and billed afresh every time. Cache eligibility and hits are assumptions, not observations or guarantees.

Crucially, cached calculations assume the entire input length, including cache reads/writes, selects the tier for all charges, including output. The supplied notes do not independently establish this cache-specific threshold metering detail, so cached scenarios are explicitly conditional. The uncached examples are directly supported by the supplied tier description.

Standard synchronous first-party API pricing only. Excludes batch discounts, taxes, negotiated discounts, regional modifiers, and server-tool charges. Actual document token counts were not measured.

## Findings
At N=100,000, U=$0.0105000; at N=100,001, U=$0.0525005. One additional input token increases the estimate by $0.0420005, or approximately 400.0048%, to 5.0000476 times the earlier cost. Both input and output rates increase fivefold at the threshold.

At N=100,000, ten uncached requests cost $0.1050000; one five-minute write plus nine hits costs $0.0265000 (74.76% less), and one one-hour write plus nine hits costs $0.0340000 (67.62% less).

At N=100,001, those totals are $0.5250050, $0.132501075 and $0.17000145, respectively. Caching lowers absolute cost but preserves the approximate fivefold boundary jump under the assumed tier rule.

The CSV contains six input lengths with per-request and ten-request estimates, rounded to eight decimal places. Decimal arithmetic was used before conversion for plotting/export. The plot separates the lower and upper branches rather than interpolating through the price discontinuity; filled markers indicate 100,000 tokens and open markers 100,001 tokens. Output remains 1,000 tokens in both panels. Plot text was checked for overlap and the saved image was visually inspected.
