On 7 October 2026 Anthropic shipped Claude Haiku 5.5 and priced prompts up to 100,000 tokens at $0.10 per million input tokens and $0.50 per million output tokens, 90 percent below Haiku 4.5's $1 and $5. Prompts over that threshold are $0.50 and $2.50, a 50 percent cut. Cache reads on the short band are $0.01 per million. The same day OpenAI started rolling GPT-6 Sol to paid ChatGPT plans and scheduled GPT-6 Luna, listed at the same $0.10 / $0.50 API band, for free and Go users on 8 October. Anthropic also halved Sonnet 5.5 cache-read prices, which it says makes typical agent work about 20 percent cheaper.
For an indie SaaS already routing classification, extraction, and drafts through a small model, that is a margin event. It shows up on next month's invoice if you do nothing, and as a support incident if you switch everything at once. Anthropic says Haiku 5.5 beat GPT-6 Luna on internal comparisons, including computer use. Treat that as a vendor claim until you run your own set.
Most bootstrapped products picked a model in 2025 and left the router alone. A support tool might send every ticket summary to a Sonnet-class model because the first eval looked fine. At Haiku 4.5 prices that habit was wasteful. At Haiku 5.5 prices it leaves most of the small-model budget on the table for the 90 percent of requests Anthropic says historically stayed under 100,000 tokens.
A meeting-notes product that processes 400,000 short jobs a month, each 1,200 input tokens and 280 output tokens, uncached and under the threshold, spends about $480 of input and $560 of output on Haiku 4.5, roughly $1,040. On Haiku 5.5 list price the same counts are about $48 and $56, roughly $104. GPT-6 Luna's published list price lands in the same neighborhood. The gap is real if those jobs are already in production. It is not free money if the updated tokenizer uses more tokens per task, which Anthropic says it does, or if quality drops on jobs that mention a refund.
Cross 100,000 tokens and Haiku 5.5 input jumps fivefold, from $0.10 to $0.50. Output jumps from $0.50 to $2.50. A founder who dumps a full customer workspace into every call will not see a 90 percent cut.
Split the catalog by failure cost, not by brand.
Safe to trial on Haiku 5.5 or Luna this week: language detection, ticket tagging, spam scoring, invoice field extraction with a schema, subject-line variants, and routing decisions a human still confirms. These jobs are repetitive, scored easily, and cheap to roll back.
Leave on Sonnet 5.5, Opus 5.5, or GPT-6 Sol until you have a shadow run: customer-visible legal text, refund calculations, production code edits without a test gate, and computer-use loops where a wrong click spends money. A vendor computer-use score is not a reason to point a billing agent at the cheapest model on day one.
A support desk can move tag and draft. The send button stays on the stronger model, or on a human, until draft acceptance holds. A bookkeeping micro-SaaS can move merchant normalization. It should not move the deductible answer customers paste into a filing.
Freeze a seven-day sample before you change a route. Export call id, model, input tokens, output tokens, cache-read tokens, latency, customer id, and the downstream label you already store: tag accepted, field corrected, draft edited, ticket reopened. If you lack the label, add it for one workflow only.
Price the sample on both cards. Use list prices, then apply your real cache rate. A 2,000-token system prompt read from cache at $0.01 per million is a rounding error now. It was not at $0.10. Run the same spreadsheet for GPT-6 Luna. The cheaper brand name is not automatically the cheaper route after cache and tokenizer differences.
Shadow 500 production calls. Send the real prompt to the incumbent and to Haiku 5.5. Store both outputs. Do not show the new output to customers. Score exact-match fields automatically. Score drafts with the rubric a support lead already uses: factual, on-policy, no invented refund. Reject the route if the shadow set fails more than the current model on any class that reaches a customer.
Cap the prompt before you celebrate the cut. If p95 input tokens on the route is over 40,000, you are one workspace dump away from the $0.50 band. Move static instructions into a cached prefix. A classifier needs the latest message and the account plan, not the last 40 emails.
Shift 10 percent of traffic, not 100. Route by a stable hash of account id so one customer does not flap between models. Watch reopen rate, edit rate, and refund tickets for 72 hours. On 400,000 jobs, 10 percent is 40,000 calls: enough to see a 2-point quality drop, small enough to cover with goodwill credits if you were wrong.
Rewrite unit cost on the pricing page only after the slice holds. If a Pro plan includes 2,000 AI actions and cost per action fell from about $0.0026 to about $0.00026, you have room. Passing the whole cut through teaches customers the next cut will be passed through too. Better uses: raise the included action count, add a workflow you could not afford, or hold price and widen margin. Opus Clip's October move, lifting Pro annual from about $14.50 to $19.43 a month while changing the credit shape, shows packaging can move the other way when usage is the constraint.
Make rollback a config flag. Model id, prompt version, and traffic percent should be one change, not a deploy. Haiku 5.5's model id is claude-haiku-5-5. Pin it. "Latest Haiku" is how a quiet quality shift becomes your incident.
Recheck batch discounts before you lock the budget. Anthropic's batch API is 50 percent off input and output. A nightly enrichment job that does not need a two-second response should not run on the synchronous rate. Interactive chat should.
InboxSort, a fictional but ordinary product, charges $29 a month, has 1,800 paying accounts, and classifies 40 threads per account per month. That is 72,000 jobs of 900 input tokens and 120 output tokens.
On Haiku 4.5 the model line is about $65 of input and $43 of output, roughly $108, before cache. On Haiku 5.5 short-prompt rates it is about $6.50 and $4.30, roughly $11. Payroll and Stripe fees still dominate. The model line stops being the reason you refuse a free trial of the classifier.
InboxSort's failure mode is a newsletter tagged as a customer. Shadow the new model on the last 500 mis-tags and the last 500 correct tags. If precision on the customer label drops, keep the incumbent for that label and move only newsletter and receipt. Put GPT-6 Luna in the same shadow. If Luna wins on tags and Haiku 5.5 wins on latency, split by workflow.
The 7 October prices do not fix a prompt that invents policy, a missing eval, or a computer-use loop that can spend a customer's money. They do not obligate you to pass the savings through. Anthropic's Claude for Startups expansion, announced 6 October, offers eligible young companies a year of Team seats and $1,000 in API credits. That is runway. Credits hide a bad route until the credit ends.
Free users getting GPT-6 Luna inside ChatGPT on 8 October will compare your assistant with a chat tab that now answers with charts and controls. Matching that app feature for feature is a trap. Matching it on the jobs you charge for, a correct tag, a grounded draft, a refund that follows your policy, is the work.
Does Haiku 5.5 cost 90 percent less on every call?
No. The 90 percent cut versus Haiku 4.5 applies to prompts up to 100,000 tokens. Over that, input is $0.50 and output is $2.50 per million tokens. Tokenizer changes can also raise tokens per task.
Is GPT-6 Luna the same price?
OpenAI's API list price for Luna is $0.10 input and $0.50 output per million tokens. Cache rules, batch discounts, and tokens per task still differ. Compare on your prompts.
Should agent loops move on day one?
Not customer-visible ones. Shadow them. Anthropic's computer-use comparison is Anthropic's evaluation. Your click-path is the one that can spend money.
Do I have to lower my SaaS price?
No. Recompute cost per included action, then add usage, hold price, or discount. A model cut is a supply change. Your price is a demand decision.
What is the minimum log line before a switch?
Call id, model id, input tokens, output tokens, cached tokens, latency, and one downstream label. Without the label you can see a cheaper bill and miss a worse product.