Mastodon Skip to content
LIVE - NYSE/-/- CRYPTO/OPEN/24/7
BTC$82,897▼ 1.34%ETH$2,565▼ 1.96%SOL$114.96▼ 3.05%TOTAL CRYPTO$2.83T▼ 3.42%S&P 5007,801.77▼ 0.22%NASDAQ27,538.69▼ 0.22%DOW51,179.87▼ 0.66%GOLD4,153.00▲ 0.30%WTI91.51▲ 3.66%BRENT104.10▲ 3.89%EUR/USD1.1198▼ 0.49%USD/JPY158.19▼ 0.07%DXY102.28▲ 0.03%
AI

Anthropic Launches Claude Haiku 5.5, Its Cheapest Model Yet

Claude Haiku 5.5 costs about 75 percent less to run than its predecessor and targets summaries, database queries and live customer support work at $0.10 per million input tokens.

Pexels – igovar igovar

Anthropic released Claude Haiku 5.5 on Wednesday, calling it the cheapest, fastest and most capable small model the company has made. The release, reported by Reuters and Decrypt, is the third model in the Claude 5.5 family to ship within a month and lands while Anthropic prepares for a widely expected initial public offering.

Pricing is the story most developers will read first. Builders plugging Claude into their own products pay $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens, which Anthropic says works out to about 75 percent below what Haiku’s predecessor cost to run. Tokens are the chunks of text that models actually read, so a page of prose runs a few thousand of them; at these rates, processing an entire novel’s worth of input costs well under a dollar.

The company positions the model for high-volume chores: summarizing documents, querying databases, and speed-sensitive work such as live customer support or operating a web browser on the user’s behalf. Those are the jobs where latency and cost dominate, where waiting two seconds for a reply is the difference between a usable product and an abandoned cart. Anthropic’s blurb for the model put it plainly: it targets “high-volume jobs like summaries and live customer support.”

A crowded release calendar

Wednesday’s launch follows Opus 5.5, released 15 days earlier, and completes a family cadence Anthropic has accelerated dramatically this year. Where model families once arrived a season apart, the 5.5 series landed as three separate releases inside 30 days, each aimed at a different tier of the market: large models for frontier work, mid-tier models for agents and coding, and now a small model that prices out commodity expenses. The staggered rollout lets Anthropic hold each tier’s news cycle separately instead of bundling all three into one keynote-style event.

Competitors will read the move as pricing pressure. OpenAI’s GPT-6 rollout began shipping interactive interfaces to paid ChatGPT tiers on Tuesday, with free and entry tiers following on October 8. Google has pushed Gemini model updates on a similar rhythm. The generative AI leaders are no longer racing only on capability benchmarks; they are racing on unit economics, because the buyers that matter, large enterprises running millions of support tickets and internal documents through models monthly, choose on price per resolved task as much as on leaderboard position.

Claude 5.5 family releases, September-October 2026 Date / figure
Opus 5.5 released Late September 2026
Sonnet 5.5 released Early October 2026
Haiku 5.5 released October 7, 2026
Haiku 5.5 input price $0.10 per million tokens
Haiku 5.5 output price $0.50 per million tokens
Cost vs predecessor About 75% lower
Max prompt length at listed price 100,000 tokens

Why the small-model market matters

Small models are where the volume is. Frontier models get the headlines, but the majority of tokens consumed across the industry get spent on summarization, extraction, routing and classification, the unglamorous workloads that run continuously in the background of every large site and enterprise. A model that does those jobs at one quarter of the previous cost changes the arithmetic for entire categories of products that were previously unprofitable to ship, from document ingestion pipelines to always-on browser agents.

Anthropic has commercial reasons to nail this tier. The company reached a $350 billion private valuation after Microsoft and Nvidia committed a combined $15 billion in November 2025, a round layered on top of roughly $30 billion of committed Microsoft cloud purchases. Those commitments carry expectations of revenue, and revenue at Anthropic’s scale comes from API usage, not primarily from subscription seats. A cheap, fast Haiku is designed to grow that usage line specifically.

There is a safety angle too. Anthropic has built its brand on caution, and small models deployed into customer-facing roles carry their own risk profile: they answer more queries per second than any human reviewer can spot-check, which makes evals and automated graders more important, not less. The company has published safety cases with each Claude release, and the Haiku tier gets its own capability testing rather than inheriting conclusions drawn from its larger siblings.

The IPO backdrop

Every release this quarter lands against the filing. Anthropic submitted IPO paperwork in June 2026 after confirming SEC review, positioning itself to become one of the first frontier labs to reach public markets, ahead of OpenAI. Analysts following the file noted that quarterly revenue growth and a defensible gross margin depend heavily on the API business, the exact segment a 75 percent price cut targets. Reuters framed Wednesday’s launch the same way: as lineup expansion before a planned listing.

Wall Street will read Haiku’s price cut two ways. Lower prices shrink margin per token, which can look bad in an S-1. But the sector’s actual dynamic has been price elasticity: each major drop across the industry so far has expanded consumption faster than it compressed revenue, a pattern OpenAI, Google and Anthropic have all cited in prior earnings discussions. Whether Haiku 5.5 keeps that streak is the number that will matter most in its first full quarter of availability.

Developer reaction is the other early signal. The model’s speed numbers, not its benchmark scores, are what teams shipping agents will test first, since an agent that calls a small model dozens of times per task feels the price and latency of every call. Given how quickly these model families keep arriving, the cost comparison spreadsheet that teams priced months ago may need a refresh before the quarter ends.

SourcesReuters (October 7, 2026); Decrypt; Anthropic release notes; TechTarget weekly roundup (June 2026) for IPO filing details.
Share: X