Claude Sonnet 5’s promotional pricing ends August 31: what a 50% token price hike changes for your SME’s AI agents
Anthropic’s own pricing page confirms it: Claude Sonnet 5’s introductory rate expires August 31, 2026, and standard pricing takes effect September 1 — 50% more per token, on the model most production AI agents actually run on. Here is what to check before the bill changes.
When Anthropic launched Claude Sonnet 5 at the end of June 2026, it shipped at an introductory price of $2 per million input tokens and $10 per million output tokens — clearly labelled as time-limited from day one. Anthropic’s own pricing documentation now confirms the date: that rate holds through August 31, 2026, and standard pricing of $3 per million input tokens and $15 per million output tokens takes effect on September 1 — a 50% increase on both input and output, three weeks from today.
Why this is not a niche pricing footnote
Sonnet-tier models are, by Anthropic’s own guidance, the ones recommended for “most production workloads” — Haiku for simple tasks, Opus for the hardest reasoning, Sonnet for everything in between. In practice, that middle tier is where most AI agents running in production actually sit: a customer support assistant, an invoice reconciliation agent, an inbox triage agent. A 50% jump on the default production model is not a rounding error on next month’s invoice — it is the second or third time in 2026 the underlying price of running an agent has moved by a meaningful margin, after Claude Opus 5 landed cheaper than its predecessor in July. Prices are not settling; they are moving in both directions several times a year now, and a budget set once in the spring is already out of date.
What it changes for your SME
How much this costs you in practice depends entirely on your own token volumes, not on the announcement’s headline percentage — the same lesson that applies every time a model’s price moves. It is precisely because that arithmetic is hard to get right that researchers now study it directly: a January 2026 paper by Jonathan Knoop and Hendrik Holtmann, Private LLM Inference on Consumer Blackwell GPUs: A Practical Guide for Cost-Effective Local Deployment in SMEs, benchmarks whether running open-weight models on consumer-grade hardware can undercut cloud API pricing for small and mid-sized businesses — evidence that inference cost has become a line item worth engineering around, not just paying. Whether or not local deployment makes sense for your case, the underlying pressure the study responds to is the same one behind this price change: token costs move, and a business running agents at scale needs to track that movement rather than assume the number it budgeted for still holds.
Whatever agents your business runs — an inbox triage agent, a WhatsApp sales assistant, a regulatory watch agent, an unpaid-invoice follow-up agent — each one calls a model on every single execution, and each of those calls is billed at whichever rate is in effect that day. An agent that ran cheaply in July is not guaranteed to run cheaply in September, and the only way to know is to check your own numbers, not the percentage in the announcement.
Four checks to run before September 1
- →Pull your actual monthly token volume per agent, input and output separately, and multiply by the new September rates to see the real euro impact — not the 50% headline, your own bill.
- →Check whether routine, low-stakes tasks (simple email sorting, basic FAQ answers) can run on a cheaper model tier while the more expensive tier is reserved for tasks where accuracy actually justifies the cost.
- →If prompt caching is not already active on your agents, turn it on — cached tokens are billed at a fraction of the standard input rate and can absorb part of the increase on repeated context.
- →Put a recurring quarterly check on the calendar for the pricing of every AI model your agents depend on — 2026 has already shown that these numbers move more than once a year, in both directions.
None of this requires switching providers or pausing a project. It requires treating model pricing the way you already treat any other recurring supplier cost: reviewed on a schedule, not assumed to be fixed the day you signed up.
Free resource
The self-assessment grid: 20 tasks AI can automate
Sales, admin, support, operations: the 20 tasks AI agents already handle in SMEs — with, for each one, the tell-tale sign that your team is concerned.
Read next
Strategy, costs & ROI
AI and jobs: should your SMB rethink junior salaries?
19 August 2026·5 min read
Strategy, costs & ROI
ChatGPT Business Premium at $125: what OpenAI’s new tier reveals about the real cost of agentic AI for your SMB
12 August 2026·5 min read
Strategy, costs & ROI
AI at work: OpenAI study finds 43% of skilled use crosses job boundaries — what it means for your SMB
29 July 2026·5 min read