Paying AI for content: what it actually costs in 2026

We have all gotten used to handing articles to AI. That habit comes with a bill. Here is what a mid-length article actually costs across the major APIs in August 2026 — and the moment that bill stops making sense.

The sticker price

Prices are per 1M tokens in USD, verified August 2026 against provider pricing pages. The table has turned over again in the single month since our July edition — check the provider pages before budgeting a pipeline.

ProviderModelInputOutput
AnthropicClaude Fable 5 (flagship)$10$50
Claude Opus 5$5$25
Claude Sonnet 5 (mid)$3$15
Claude Haiku 4.5 (budget)$1$5
OpenAIGPT-5.5 Pro (flagship)$30$180
GPT-5.6 Sol$5$30
GPT-5.6 Terra (mid)$2$12
GPT-5.6 Luna (budget)$0.20$1.20
GoogleGemini 3.1 Pro$2.00$12
Gemini 3.7 Flash$0.75$3.75
Gemini 3.5 Flash-Lite$0.30$2.50
xAIGrok 4.6$2$6
Grok 4.3$1.25$2.50
DeepSeekV4 Pro (off-peak)$0.66$1.98
V4 Flash (off-peak)$0.22$0.66
AlibabaQwen3.8 Max$2$6
Qwen3.7 Plus$0.32$1.28
Qwen3.7 Flash$0.03$0.13
MoonshotKimi K3$3$15
Kimi K2.5$0.60$3.00
MistralMedium 3.5$1.50$7.50
Large 3$0.50$1.50
Small 4$0.15$0.60

Mistral is the one block that did not move this month — every other provider’s prices above are new since July, and Moonshot’s Kimi joins the table for the first time (its flagship K3 lands at exactly Claude Sonnet 5’s price). Output runs 2–6× more expensive than input across every tier. That matters: a long article is mostly output tokens. A short brief with a long system prompt is mostly input. More footnotes than ever ride on the headline numbers — Claude Sonnet 5’s introductory pricing ($2/$10) ends 2026-08-31; Gemini 3.7 Flash is introductory too ($1.50/$7.50 from January 2027); Gemini 3.1 Pro and Grok 4.6 roughly double above a 200K-token context ($4/$18 and $4/$12); DeepSeek is priced by the clock — the table shows off-peak, and peak hours (01:00–04:00 and 06:00–10:00 UTC) double it; and Qwen3.7 Flash is tiered by request length ($0.03/$0.13 under 32K tokens, up to $0.20/$0.80).

What one article actually costs

A mid-length article — roughly 2,500 words — lands at about 3,000 input tokens (system prompt, brief, template guidance, source material) and 3,500 output tokens (~2,500 words). Using those numbers:

ModelList per articleWith cache (read discount)With cache + batch
GPT-5.5 Pro~$0.72n/a~$0.36
Claude Fable 5~$0.21~$0.18~$0.089
GPT-5.6 Sol~$0.12~$0.11~$0.053
Claude Opus 5~$0.10~$0.089~$0.045
Claude Sonnet 5~$0.062~$0.053~$0.027
Kimi K3~$0.062~$0.053n/a
GPT-5.6 Terra~$0.048~$0.043~$0.021
Gemini 3.1 Pro~$0.048~$0.043~$0.021
Mistral Medium 3.5~$0.031~$0.027~$0.013
Qwen3.8 Max~$0.027~$0.024~$0.014 (batch only)
Grok 4.6~$0.027~$0.023~$0.018 (batch −20%)
Claude Haiku 4.5~$0.021~$0.018~$0.009
Gemini 3.7 Flash~$0.015~$0.013~$0.0067
Kimi K2.5~$0.012n/a~$0.0074 (batch −40%)
Gemini 3.5 Flash-Lite~$0.0097~$0.0088~$0.0044
DeepSeek V4 Pro (off-peak)~$0.0089~$0.0070n/a
Mistral Large 3~$0.0068~$0.0054~$0.0027
GPT-5.6 Luna~$0.0048~$0.0043~$0.0021
DeepSeek V4 Flash (off-peak)~$0.003~$0.0023n/a
Mistral Small 4~$0.0026~$0.0021~$0.0011
Qwen3.7 Flash~$0.00055n/a~$0.00027

A budget tier like Luna, Gemini Flash-Lite, or Haiku brings an article to a cent or two at list. Qwen3.7 Flash lands near a twentieth of a cent — the new floor now that DeepSeek’s repricing took it out of that seat. A true flagship like GPT-5.5 Pro is roughly 1,300× the cost of Qwen Flash and 12–15× a mid-tier like Sonnet 5 or GPT-5.6 Terra.

The discounts that actually apply

Two mechanisms matter for anyone running AI content at volume — and both got more conditional again this month:

  • Prompt caching. Anthropic, OpenAI, Google, xAI, DeepSeek, Alibaba, Moonshot, and Mistral all discount cached input reads — typically 90% off (Anthropic, OpenAI, Google, Mistral, Moonshot), 75% off (xAI), and around 97% off at DeepSeek, whose cache-hit price is still the biggest single lever on this page even after the August repricing. The catch is on the other side: cache writes carry a surcharge (Anthropic bills 1.25× for a 5-minute TTL and 2× for an hour; OpenAI added a 1.25× write charge on the GPT-5.6 family), and Google bills hourly storage. Caching only pays off once you actually re-read the cached prefix a few times.
  • Batch API. Anthropic, OpenAI, Google, Alibaba, and Mistral offer 50% off for async jobs completing within 24 hours. It is not a universal 50%: xAI discounts by only 20%, Moonshot bills batch at 60% of list (−40%, and does not list its flagship K3 on the batch page at all), Alibaba does not let batch and cache combine on the same request, and DeepSeek’s discount is the clock, not a queue — the August 16 repricing brought back peak/off-peak billing, with off-peak (most of the day) at half the peak rate.

The two stack where both exist. Batch + cache gets you roughly 95% off list on input and 50% off on output. If you are generating content in scheduled batches with a fixed system prompt, you should be using both.

The math at scale

Pick a reasonable mid-tier model — Claude Sonnet 5 at list, $0.062 per article:

VolumePure-AI (regenerate each time)AI template once + local renders
1 article$0.06$0.06
100 articles~$6.15~$0.06
1,000 articles~$61.50~$0.06
10,000 articles~$615~$0.06

Against a flagship like GPT-5.5 Pro the gap gets wider: 10,000 articles costs about $7,200 at list, versus the cost of generating one good template (well under $1) and running it through a local renderer for free. That is the moment the bill stops making sense.

The lineup churns faster than your content plan

There is a second cost here that no pricing page lists. Every model named in our April 2026 edition of this article — GPT-5.4, Gemini 3 Flash, Grok 4, DeepSeek V3 and R1, Qwen3 Max, Mistral Small 3.1, Claude Opus 4.7 and Sonnet 4.6 — has since been renamed, superseded, or deprecated. Three months. OpenAI dropped numeric tiers for named ones (Sol / Terra / Luna). DeepSeek retired the deepseek-chat and deepseek-reasoner aliases outright.

The July edition survived barely a month. Since then: Anthropic shipped Claude Opus 5 (2026-07-24) and the Opus the table listed became last generation; OpenAI cut GPT-5.6 Terra by 20% and Luna by 80% (2026-07-30); Google launched Gemini 3.6 Flash and then 3.7 Flash five weeks apart; xAI shipped Grok 4.6 (2026-08-12); Alibaba replaced Qwen3.7 Max with Qwen3.8 Max (2026-08-03); and DeepSeek raised prices by up to 1,100% at peak (2026-08-16), reintroducing the time-of-day billing it had dropped with V4. One month; six of the seven providers changed their price list.

If your content pipeline calls a model on every render, every one of those changes is a migration: new model IDs, re-tuned prompts, re-baselined costs, and output that reads differently than last quarter’s. If your pipeline calls a model once to author a template and then renders locally, none of it touches you. Model churn is an argument for the template, not just an inconvenience.

When you should pay per article

Not every article is a repeat. Paying the API bill makes sense for:

  • one-off editorial pieces with a unique voice or angle;
  • new topics you have never covered, where the model does actual research;
  • long-form features where nuance matters more than volume;
  • drafts you will heavily edit by hand anyway.

For this kind of work you want the best model you can afford. Flagship-tier output pays for itself in reduced editing time.

When you should pay once and render many

The cost math flips the moment you have to produce the same shape of article more than a few times:

  • product pages across a catalog of SKUs;
  • localized variants across languages and regions;
  • SEO landing pages for a long keyword list;
  • multi-tenant SaaS marketing sites that share structure;
  • anything where the structure is constant and the facts change.

Here, paying per render is pure waste. Write the template once with AI, then hand rendering to a tool that does it for free.

The hybrid workflow

This is the workflow Spintax is built for:

  1. Pay the flagship once. Use the best model you can afford — Fable 5, Opus 5, GPT-5.6 Sol, Gemini 3.1 Pro — to author the template. You pay $0.05–$1 and you get a high-quality, grammar-safe, multi-variant template.
  2. Render forever, locally. Spintax resolves the template on your CPU. Thousands of variants cost effectively zero.
  3. Update when meaning changes. Regenerate the template only when the facts or positioning change — not when the vendor renames a model. For a stable product, that might be quarterly.

You pay for quality where it matters. You stop paying for quantity.

Caveats

  • Prices change, and so do names. The entire table above turned over between April and July 2026 — and then again between July and August. Recheck the provider pages before budgeting, and never hardcode a model ID you have not re-verified this quarter.
  • Cache writes are not free. The headline “90% off” applies to cache reads. Anthropic and OpenAI both bill a premium on the write, and Google bills hourly storage. A cache that gets read once is more expensive than no cache.
  • Batch is not universally 50%. xAI is 20%. Moonshot is 40% and excludes its flagship. Alibaba refuses to stack batch with cache. DeepSeek’s discount is the time of day, not a batch queue. Do not assume the stack applies before you check.
  • Long context costs extra on some providers. Gemini Pro tiers and Grok 4.6 roughly double past 200K tokens. Anthropic and OpenAI charge flat rates across their context windows.
  • Quality is not the same at every tier. Budget tiers are fine for scaffolding and rewrites. Nuanced voice, long-context fidelity, and factual grounding still favour flagship. Use cheap models for local renders, not for the template itself.

Where to go next

Ready to turn one good AI generation into thousands of renders? Start with the authoring series:

Data sources: Anthropic, OpenAI, Google, xAI, DeepSeek, Alibaba, Moonshot, Mistral pricing pages — verified 2026-08-20. Per-article calculations assume 3,000 input tokens + 3,500 output tokens per article.