# Paying AI for content: what it actually costs in 2026 We have all gotten used to handing articles to AI. That habit comes with a bill. Here is what a mid-length article actually costs across the major APIs in August 2026 — and the moment that bill stops making sense. ## The sticker price Prices are per 1M tokens in USD, verified August 2026 against provider pricing pages. The table has turned over again in the single month since our July edition — check the provider pages before budgeting a pipeline. | Provider | Model | Input | Output | | --- | --- | --- | --- | | Anthropic | Claude Fable 5 (flagship) | $10 | $50 | | Claude Opus 5 | $5 | $25 | | Claude Sonnet 5 (mid) | $3 | $15 | | Claude Haiku 4.5 (budget) | $1 | $5 | | OpenAI | GPT-5.5 Pro (flagship) | $30 | $180 | | GPT-5.6 Sol | $5 | $30 | | GPT-5.6 Terra (mid) | $2 | $12 | | GPT-5.6 Luna (budget) | $0.20 | $1.20 | | Google | Gemini 3.1 Pro | $2.00 | $12 | | Gemini 3.7 Flash | $0.75 | $3.75 | | Gemini 3.5 Flash-Lite | $0.30 | $2.50 | | xAI | Grok 4.6 | $2 | $6 | | Grok 4.3 | $1.25 | $2.50 | | DeepSeek | V4 Pro (off-peak) | $0.66 | $1.98 | | V4 Flash (off-peak) | $0.22 | $0.66 | | Alibaba | Qwen3.8 Max | $2 | $6 | | Qwen3.7 Plus | $0.32 | $1.28 | | Qwen3.7 Flash | $0.03 | $0.13 | | Moonshot | Kimi K3 | $3 | $15 | | Kimi K2.5 | $0.60 | $3.00 | | Mistral | Medium 3.5 | $1.50 | $7.50 | | Large 3 | $0.50 | $1.50 | | Small 4 | $0.15 | $0.60 | Mistral is the one block that did not move this month — every other provider’s prices above are new since July, and Moonshot’s Kimi joins the table for the first time (its flagship K3 lands at exactly Claude Sonnet 5’s price). Output runs 2–6× more expensive than input across every tier. That matters: a long article is mostly output tokens. A short brief with a long system prompt is mostly input. More footnotes than ever ride on the headline numbers — Claude Sonnet 5’s introductory pricing ($2/$10) ends 2026-08-31; Gemini 3.7 Flash is introductory too ($1.50/$7.50 from January 2027); Gemini 3.1 Pro and Grok 4.6 roughly double above a 200K-token context ($4/$18 and $4/$12); DeepSeek is priced by the clock — the table shows off-peak, and peak hours (01:00–04:00 and 06:00–10:00 UTC) double it; and Qwen3.7 Flash is tiered by request length ($0.03/$0.13 under 32K tokens, up to $0.20/$0.80). ## What one article actually costs A mid-length article — roughly 2,500 words — lands at about **3,000 input tokens** (system prompt, brief, template guidance, source material) and **3,500 output tokens** (~2,500 words). Using those numbers: | Model | List per article | With cache (read discount) | With cache + batch | | --- | --- | --- | --- | | GPT-5.5 Pro | ~$0.72 | n/a | ~$0.36 | | Claude Fable 5 | ~$0.21 | ~$0.18 | ~$0.089 | | GPT-5.6 Sol | ~$0.12 | ~$0.11 | ~$0.053 | | Claude Opus 5 | ~$0.10 | ~$0.089 | ~$0.045 | | Claude Sonnet 5 | ~$0.062 | ~$0.053 | ~$0.027 | | Kimi K3 | ~$0.062 | ~$0.053 | n/a | | GPT-5.6 Terra | ~$0.048 | ~$0.043 | ~$0.021 | | Gemini 3.1 Pro | ~$0.048 | ~$0.043 | ~$0.021 | | Mistral Medium 3.5 | ~$0.031 | ~$0.027 | ~$0.013 | | Qwen3.8 Max | ~$0.027 | ~$0.024 | ~$0.014 (batch only) | | Grok 4.6 | ~$0.027 | ~$0.023 | ~$0.018 (batch −20%) | | Claude Haiku 4.5 | ~$0.021 | ~$0.018 | ~$0.009 | | Gemini 3.7 Flash | ~$0.015 | ~$0.013 | ~$0.0067 | | Kimi K2.5 | ~$0.012 | n/a | ~$0.0074 (batch −40%) | | Gemini 3.5 Flash-Lite | ~$0.0097 | ~$0.0088 | ~$0.0044 | | DeepSeek V4 Pro (off-peak) | ~$0.0089 | ~$0.0070 | n/a | | Mistral Large 3 | ~$0.0068 | ~$0.0054 | ~$0.0027 | | GPT-5.6 Luna | ~$0.0048 | ~$0.0043 | ~$0.0021 | | DeepSeek V4 Flash (off-peak) | ~$0.003 | ~$0.0023 | n/a | | Mistral Small 4 | ~$0.0026 | ~$0.0021 | ~$0.0011 | | Qwen3.7 Flash | ~$0.00055 | n/a | ~$0.00027 | A budget tier like Luna, Gemini Flash-Lite, or Haiku brings an article to a cent or two at list. Qwen3.7 Flash lands near a twentieth of a cent — the new floor now that DeepSeek’s repricing took it out of that seat. A true flagship like GPT-5.5 Pro is roughly **1,300×** the cost of Qwen Flash and 12–15× a mid-tier like Sonnet 5 or GPT-5.6 Terra. ## The discounts that actually apply Two mechanisms matter for anyone running AI content at volume — and both got more conditional again this month: - **Prompt caching.** Anthropic, OpenAI, Google, xAI, DeepSeek, Alibaba, Moonshot, and Mistral all discount cached input _reads_ — typically **90%** off (Anthropic, OpenAI, Google, Mistral, Moonshot), **75%** off (xAI), and around **97%** off at DeepSeek, whose cache-hit price is still the biggest single lever on this page even after the August repricing. The catch is on the other side: cache _writes_ carry a surcharge (Anthropic bills 1.25× for a 5-minute TTL and 2× for an hour; OpenAI added a 1.25× write charge on the GPT-5.6 family), and Google bills hourly storage. Caching only pays off once you actually re-read the cached prefix a few times. - **Batch API.** Anthropic, OpenAI, Google, Alibaba, and Mistral offer **50% off** for async jobs completing within 24 hours. It is not a universal 50%: **xAI discounts by only 20%**, **Moonshot bills batch at 60% of list** (−40%, and does not list its flagship K3 on the batch page at all), **Alibaba does not let batch and cache combine** on the same request, and **DeepSeek’s discount is the clock, not a queue** — the August 16 repricing brought back peak/off-peak billing, with off-peak (most of the day) at half the peak rate. The two stack where both exist. Batch + cache gets you roughly 95% off list on input and 50% off on output. If you are generating content in scheduled batches with a fixed system prompt, _you should be using both_. **Google’s cache is not free to hold.** Gemini context caching charges an hourly storage fee on top of the token discount — $0.50 per 1M tokens per hour on 3.7 Flash (introductory; $1.00 from January 2027), $4.50 on 3.1 Pro. For a short-lived job the storage fee can exceed the token savings. Cache when the same prefix is hit repeatedly within the hour; do not cache speculatively. If you are calling the API _synchronously_ for one article at a time, with no prompt cache and no batch, you are paying full list. That is fine for a handful of articles a week. At higher volumes it is just waste. ## The math at scale Pick a reasonable mid-tier model — Claude Sonnet 5 at list, $0.062 per article: | Volume | Pure-AI (regenerate each time) | AI template once + local renders | | --- | --- | --- | | 1 article | $0.06 | $0.06 | | 100 articles | ~$6.15 | ~$0.06 | | 1,000 articles | ~$61.50 | ~$0.06 | | 10,000 articles | ~$615 | ~$0.06 | Against a flagship like GPT-5.5 Pro the gap gets wider: 10,000 articles costs about $7,200 at list, versus the cost of generating one good template (well under $1) and running it through a local renderer for free. That is the moment the bill stops making sense. ## The lineup churns faster than your content plan There is a second cost here that no pricing page lists. Every model named in our April 2026 edition of this article — GPT-5.4, Gemini 3 Flash, Grok 4, DeepSeek V3 and R1, Qwen3 Max, Mistral Small 3.1, Claude Opus 4.7 and Sonnet 4.6 — has since been renamed, superseded, or deprecated. Three months. OpenAI dropped numeric tiers for named ones (Sol / Terra / Luna). DeepSeek retired the `deepseek-chat` and `deepseek-reasoner` aliases outright. The July edition survived barely a month. Since then: Anthropic shipped Claude Opus 5 (2026-07-24) and the Opus the table listed became last generation; OpenAI cut GPT-5.6 Terra by 20% and Luna by 80% (2026-07-30); Google launched Gemini 3.6 Flash and then 3.7 Flash five weeks apart; xAI shipped Grok 4.6 (2026-08-12); Alibaba replaced Qwen3.7 Max with Qwen3.8 Max (2026-08-03); and DeepSeek raised prices by up to 1,100% at peak (2026-08-16), reintroducing the time-of-day billing it had dropped with V4. One month; six of the seven providers changed their price list. If your content pipeline calls a model on every render, every one of those changes is a migration: new model IDs, re-tuned prompts, re-baselined costs, and output that reads differently than last quarter’s. If your pipeline calls a model _once_ to author a template and then renders locally, none of it touches you. Model churn is an argument for the template, not just an inconvenience. ## When you should pay per article Not every article is a repeat. Paying the API bill makes sense for: - one-off editorial pieces with a unique voice or angle; - new topics you have never covered, where the model does actual research; - long-form features where nuance matters more than volume; - drafts you will heavily edit by hand anyway. For this kind of work you want the _best_ model you can afford. Flagship-tier output pays for itself in reduced editing time. ## When you should pay once and render many The cost math flips the moment you have to produce _the same shape of article_ more than a few times: - product pages across a catalog of SKUs; - localized variants across languages and regions; - SEO landing pages for a long keyword list; - multi-tenant SaaS marketing sites that share structure; - anything where the structure is constant and the facts change. Here, paying per render is pure waste. Write the template once with AI, then hand rendering to a tool that does it for free. ## The hybrid workflow This is the workflow Spintax is built for: 1. **Pay the flagship once.** Use the best model you can afford — Fable 5, Opus 5, GPT-5.6 Sol, Gemini 3.1 Pro — to author the template. You pay $0.05–$1 and you get a high-quality, grammar-safe, multi-variant template. 2. **Render forever, locally.** Spintax resolves the template on your CPU. Thousands of variants cost effectively zero. 3. **Update when meaning changes.** Regenerate the template only when the facts or positioning change — not when the vendor renames a model. For a stable product, that might be quarterly. You pay for quality where it matters. You stop paying for quantity. ### Pick your tier with a model Paste this page into your chat with Claude, GPT, or Gemini, then add: > _Given N articles per month, a budget of $X, and a quality bar of (draft / publish-with-edits / publish-raw), recommend a model tier and discount stack. Show the math._ The model has every number it needs to run the calculation. Pair it with the [authoring-guide series](https://spintax.net/docs/authoring-mindset.md) if you also want a templatized output. Agents can read this page as clean Markdown — [spintax.net/ai-content-costs.md](https://spintax.net/ai-content-costs.md): the same numbers without the page chrome. ## Caveats - **Prices change, and so do names.** The entire table above turned over between April and July 2026 — and then again between July and August. Recheck the provider pages before budgeting, and never hardcode a model ID you have not re-verified this quarter. - **Cache writes are not free.** The headline “90% off” applies to cache _reads_. Anthropic and OpenAI both bill a premium on the write, and Google bills hourly storage. A cache that gets read once is more expensive than no cache. - **Batch is not universally 50%.** xAI is 20%. Moonshot is 40% and excludes its flagship. Alibaba refuses to stack batch with cache. DeepSeek’s discount is the time of day, not a batch queue. Do not assume the stack applies before you check. - **Long context costs extra on some providers.** Gemini Pro tiers and Grok 4.6 roughly double past 200K tokens. Anthropic and OpenAI charge flat rates across their context windows. - **Quality is not the same at every tier.** Budget tiers are fine for scaffolding and rewrites. Nuanced voice, long-context fidelity, and factual grounding still favour flagship. Use cheap models for local renders, not for the template itself. ## Where to go next Ready to turn one good AI generation into thousands of renders? Start with the authoring series: - [Reverse authoring mindset](https://spintax.net/docs/authoring-mindset.md): Write the text first, add markup last. - [Variables & multi-site reuse](https://spintax.net/docs/variables.md): The mechanism that turns one template into a network. - [Permutations in practice](https://spintax.net/docs/permutations.md): Where the real variety lives. - [Grammar-safe synonymization](https://spintax.net/docs/grammar-safe-spintax.md): Keep every render readable. Data sources: Anthropic, OpenAI, Google, xAI, DeepSeek, Alibaba, Moonshot, Mistral pricing pages — verified 2026-08-20. Per-article calculations assume 3,000 input tokens + 3,500 output tokens per article. --- [Open in playground](https://spintax.net/play/) [Back to all guides](https://spintax.net/docs.md) --- Source: https://spintax.net/ai-content-costs/