Photo by Kindel Media on Pexels
Choosing an AI model for an app now means reading a price sheet as carefully as a feature list. This guide puts the list prices of seven current models side by side, explains the billing terms that trip up most newcomers, and ends with a simple way to choose based on who you are and how much you plan to send. Every price is in US dollars per million tokens and comes from vendor announcements and coverage published in September and October 2026.
The article compares cost only. It does not rank quality, and it shows no benchmarks. A cheap model that fails your task is an expensive purchase, so treat these numbers as the starting point for your own testing rather than the final answer.
What each model costs in October 2026
The seven models below range from $0.30 per million input tokens at the low end to $10 at the top. The vendors publish these figures on their own pricing pages, including OpenAI's API pricing page and Anthropic's pricing page, and the table collects them in one place.
| Model | Company | Input | Output | Example job cost |
|---|---|---|---|---|
| GPT-6 Astra | OpenAI | $10 | $50 | $20.00 |
| Claude Fable 5.1 | Anthropic | $10 | $50 | $20.00 |
| Claude Opus 5.5 | Anthropic | $4 | $20 | $8.00 |
| GPT-6.1 Sol | OpenAI | $2 | $10 | $4.00 |
| Muse Spark 1.3 | Meta | $1.25 | $4.25 | $2.10 |
| Gemini 3.8 Flash | $0.75 | $3.75 | $1.50 | |
| DeepSeek V4.1 Flash | DeepSeek | up to $0.30 | up to $1.20 | $0.54 |
These are list prices at the time of writing and can change, so confirm each figure on the vendor's own page before you budget. The example job column is explained in the next section.
A worked example: one million tokens in, 200,000 out
The same job costs anywhere from $0.54 to $20.00 depending on the model, a gap of roughly 37 times. The job in the table sends 1 million input tokens and receives 200,000 output tokens, priced at list rates with no caching and no discounts.
Take GPT-6 Astra as the arithmetic template. One million input tokens at $10 is $10.00, and 200,000 output tokens at $50 per million is $10.00, so the job costs $20.00. Claude Fable 5.1 has the same rates and lands on the same $20.00. Claude Opus 5.5 comes to $8.00, GPT-6.1 Sol to $4.00, Muse Spark 1.3 to $2.10, and Gemini 3.8 Flash to $1.50. At the bottom, DeepSeek V4.1 Flash costs $0.54 using its listed peak rates of $0.30 and $1.20.
What a token is
A token is a small chunk of text, and in English it works out to roughly three quarters of a word, or a few characters. Models read and write in these chunks rather than in whole words, which is why pricing is quoted per million of them.
Context window size is measured in the same unit. Gemini 3.8 Flash, Muse Spark 1.3 and DeepSeek V4.1 Flash each list a 1 million token window, and DeepSeek lists a 384K maximum output. The previous GPT-6 Sol had a 1,050,000 token window. A bigger window lets you send more in one request, but every token you send is billed, so a large window is a ceiling and not a recommendation to fill it.
Why input and output are priced differently
Output tokens cost more because the model has to generate them one at a time, while input tokens can be processed in bulk. Every model in the table follows this pattern, with output priced at roughly three to five times the input rate.
Look at the ratios. GPT-6 Astra, Claude Fable 5.1, Claude Opus 5.5 and GPT-6.1 Sol all charge five times as much for output as for input. Gemini 3.8 Flash is also five times. Muse Spark 1.3 is closer to 3.4 times, and DeepSeek V4.1 Flash is four times at its listed peak rates. The practical lesson is that the length of the answer you ask for matters as much as the length of the document you send.
Anything that makes answers shorter, such as asking for a summary in a set number of sentences or requesting structured fields instead of free prose, reduces the expensive side of the bill. Tasks that mostly read and rarely write, such as classification or search over documents, are cheaper per unit of work than tasks that write long drafts.
What cached input means and why it matters
Cached input is text the vendor has already processed in a recent request, and it is billed at a steep discount when you send it again. Four of the models in this guide publish a cached rate: GPT-6 Astra at $1 per million, GPT-6.1 Sol at $0.10, Claude Fable 5.1 at $0.25 for cache reads, and Claude Opus 5.5 at $0.20 for cache reads.
Caching matters most for agents and chat. An agent that works through a long task re-sends its instructions, tool descriptions and earlier steps on every turn. A chat app re-sends the whole conversation each time the user writes a new message. Without caching you pay the full input rate for the same opening text again and again. With caching, that repeated prefix drops to a fraction of the price.
As an illustration, the $10 input rate for GPT-6 Astra falls to $1 for cached tokens, a tenth of the price. For a long-running agent where most of each request is repeated context, the realistic input cost can be far below the list figure. The Muse Spark, Gemini and DeepSeek figures in this guide are plain list prices, so check each vendor's page for any caching terms that apply to them.
Why the cheapest model is not always the cheapest job
The cheapest rate per token only gives the cheapest job if the model finishes the work in one clean pass. Three things erode the saving: retries, longer outputs and quality gaps.
Retries come first. If a low-priced model fails to follow your format one time in five, you pay for the failed attempt and then pay again, plus any code you wrote to detect and repair the failure. Longer outputs come second, since a model that writes more words to reach the same answer bills more output tokens, and output is the pricey side. Quality gaps come third. If a person has to review or fix the result, the labor cost can dwarf the model bill.
The reverse also happens. A premium model that solves a hard task on the first try can cost less per finished result than a bargain model that needs four attempts. The only reliable way to know is to run your own task through two or three candidates and compare the total spend per accepted result, instead of the price per million tokens.
Tier logic: flagship, workhorse and fast and cheap
Vendors sort their models into rough tiers by price, and the seven here fall into three groups. The top group is the flagship tier, with GPT-6 Astra and Claude Fable 5.1 both at $10 input and $50 output. The middle group is the workhorse tier, with Claude Opus 5.5 at $4 / $20 and GPT-6.1 Sol at $2 / $10. The bottom group is the fast and cheap tier, with Muse Spark 1.3 at $1.25 / $4.25, Gemini 3.8 Flash at $0.75 / $3.75 and DeepSeek V4.1 Flash at up to $0.30 / $1.20.
Most requests do not need the most expensive model, so a common setup sends easy, high-volume work to a fast tier model and reserves the flagship for the hard share.
For details on each model, see our explainers on OpenAI GPT-6 Astra, GPT-6.1 Sol, Claude Fable 5.1, Claude Opus 5.5, Meta Muse Spark 1.3, Gemini 3.8 Flash and DeepSeek V4.1 Flash. Google describes the Gemini 3.8 Flash figures as introductory rates, so they may rise, and the Gemini API pricing page shows the current numbers.
Open weights versus closed models
Open weights means the trained model files can be downloaded and run on your own hardware, while a closed model is reachable only through the vendor's service. DeepSeek V4.1 Flash is released under the MIT license, which is a permissive license that allows commercial use and modification.
The difference matters for cost in two ways. First, an open model can be hosted by several providers, so competition can push the price down, and you can choose a host that fits your region or privacy needs. Second, if your volume is high enough you can run the model yourself and pay for hardware instead of per token. That trade comes with real work, including servers, scaling, monitoring and updates, so it suits teams that already have that skill.
Closed models give up that control in exchange for convenience. You get a single bill, a managed service and support, and you do not maintain anything. DeepSeek also publishes its own hosted rates, and its pricing documentation lists the peak figures used in the table above. Because the license is open, third-party hosts may quote different prices, which is another reason to treat any single number as a reference point.
When paying for a flagship is worth it
A flagship model is worth its price when a wrong answer is costly, when the task is hard enough that cheaper models fail often, or when volume is low enough that the bill stays small anyway. Contract analysis, difficult code changes, and decisions with legal or financial weight are typical examples.
Volume is the factor people forget. At $20.00 for a job of one million input tokens and 200,000 output tokens, a flagship is easy to afford when you run that job a few times a week. The same job run ten thousand times a day is a very different budget line. Do the multiplication before you commit, and remember that caching can cut the repeated input portion sharply.
A simple decision guide
Your best choice depends on how much you send and how much a mistake costs you, and the guide below covers three common situations.
If you are a solo developer building a side project, start with a workhorse or fast tier model. GPT-6.1 Sol at $2 / $10 or Gemini 3.8 Flash at $0.75 / $3.75 keeps experiments cheap, and your monthly bill will probably stay small. Move one specific, difficult feature to a flagship only if testing shows the cheaper model fails it.
If you run a startup with high volume, price per token compounds fast. Route most traffic to a fast tier model such as Muse Spark 1.3 or DeepSeek V4.1 Flash, add caching wherever your prompts repeat, and measure cost per accepted result instead of cost per call. Reserve Claude Opus 5.5 or a $10 / $50 flagship for the share of requests where quality drives revenue. Track spend per feature from day one so a surprise does not arrive at the end of the month.
If you are a non-technical user who just chats with an AI through a subscription, API pricing does not apply to you. You pay a flat monthly fee to the app, and the per-token numbers in this guide are what developers pay. To compare chat products and what they include, browse the ForgeChatAI tools directory.
Caveats: list price versus the real bill
A list price is where a budget starts, and the final bill is usually different. Several things sit between the two.
Rate limits can force you onto a higher plan or a second vendor before you reach the volume you expected. Regional pricing and data residency options may carry a surcharge in some places, and taxes or currency conversion change the number you actually pay. Discounts for committed spend or batch processing can lower the rate, and extra charges for tools such as search or code execution can raise it. Introductory rates, like those on Gemini 3.8 Flash, can end. DeepSeek's figures are its listed peak prices, so the number you pay can vary with the time of use.
Finally, this guide shows no benchmarks and makes no claim about which model is better at any task. Price and quality are separate questions, and you answer the second one by testing on your own data. Check each vendor's page for the current terms before you build a budget around any figure here.
Frequently asked questions
Which model is cheapest per token in October 2026? Among the seven models covered here, DeepSeek V4.1 Flash has the lowest listed rates, at up to $0.30 per million input tokens and $1.20 per million output tokens. Gemini 3.8 Flash follows at $0.75 and $3.75, though Google describes those as introductory rates.
Why do GPT-6 Astra and Claude Fable 5.1 cost the same? Both list $10 per million input tokens and $50 per million output tokens, so a job costs the same on either at list prices. Their cached input rates differ, with Astra at $1 and Fable 5.1 cache reads at $0.25, which can change the bill for apps that repeat long prompts.
Does a chat subscription use these API prices? No. A chat subscription is a flat monthly fee charged by the app, and the per-token prices here apply to developers who call a model through an API. If you only use a chat app, you can skip the cost math and compare products in the tools directory instead.
How many tokens is a page of text? Roughly 300 to 500 tokens for a page of ordinary English, since a token is about three quarters of a word. The exact count varies by model and by the kind of text, with code and non-English languages usually using more tokens per word.
Will these prices change? They can. These are list prices at the time of writing, and vendors adjust them, retire models and end introductory offers. Confirm the current numbers on each vendor's pricing page before committing to a budget.
Is the cheapest model always the best value? No. Retries, longer outputs and results that need human review can erase a low token price. Compare the total cost per accepted result on your own task, using two or three candidate models, before you settle on one.