Enter your token volume once and see every model ranked from cheapest to priciest — GPT-4o, GPT-4.1, o1, Claude, Gemini, DeepSeek, and Mistral, side by side.
Enter your details and click Calculate to see results
The LLM API cost calculator exists because picking a model by memory alone stopped being realistic once OpenAI, Anthropic, Google, DeepSeek, and Mistral were each shipping several priced tiers at once. Instead of estimating one model's bill in isolation, this tool takes a single input/output token profile and request volume and instantly ranks all 14 models — GPT-4o, GPT-4.1, o1, o3-mini, Claude Sonnet 4.6, Claude Opus 4.8, Gemini 1.5 Pro/Flash, DeepSeek V3/R1, Mistral Large, and more — from cheapest to priciest, so you see exactly where your chosen model sits relative to every real alternative.
Every LLM provider bills per token, but the input price and output price are set independently, and the ratio between them differs by provider and even by model generation. This calculator turns your average input tokens, output tokens, and daily request count into a per-request dollar figure and a projected monthly cost for all 14 supported models simultaneously, then sorts the list so the cheapest and priciest options — and your own pick — are immediately visible side by side.
It's built for developers choosing a model for a new AI feature, engineering leads weighing a provider switch to cut spend, product managers scoping a vendor comparison for a roadmap review, and startup founders projecting API burn rate before committing to an architecture. It's equally useful for anyone auditing an existing integration after a provider changes its published rates.
Headline "per 1M tokens" pricing is misleading on its own, because a model with cheap input tokens can still be one of the pricier options once your typical response length grows — output tokens are billed at 3-5x the input rate on nearly every provider in this list. Comparing all 14 models at your actual token mix, rather than eyeballing separate pricing pages, avoids the common mistake of picking a "cheap" model that turns out expensive for an output-heavy workload.
How this LLM API cost calculator turns a token profile into a ranked cost comparison
Input Price and Output Price are each model's published dollars-per-million-token rate, shown directly in the model dropdown. The calculator applies the identical formula to all 14 models at once — only the two price constants change per model — which is what makes the ranking an apples-to-apples comparison.
Input tokens cover your prompt, system message, and any retrieved context sent with each call. Output tokens are the model's generated response — priced separately, and almost always at a higher per-token rate than input.
The results highlight the single cheapest model by monthly cost, the priciest, and whichever model you selected as "Your Selection" — so you can read off the exact dollar gap between your current choice and every alternative.
Cost per request is multiplied by requests per day, then by 30, to approximate a 30-day month. This is a projection, not a calendar-exact billing cycle, so actual invoices may differ slightly by a day or two of usage.
From entering your token profile to reading the full 14-model ranking
Type the average input tokens per call — your prompt, system message, and any retrieved context combined. The default is 1,000 tokens.
Type the expected length of the model's response in tokens. The default is 500 tokens — output tokens are usually the pricier half of the bill.
Enter your expected daily call volume, so a single request's cost can be rolled up into a projected monthly figure for every model.
Pick the model you're using or considering from the Your Selection dropdown. It doesn't change the ranking — it just highlights that row as "your pick."
The calculator computes cost per request and monthly cost for all 14 models at your exact numbers and sorts them cheapest to priciest.
Review the cheapest and your-selection cost boxes, the full ranked list with badges, and the bar chart (monthly cost by model) and doughnut chart (your selection's input vs output split).
Using the calculator's own default scenario — 1,000 input / 500 output tokens, 1,000 requests/day, Claude Sonnet 4.6 selected
Suppose you're comparing providers for a feature that sends 1,000 input tokens and expects 500 output tokens per request, at 1,000 requests per day, with Claude Sonnet 4.6 ($3.00 / $15.00 per 1M input/output tokens) as your current pick.
Explanation: At this exact volume, Claude Sonnet 4.6 lands 11th out of 14 by price — output tokens (500 × $15.00/1M = $0.0075) make up roughly 71% of its per-request cost even though it's outnumbered two-to-one by input tokens, because its output rate is 5× its input rate. Switching to Gemini 1.5 Flash would cut the monthly bill from $315.00 to $6.75, a 97.9% reduction — but that comparison alone doesn't confirm Flash produces comparable output quality for the task at hand, which is why price ranking should be a shortlist step, not the final decision.
What your selected model's projected monthly cost generally implies
| Your Selection's Monthly Cost | What It Generally Means | Recommended Next Step |
|---|---|---|
| Under $25 | Ultra-low-cost tier — likely a mini/flash-class model | Proceed as-is; the price gap to switch further is usually not worth the effort |
| $25 – $150 | Budget-friendly production tier | Good default for most features; re-check ranking only if volume grows sharply |
| $150 – $400 | Mid-tier flagship pricing | Compare against the top 2-3 cheaper models in the ranking before committing at scale |
| $400 – $1,000 | Premium model tier | Justify the premium with a documented quality difference, or test a cheaper alternative |
| Over $1,000 | Frontier/top-tier pricing | Rigorously validate that no cheaper model in the ranking can do the job before scaling |
If your selection is far from the cheapest row: that gap is the real, quantified cost of choosing capability over price. Use it to decide whether the quality difference is worth paying for, rather than assuming the priciest model is automatically the best fit.
If the ranking shifts a lot when you change output tokens: that's expected — output pricing varies more between providers than input pricing does, so an output-heavy workload (long-form writing, detailed summaries) tends to separate cheap and expensive models further than a short-answer workload does.
These are planning-stage estimates based on published standard pricing, not exact invoices. Always reconcile against your provider's live billing dashboard once a feature is in production.
This calculator provides planning estimates only. Actual charges depend on your provider's live pricing, any negotiated or promotional rates on your account, and provider-side rounding. Always verify against your provider's billing dashboard before finalizing a budget.
Where ranking per-request LLM cost across providers genuinely helps
Rank all 14 models before writing production code, so the first line of code targets a model you've already price-checked.
Quantify the exact monthly savings (or cost increase) of moving from your current model to an alternative before committing engineering time to a migration.
Run the same token profile across OpenAI, Anthropic, and Google's flagship models to see which wins on price for your specific workload.
Show a manager or client the concrete dollar gap between a frontier model like o1 and a mid-tier alternative to support a build vs. cost tradeoff discussion.
Check how the ranking holds up as you project requests-per-day growth from a pilot to full production traffic.
See exactly how much a lightweight model like GPT-4o mini or Gemini 1.5 Flash saves versus its flagship sibling at your real volume.
Test how a verbose response style versus a terse, structured one shifts your ranking, since output tokens are priced highest across every provider here.
Bring a ranked, apples-to-apples price comparison into a vendor evaluation instead of relying on each provider's own pricing page in isolation.
Identify which cheap model can handle simple requests and which premium model should be reserved for complex ones in a multi-model routing architecture.
Fold a specific, ranked LLM API cost figure into a pitch deck or burn-rate model instead of an unverified round number.
Estimate the API spend for a research project or coursework assignment that calls multiple LLMs before requesting a grant or lab budget.
Re-run the comparison whenever a provider updates its published rates to confirm your current model is still the right price/quality tradeoff.
What this LLM API cost calculator does well, and where it can't replace live billing data
All 14 supported models, ranked cheapest to priciest at the calculator's default volume: 1,000 input / 500 output tokens per request, 1,000 requests/day
| Rank | Model | Input / Output ($ per 1M) | Cost / Request | Monthly Cost |
|---|---|---|---|---|
| 1 | Gemini 1.5 Flash | $0.075 / $0.30 | <$0.01 | $6.75 |
| 2 | GPT-4o mini | $0.15 / $0.60 | <$0.01 | $13.50 |
| 3 | DeepSeek V3 | $0.27 / $1.10 | <$0.01 | $24.60 |
| 4 | DeepSeek R1 | $0.55 / $2.19 | <$0.01 | $49.35 |
| 5 | o3-mini | $1.10 / $4.40 | <$0.01 | $99.00 |
| 6 | Claude Haiku 4.5 | $1.00 / $5.00 | <$0.01 | $105.00 |
| 7 | Mistral Large | $2.00 / $6.00 | <$0.01 | $150.00 |
| 8 | GPT-4.1 | $2.00 / $8.00 | <$0.01 | $180.00 |
| 9 | GPT-4o | $2.50 / $10.00 | <$0.01 | $225.00 |
| 10 | Gemini 1.5 Pro | $3.50 / $10.50 | <$0.01 | $262.50 |
| 11 | Claude Sonnet 4.6 | $3.00 / $15.00 | $0.0105 | $315.00 |
| 12 | Claude Opus 4.8 | $5.00 / $25.00 | $0.0175 | $525.00 |
| 13 | Claude Fable 5 | $10.00 / $50.00 | $0.0350 | $1,050.00 |
| 14 | o1 | $15.00 / $60.00 | $0.0450 | $1,350.00 |
Summary: This LLM API cost calculator gives you an instant, free ranking of per-request and monthly cost across 14 models from OpenAI, Anthropic, Google, DeepSeek, and Mistral, so you can see the real dollar gap between providers at your exact token volume instead of comparing pricing pages by eye. Pair it with the AI Token Calculator and Fine-tuning Cost Estimator for a fuller AI cost picture.
Common questions about comparing LLM API pricing
Official documentation to complement this calculator — always verify live rates before finalizing a budget
Explore other AI & tech tools