Token prices are hard to compare across providers — enter your usage once and see the real monthly bill for every major model.
Live data · updated 2026-09-15 18:18 · sources: LMArena + OpenRouter
For a typical indie workload (50M in / 10M out per month), the spread between the cheapest and most expensive frontier model is over 40× — most apps can run on muse-spark-1.2 (xHigh) for single-digit dollars.
| # | Model | Arena | $/1M in | $/1M out | Monthly cost |
|---|---|---|---|---|---|
| 1 | muse-spark-1.2 (xHigh) Meta | 1500 | $1.25 | $4.25 | $105.00 |
| 2 | qwen3.8-max Alibaba | 1302 | $2 | $6 | $160.00 |
| 3 | gemini-omni-1.1-flash Google | 1515 | $1.5 | $9 | $165.00 |
| 4 | gpt-5.6-sol-xhigh OpenAI | 1257 | $4 | $20 | $400.00 |
| 5 | claude-opus-5-high Anthropic | 1516 | $5 | $25 | $500.00 |
| 6 | claude-opus-4-6-high Anthropic | 1505 | $5 | $25 | $500.00 |
| 7 | claude-opus-4-6-search Anthropic | 1253 | $5 | $25 | $500.00 |
| 8 | claude-opus-4-7-high Anthropic | 1502 | $5 | $25 | $500.00 |
| 9 | claude-opus-5-max Anthropic | 1687 | $5 | $25 | $500.00 |
| 10 | claude-opus-4-6 Anthropic | 1507 | $5 | $25 | $500.00 |
| 11 | gpt-5.5-search OpenAI | 1242 | $5 | $30 | $550.00 |
| 12 | gpt-image-2 (medium) OpenAI | 1381 | $8 | $30 | $700.00 |
| 13 | claude-fable-5 Anthropic | 1506 | $10 | $50 | $1,000 |
| 14 | gpt-6-astra-max OpenAI | 1800 | $10 | $50 | $1,000 |
| 15 | claude-fable-5.1-max Anthropic | 1758 | $10 | $50 | $1,000 |
50M input + 10M output tokens per month. Cheapest: muse-spark-1.2 (xHigh) at $105.00. Prices from OpenRouter, arena ratings from LMArena.
One English word ≈ 1.3 tokens; a typical chat turn (your prompt + history) runs 500–2,000 input tokens. RAG pipelines multiply input fast — every retrieval re-sends context. Measure one week of real traffic before committing to a model tier.
As of 2026-09-15 18:18, output tokens cost 3–5× more than input on most models. Chat-heavy products should weight output price; summarization and RAG products should weight input. Batch APIs (off-peak) cut costs ~50% on several providers if latency is flexible.
A model 20% cheaper that fails 5% more often costs more once you count retries and bad user experiences. Use the arena rating column as a quality floor, then optimize price above it.