calculator
Workload in, monthly bill out.
Change any input and every row updates. The share link carries your numbers with it, so you can paste a scenario into a thread instead of describing it.
- Requests
- 6k/mo
- Input per request
- 2.5k tokens
- Output per request
- 350 tokens
- Cache hit
- 0%
- Batch
- no
- Cheapest monthly
- $0.348
| Model | In /1M | Out /1M | Per request | Per month | Share | Buy |
|---|---|---|---|---|---|---|
| Mistral NemoMistralcheapestopen weights | $0.019 | $0.030 | $0.000058 | $0.348 | Open | |
| GPT-OSS 20BOpenAIopen weights | $0.018 | $0.090 | $0.000076 | $0.459 | Openaff | |
| Qwen3.7 FlashAlibaba | $0.030 | $0.130 | $0.000121 | $0.723 | Open | |
| Llama 3.1 8BMetaopen weights | $0.050 | $0.080 | $0.000153 | $0.918 | Openaff | |
| DeepSeek V4 FlashDeepSeek | $0.086 | $0.171 | $0.000275 | $1.65 | Open | |
| Gemini 2.5 Flash-LiteGoogle | $0.100 | $0.400 | $0.00039 | $2.34 | Open | |
| GLM 5.3 FlashZ.ai | $0.150 | $0.500 | $0.00055 | $3.30 | Open | |
| Gemini 3.8 FlashGoogle | $0.750 | $3.75 | $0.0032 | $19.13 | Open | |
| Grok 4.3xAI | $1.25 | $2.50 | $0.004 | $24.00 | Open | |
| Claude Haiku 4.5Anthropic | $1.00 | $5.00 | $0.0043 | $25.50 | Openaff | |
| GPT-5.1OpenAI | $1.25 | $10.00 | $0.0066 | $39.75 | Open | |
| Grok 4.6xAI | $2.00 | $6.00 | $0.0071 | $42.60 | Open | |
| Claude Sonnet 5Anthropic | $2.00 | $10.00 | $0.0085 | $51.00 | Openaff | |
| Gemini 3.1 ProGoogle | $2.00 | $12.00 | $0.0092 | $55.20 | Open | |
| GPT-5.2 CodexOpenAI | $1.75 | $14.00 | $0.0093 | $55.65 | Open | |
| GPT-5.4OpenAI | $2.50 | $15.00 | $0.0115 | $69.00 | Open | |
| Kimi K3Moonshot | $3.00 | $15.00 | $0.0128 | $76.50 | Open | |
| Claude Opus 5.5Anthropic | $4.00 | $20.00 | $0.017 | $102 | Openaff | |
| GPT-5.5OpenAI | $5.00 | $30.00 | $0.023 | $138 | Open | |
| Claude Fable 5.1Anthropic | $10.00 | $50.00 | $0.0425 | $255 | Openaff | |
| GPT-6 AstraOpenAI | $10.00 | $50.00 | $0.0425 | $255 | Open |
Rent a card instead?
Open weights change the math. Rent a GPU by the hour and your cost is throughput, not tokens.
- Tokens per month
- 17.1M
- Hours at that speed
- 95
- Rented GPU cost
- $70.30
- Cheapest API above
- $0.348
The API is cheaper by $69.95/mo at 50 tokens/sec. Raise throughput or run the card longer only if you have other jobs for it. Throughput is yours to measure: model size, quantisation, and batching move it by an order of magnitude. Rented prices are Runpod community rates, verified 2026-09-13.
How each row is calculated
fresh_input = input_tokens x (1 - cache_hit)
cached_input = input_tokens x cache_hit
per_request = (fresh_input / 1e6) x price_in
+ (cached_input / 1e6) x price_cache_read
+ (output_tokens / 1e6) x price_out
monthly = per_request x requests
batch = monthly x discount (only where the provider publishes one)- Caching only reduces a row when the model publishes a cache-read rate. Otherwise it is ignored.
- Batch applies the provider's published discount for that model, not a flat 50%.
- Prices are list rates without tax, read on Sep 24, 2026 from 6 sources.
- Not modelled: retries, failed calls, tool calls, embeddings, image and audio tokens, or negotiated discounts.
Prices read from 6 provider pricing pages on Sep 24, 2026. Some outbound links may be affiliate links. How that works.