calculator

Workload in, monthly bill out.

Change any input and every row updates. The share link carries your numbers with it, so you can paste a scenario into a thread instead of describing it.

Requests
6k/mo
Input per request
2.5k tokens
Output per request
350 tokens
Cache hit
0%
Batch
no
Cheapest monthly
$0.348

21 models ranked for 6k requests at 2.5k in / 350 out. Cheapest first.
ModelIn /1MOut /1MPer requestPer monthShareBuy
Mistral NemoMistralcheapestopen weights$0.019$0.030$0.000058$0.348Open
GPT-OSS 20BOpenAIopen weights$0.018$0.090$0.000076$0.459Openaff
Qwen3.7 FlashAlibaba$0.030$0.130$0.000121$0.723Open
Llama 3.1 8BMetaopen weights$0.050$0.080$0.000153$0.918Openaff
DeepSeek V4 FlashDeepSeek$0.086$0.171$0.000275$1.65Open
Gemini 2.5 Flash-LiteGoogle$0.100$0.400$0.00039$2.34Open
GLM 5.3 FlashZ.ai$0.150$0.500$0.00055$3.30Open
Gemini 3.8 FlashGoogle$0.750$3.75$0.0032$19.13Open
Grok 4.3xAI$1.25$2.50$0.004$24.00Open
Claude Haiku 4.5Anthropic$1.00$5.00$0.0043$25.50Openaff
GPT-5.1OpenAI$1.25$10.00$0.0066$39.75Open
Grok 4.6xAI$2.00$6.00$0.0071$42.60Open
Claude Sonnet 5Anthropic$2.00$10.00$0.0085$51.00Openaff
Gemini 3.1 ProGoogle$2.00$12.00$0.0092$55.20Open
GPT-5.2 CodexOpenAI$1.75$14.00$0.0093$55.65Open
GPT-5.4OpenAI$2.50$15.00$0.0115$69.00Open
Kimi K3Moonshot$3.00$15.00$0.0128$76.50Open
Claude Opus 5.5Anthropic$4.00$20.00$0.017$102Openaff
GPT-5.5OpenAI$5.00$30.00$0.023$138Open
Claude Fable 5.1Anthropic$10.00$50.00$0.0425$255Openaff
GPT-6 AstraOpenAI$10.00$50.00$0.0425$255Open

Rent a card instead?

Open weights change the math. Rent a GPU by the hour and your cost is throughput, not tokens.


Tokens per month
17.1M
Hours at that speed
95
Rented GPU cost
$70.30
Cheapest API above
$0.348

The API is cheaper by $69.95/mo at 50 tokens/sec. Raise throughput or run the card longer only if you have other jobs for it. Throughput is yours to measure: model size, quantisation, and batching move it by an order of magnitude. Rented prices are Runpod community rates, verified 2026-09-13.

How each row is calculated

fresh_input   = input_tokens x (1 - cache_hit)
cached_input  = input_tokens x cache_hit
per_request   = (fresh_input / 1e6) x price_in
              + (cached_input / 1e6) x price_cache_read
              + (output_tokens / 1e6) x price_out
monthly       = per_request x requests
batch         = monthly x discount   (only where the provider publishes one)
  • Caching only reduces a row when the model publishes a cache-read rate. Otherwise it is ignored.
  • Batch applies the provider's published discount for that model, not a flat 50%.
  • Prices are list rates without tax, read on Sep 24, 2026 from 6 sources.
  • Not modelled: retries, failed calls, tool calls, embeddings, image and audio tokens, or negotiated discounts.

Full method, sources, and known gaps

Prices read from 6 provider pricing pages on Sep 24, 2026. Some outbound links may be affiliate links. How that works.