LLM cost calculator: cloud API vs local GPU

For a given workload, estimate monthly cost on a cloud API offering and the cost of running an equivalent open-weight model on your own GPU hardware, with amortisation, energy and labour made explicit. Not a benchmark: local throughput is estimated from hardware specifications with efficiency factors you can override below. All amounts in USD. EUR display

Fairness controls

fp8 is approx. -0.5 benchmark points vs fp16 (approximate until Phase 2).

Amortisation horizon
Workload
Or compute from requests/day
Cloud

$10.58 / month ($1.06 / Mtok effective)

$0.30 / 1M chars (o200k, blended slice)

Breakdown
Input$1.38
Output$9.20
Cache read$0.00
Cache write$0.00
Batch saving-$0.00
Per-request fees$0.00
Other hosts of this model
Host$/month$/Mtok
Qwen: Qwen3 235B A22B (DeepInfra)$2.74$0.27
Qwen: Qwen3 235B A22B (Novita AI)$2.86$0.29
Qwen: Qwen3 235B A22B (DeepInfra)$3.24$0.32
Qwen: Qwen3 235B A22B (Nebius AI Studio)$3.60$0.36
Qwen: Qwen3 235B A22B (Nebius AI Studio)$3.60$0.36
Qwen: Qwen3 235B A22B (Novita AI)$4.40$0.44
Qwen: Qwen3 235B A22B (AWS Bedrock)$4.84$0.48
Qwen: Qwen3 235B A22B (Fireworks AI)$4.84$0.48
Qwen: Qwen3 235B A22B (Fireworks AI)$4.84$0.48
Qwen: Qwen3 235B A22B (AWS Bedrock)$4.84$0.48
Qwen: Qwen3 235B A22B (OpenRouter)$10.01$1.00
Qwen: Qwen3 235B A22B (Alibaba Cloud, unknown)$10.01$1.00
Qwen: Qwen3 235B A22B (Alibaba Cloud)$10.01$1.00
Qwen: Qwen3 235B A22B (OpenRouter)$10.58$1.06
Qwen: Qwen3 235B A22B (Alibaba Cloud)$10.58$1.06
Qwen: Qwen3 235B A22B (DeepInfra)$13.40$1.34
Qwen: Qwen3 235B A22B (Novita AI, fp8)$13.80$1.38
Qwen: Qwen3 235B A22B (Novita AI)$13.80$1.38
Qwen: Qwen3 235B A22B (Alibaba Cloud)$15.40$1.54
Qwen: Qwen3 235B A22B (venice, fp8)$16.70$1.67
Qwen: Qwen3 235B A22B (Hyperbolic)$20.00$2.00
Local hardware
Custom hardware (used only when Hardware = Custom)
Cost assumptions
Advanced: measured throughput overrides
Cloud$10.58 / mo
Local, workload$– / Mtok
Local, capacity$– / Mtok
Cheapercloud

Breakeven and sensitivity

Breakeven volume: – tokens/month. Breakeven utilisation: –.

Monthly cost vs. tokens/month

Cloud cost scales linearly with volume; local cost steps up each time another box is needed. Log scale on tokens/month.

Assumptions

KeyValueSourceOverridden
hoursPerMonth730CALCULATOR_SPEC.md §2no
daysPerMonth30.4CALCULATOR_SPEC.md §3no
weightsOverheadFactor1.05CALCULATOR_SPEC.md §6.1no
fixedOverheadGb2CALCULATOR_SPEC.md §6.1no
tpEffPcie0.85CALCULATOR_SPEC.md §6.1no
tpEffNvlink0.95CALCULATOR_SPEC.md §6.1no
tpEffUnified1CALCULATOR_SPEC.md §6.1no
defaultContextTokens8192T08 default (not specified in CALCULATOR_SPEC.md)no
decodeEff0.7CALCULATOR_SPEC.md §6.2yes
prefillEff0.5CALCULATOR_SPEC.md §6.2yes
batchEffK0.04CALCULATOR_SPEC.md §6.2yes
utilisation0.3CALCULATOR_SPEC.md §6.2no
maxUtilisation0.95CALCULATOR_SPEC.md §6.2no
sparesPct0.1CALCULATOR_SPEC.md §6.3no
residualPct0.2CALCULATOR_SPEC.md §6.3no
loadFactor0.8CALCULATOR_SPEC.md §6.3yes
horizonMonths36CALCULATOR_SPEC.md §6.3no
labourHours2CALCULATOR_SPEC.md §6.3no
labourRate60CALCULATOR_SPEC.md §6.3no
hostingMonth0CALCULATOR_SPEC.md §6.3no
softwareMonth0CALCULATOR_SPEC.md §6.3no
electricityUsdKwh0.1419src/data/electricity_defaults.json (US); EIA Electric Power Monthly via commercialenergyadvisors.comno
latencyTargetTps20CALCULATOR_SPEC.md §3no
tokenizerSliceblendedCALCULATOR_SPEC.md §7yes