LLM cost calculator: cloud API vs local GPU

For a given workload, estimate monthly cost on a cloud API offering and the cost of running an equivalent open-weight model on your own GPU hardware, with amortisation, energy and labour made explicit. Not a benchmark: local throughput is estimated from hardware specifications with efficiency factors you can override below. All amounts in USD. EUR display

Fairness controls

fp8 is approx. -0.5 benchmark points vs fp16 (approximate until Phase 2).

Amortisation horizon
Workload
Or compute from requests/day
Cloud

$11.00 / month ($1.10 / Mtok effective)

$0.31 / 1M chars (o200k, blended slice)

Breakdown
Input$3.00
Output$8.00
Cache read$0.00
Cache write$0.00
Batch saving-$0.00
Per-request fees$0.00
Other hosts of this model
Host$/month$/Mtok
MoonshotAI: Kimi K2 0905 (DeepInfra)$10.82$1.08
MoonshotAI: Kimi K2 0905 (Novita AI)$12.62$1.26
MoonshotAI: Kimi K2 0905 (Novita AI)$12.79$1.28
MoonshotAI: Kimi K2 0905 (Novita AI)$12.79$1.28
MoonshotAI: Kimi K2 0905 (OpenRouter)$13.60$1.36
MoonshotAI: Kimi K2 0905 (OpenRouter)$13.60$1.36
MoonshotAI: Kimi K2 0905 (Google Vertex AI, unknown)$13.60$1.36
MoonshotAI: Kimi K2 0905 (Novita AI, fp8)$13.60$1.36
MoonshotAI: Kimi K2 0905 (Fireworks AI)$13.60$1.36
MoonshotAI: Kimi K2 0905 (Fireworks AI)$13.60$1.36
MoonshotAI: Kimi K2 0905 (Fireworks AI)$13.60$1.36
MoonshotAI: Kimi K2 0905 (AWS Bedrock)$13.60$1.36
MoonshotAI: Kimi K2 0905 (Novita AI)$13.60$1.36
MoonshotAI: Kimi K2 0905 (Google Vertex AI)$13.60$1.36
MoonshotAI: Kimi K2 0905 (AWS Bedrock)$16.50$1.65
MoonshotAI: Kimi K2 0905 (Together AI)$18.00$1.80
MoonshotAI: Kimi K2 0905 (Hyperbolic)$20.00$2.00
Local hardware
Custom hardware (used only when Hardware = Custom)
Cost assumptions
Advanced: measured throughput overrides
Cloud$11.00 / mo
Local, workload$– / Mtok
Local, capacity$– / Mtok
Cheapercloud

Breakeven and sensitivity

Breakeven volume: – tokens/month. Breakeven utilisation: –.

Monthly cost vs. tokens/month

Cloud cost scales linearly with volume; local cost steps up each time another box is needed. Log scale on tokens/month.

Assumptions

KeyValueSourceOverridden
hoursPerMonth730CALCULATOR_SPEC.md §2no
daysPerMonth30.4CALCULATOR_SPEC.md §3no
weightsOverheadFactor1.05CALCULATOR_SPEC.md §6.1no
fixedOverheadGb2CALCULATOR_SPEC.md §6.1no
tpEffPcie0.85CALCULATOR_SPEC.md §6.1no
tpEffNvlink0.95CALCULATOR_SPEC.md §6.1no
tpEffUnified1CALCULATOR_SPEC.md §6.1no
defaultContextTokens8192T08 default (not specified in CALCULATOR_SPEC.md)no
decodeEff0.7CALCULATOR_SPEC.md §6.2yes
prefillEff0.5CALCULATOR_SPEC.md §6.2yes
batchEffK0.04CALCULATOR_SPEC.md §6.2yes
utilisation0.3CALCULATOR_SPEC.md §6.2no
maxUtilisation0.95CALCULATOR_SPEC.md §6.2no
sparesPct0.1CALCULATOR_SPEC.md §6.3no
residualPct0.2CALCULATOR_SPEC.md §6.3no
loadFactor0.8CALCULATOR_SPEC.md §6.3yes
horizonMonths36CALCULATOR_SPEC.md §6.3no
labourHours2CALCULATOR_SPEC.md §6.3no
labourRate60CALCULATOR_SPEC.md §6.3no
hostingMonth0CALCULATOR_SPEC.md §6.3no
softwareMonth0CALCULATOR_SPEC.md §6.3no
electricityUsdKwh0.1419src/data/electricity_defaults.json (US); EIA Electric Power Monthly via commercialenergyadvisors.comno
latencyTargetTps20CALCULATOR_SPEC.md §3no
tokenizerSliceblendedCALCULATOR_SPEC.md §7yes