LLM cost calculator: cloud API vs local GPU

For a given workload, estimate monthly cost on a cloud API offering and the cost of running an equivalent open-weight model on your own GPU hardware, with amortisation, energy and labour made explicit. Not a benchmark: local throughput is estimated from hardware specifications with efficiency factors you can override below. All amounts in USD. EUR display

Fairness controls

fp8 is approx. -0.5 benchmark points vs fp16 (approximate until Phase 2).

Amortisation horizon
Workload
Or compute from requests/day
Cloud

$0.70 / month ($0.0700 / Mtok effective)

$0.0197 / 1M chars (o200k, blended slice)

Breakdown
Input$0.30
Output$0.40
Cache read$0.00
Cache write$0.00
Batch saving-$0.00
Per-request fees$0.00
Other hosts of this model
Host$/month$/Mtok
Meta: Llama 4 Scout (OpenRouter)$1.80$0.18
Meta: Llama 4 Scout (DeepInfra, fp8)$1.80$0.18
Meta: Llama 4 Scout (DeepInfra)$1.80$0.18
Meta: Llama 4 Scout (Novita AI)$3.44$0.34
Meta: Llama 4 Scout (Novita AI)$3.44$0.34
Meta: Llama 4 Scout (Together AI)$3.44$0.34
Meta: Llama 4 Scout (Google Vertex AI, unknown)$4.30$0.43
Meta: Llama 4 Scout (Google Vertex AI)$4.30$0.43
Meta: Llama 4 Scout (Google Vertex AI)$4.30$0.43
Meta: Llama 4 Scout (Google Vertex AI)$4.30$0.43
Meta: Llama 4 Scout (Azure AI Foundry)$4.32$0.43
Meta: Llama 4 Scout (Azure AI Foundry)$4.32$0.43
Local hardware
Custom hardware (used only when Hardware = Custom)
Cost assumptions
Advanced: measured throughput overrides
Cloud$0.70 / mo
Local, workload$– / Mtok
Local, capacity$– / Mtok
Cheapercloud

Breakeven and sensitivity

Breakeven volume: – tokens/month. Breakeven utilisation: –.

Monthly cost vs. tokens/month

Cloud cost scales linearly with volume; local cost steps up each time another box is needed. Log scale on tokens/month.

Assumptions

KeyValueSourceOverridden
hoursPerMonth730CALCULATOR_SPEC.md §2no
daysPerMonth30.4CALCULATOR_SPEC.md §3no
weightsOverheadFactor1.05CALCULATOR_SPEC.md §6.1no
fixedOverheadGb2CALCULATOR_SPEC.md §6.1no
tpEffPcie0.85CALCULATOR_SPEC.md §6.1no
tpEffNvlink0.95CALCULATOR_SPEC.md §6.1no
tpEffUnified1CALCULATOR_SPEC.md §6.1no
defaultContextTokens8192T08 default (not specified in CALCULATOR_SPEC.md)no
decodeEff0.7CALCULATOR_SPEC.md §6.2yes
prefillEff0.5CALCULATOR_SPEC.md §6.2yes
batchEffK0.04CALCULATOR_SPEC.md §6.2yes
utilisation0.3CALCULATOR_SPEC.md §6.2no
maxUtilisation0.95CALCULATOR_SPEC.md §6.2no
sparesPct0.1CALCULATOR_SPEC.md §6.3no
residualPct0.2CALCULATOR_SPEC.md §6.3no
loadFactor0.8CALCULATOR_SPEC.md §6.3yes
horizonMonths36CALCULATOR_SPEC.md §6.3no
labourHours2CALCULATOR_SPEC.md §6.3no
labourRate60CALCULATOR_SPEC.md §6.3no
hostingMonth0CALCULATOR_SPEC.md §6.3no
softwareMonth0CALCULATOR_SPEC.md §6.3no
electricityUsdKwh0.1419src/data/electricity_defaults.json (US); EIA Electric Power Monthly via commercialenergyadvisors.comno
latencyTargetTps20CALCULATOR_SPEC.md §3no
tokenizerSliceblendedCALCULATOR_SPEC.md §7yes