Methodology

How Token Barometer sources, prioritises, confidence-scores and normalises LLM API prices, and how its cloud-vs-local calculator works. See About for the source list and licence.

Sources

Prices are collected daily by a scheduled Workflow from three kinds of source: official vendor pricing pages (parsed with cheerio, one parser per vendor, anchored on stable headings and column names rather than row position), JSON aggregators (OpenRouter, BerriAI's LiteLLM price list, models.dev), and a Haiku-class LLM fallback used only when a page's structure changes and the deterministic parser breaks or returns zero prices.

Every source has a priority (lower number wins on conflict), a kind (api, json, html, backfill or manual) and is enabled or disabled independently. A handful of official pages are disabled because they render their pricing table client-side with no server-rendered fallback, or point at a page with no extractable structured price at the time they were checked (listed under Limitations below); those vendors' prices still reach this site through the JSON aggregators, at medium rather than high confidence.

Historical prices before this project started tracking a vendor come from a one-time backfill that walks that source's own git history (LiteLLM's price file, models.dev's per-model data files) and parses each past revision with the same parser used for live data, then attaches the result to whichever offering a live refresh has already discovered. A backfill never invents a model, host or offering.

Priority, conflicts and the canonical price

Each offering (a model as served by one host, in one quantisation/variant, optionally one region) can be priced by more than one source at once, for example both the vendor's own pricing page and an aggregator. The canonical price shown everywhere on this site is the open price (still in effect, not superseded by a later one) from the highest-priority source that currently has one for that offering, among sources whose last successful run finished within the last 14 days.

When two sources of medium or higher confidence disagree by more than 100 basis points (1%) on the same offering, both prices are flagged as a conflict and held for a human to check, rather than one silently overwriting the other without comment.

When a single source's own new price differs from its last known price by more than a configurable threshold (50% by default) with no other corroborating change, the new claim is held in a review queue instead of being applied automatically: large, sudden price moves are more often a parsing break than a real price cut.

Confidence

Every stored price carries a confidence level: high (parsed deterministically from an official vendor page), medium (parsed deterministically from a third-party aggregator), or low (extracted by the LLM fallback from unstructured page text). A low-confidence price is never applied as a canonical price automatically; it always lands in the review queue first, and only becomes canonical once a human approves it.

The confidence badge shown next to a price in every table is this same field, read directly off the stored price record, not a separate estimate layered on top.

The blended price

Where a single number is needed to rank or compare models (sorting a table, computing a price change's magnitude, picking the "cheapest" offering), this site uses a blended rate of 75% input price plus 25% output price per 1M tokens. The weighting approximates a typical chat/agent workload, which reads more context than it generates; it is a ranking convenience only, never a substitute for the full input/output/cache/batch breakdown shown on every model and offering page.

A price change's percentage magnitude (shown on change pages and in the changes feed) is computed the same way: the blended rate before versus after.

Normalisation

All prices are stored as integer micro-USD per 1 million tokens (a decimal USD price multiplied by 1,000,000), parsed from each source's own string representation without ever passing a price through a floating-point division. This avoids rounding drift across the full history this project keeps.

Cache read/write pricing (including multi-tier cache TTLs where a vendor publishes them, such as Anthropic's 5-minute and 1-hour cache writes), batch-API discounts, long-context pricing tiers, and per-request or per-image pricing are each normalised into their own dedicated field where a source publishes them; anything that does not fit this common schema is kept verbatim in a per-offering "extras" field rather than discarded.

Tokenizer normalisation

Different model families tokenize the same text differently, so a USD-per-1M-tokens price is not directly comparable across tokenizer families without a common yardstick. This project measures characters-per-token for each tracked tokenizer family (o200k and cl100k via OpenAI's published tiktoken rank data; Llama 3, Qwen2, Mistral, DeepSeek and Gemma via each model family's own public Hugging Face tokenizer) against a fixed, roughly 1 MB text corpus: English prose, source code (Python, TypeScript, Rust), German, Croatian, Chinese and JSON, every file either public domain or self-authored for this project.

The measured ratio is published per corpus slice and as a single blended figure across all slices; the price table's USD-per-1M-characters view defaults to the blended figure and lets the slice be changed (plain prose vs. code skews the ratio noticeably for some families). A tokenizer family with no measured ratio (for example one needing a live API call this project does not make at build time) shows the native per-token price only, with no invented conversion.

Calculator formulas and defaults

The cloud-vs-local calculator answers one question: for a given monthly workload (input/output tokens, cache hit ratio, batch share, concurrency, an interactive latency target), what does a specific cloud API offering cost per month, and what would an equivalent open-weight model cost to run on specific local hardware, with every cost on both sides made explicit?

Cloud cost multiplies the workload's input, output, cache-read and cache-write token volumes by that offering's own priced rates, applying batch discounts and long-context tiers where published; nothing here is estimated, it is the offering's stored prices applied to the stated workload.

Local cost has three parts, each a named default the calculator shows and lets you override: hardware amortisation (purchase price plus a spares allowance, spread over a chosen horizon of 24, 36 or 48 months, minus an estimated residual value), electricity (a chosen USD/kWh tariff, defaulting to a small-business US rate from the EIA's Electric Power Monthly, applied to GPU and host power draw at load and at idle), and labour (an hourly rate times a monthly hours estimate for upkeep). Achievable throughput is derived from the hardware's published memory bandwidth and compute, the model's parameter count and quantisation, and a decode/prefill efficiency factor calibrated against reported real-world serving numbers, not measured on this site's own hardware; a measured tokens/second figure can be substituted for any of these at any time.

Every default value used in a given calculation, and its source, is listed in that calculation's own assumptions panel, not only here: hardware street prices are dated and re-verified periodically rather than being a one-time estimate.

Limitations

Coverage is not exhaustive. A small number of official vendor pricing pages render entirely client-side with no server-rendered fallback and are not parsed directly (their prices still reach this site through the JSON aggregators, at medium rather than high confidence); the full list of enabled and disabled sources, and why, is in this project's source documentation.

The review queue holds anything this pipeline is not confident enough to apply automatically: a brand-new model or host this project has not mapped to a known entry yet, a large unexplained price move, a disagreement between two sources, or an LLM-extracted price. An item sitting in the queue does not appear as a canonical price until a human resolves it, so a very new offering can take longer than one refresh cycle to show up.

Local-hardware cost estimates model throughput from published hardware specifications and a calibrated efficiency factor; they are not a benchmark of real serving software on real hardware, and will diverge from measured numbers for any specific inference stack, batching strategy or driver version. Hardware and electricity prices move and are dated at the time they were recorded, not live-quoted.

No quality or capability signal (benchmark scores, "price per intelligence point") is factored into any price, ranking or comparison on this site today; prices are compared as published, not adjusted for how good a model is. EUR and other non-USD display, rate limits, a deprecation calendar and per-host latency/throughput are tracked as future work, not built yet.