Offerings
Every active offering: a model as served by one host, with its current canonical price. Sort by clicking a column header; filter with the row below. Prices are USD per 1M tokens unless the unit toggle is set to per 1M characters (an approximation using a measured, blended tokens-per-character ratio).
Showing tier-tagged and vendor-direct models. Show all offerings, including the long tail
| Model | Vendor | Host | Variant | Input ↑ | Output | Cache read | Cache write | Batch in | Batch out | Context | Tokenizer | Latency / throughput soon | Last change | Confidence | Source |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Google: Gemini 2.5 Flash Lite | unknown | 0.100000 | 0.400000 | 0.010000 | 0.083333 | – | – | 1048576 | gemini | – | – | high | openrouter | ||
| Meta: Llama 4 Scout | meta | deepinfra | fp8 | 0.100000 | 0.300000 | – | – | – | – | 131072 | llama3 | – | – | high | openrouter |
| Meta: Llama 3.3 70B Instruct | meta | deepinfra | fp8 | 0.100000 | 0.320000 | – | – | – | – | 128000 | llama3 | – | – | high | openrouter |
| OpenAI: GPT-6 Luna | openai | azure | – | 0.100000 | 0.500000 | 0.010000 | 0.125000 | – | – | 1050000 | o200k | – | – | medium | litellm |
| OpenAI: GPT-6 Luna | openai | azure | – | 0.100000 | 0.500000 | 0.010000 | 0.125000 | – | – | 1050000 | o200k | – | – | medium | litellm |
| Mistral Small (latest) | mistral | azure | – | 0.100000 | 0.300000 | – | – | – | – | 256000 | mistral | – | – | medium | litellm |
| Meta: Llama 3.3 70B Instruct | meta | deepinfra | – | 0.100000 | 0.320000 | – | – | – | – | 128000 | llama3 | – | – | medium | litellm |
| Meta: Llama 4 Scout | meta | deepinfra | – | 0.100000 | 0.300000 | – | – | – | – | 131072 | llama3 | – | – | medium | litellm |
| Llama 3.3 Nemotron Super 49B v1.5 | nvidia | deepinfra | – | 0.100000 | 0.400000 | – | – | – | – | 131072 | llama3 | – | – | medium | litellm |
| Google: Gemini 2.5 Flash Lite | google-vertex | – | 0.100000 | 0.400000 | 0.010000 | – | 0.050000 | 0.200000 | 1048576 | gemini | – | – | medium | litellm | |
| Google: Gemini 2.5 Flash Lite | – | 0.100000 | 0.400000 | 0.010000 | – | 0.050000 | 0.200000 | 1048576 | gemini | – | – | high | official-google | ||
| OpenAI: GPT-6 Luna | openai | openai | – | 0.100000 | 0.500000 | 0.010000 | 0.125000 | – | – | 1050000 | o200k | – | – | high | official-openai |
| Llama 3.3 Nemotron Super 49B v1.5 | nvidia | nebius | – | 0.100000 | 0.400000 | – | – | – | – | 131072 | llama3 | – | – | medium | litellm |
| Mistral Small (latest) | mistral | google-vertex | – | 0.100000 | 0.300000 | – | – | – | – | 256000 | mistral | – | – | medium | litellm |
| MiniMax: MiniMax M1 | minimax | fireworks | – | 0.100000 | 0.100000 | – | – | – | – | 1000000 | – | – | – | medium | litellm |
| OpenAI: GPT-6 Luna | openai | perplexity | – | 0.100000 | 0.500000 | 0.010000 | – | – | – | 1050000 | o200k | – | – | medium | litellm |
| OpenAI: GPT-6 Luna | openai | aws-bedrock | – | 0.100000 | 0.500000 | 0.010000 | 0.125000 | – | – | 1050000 | o200k | – | – | medium | modelsdev |
| OpenAI: GPT-6 Luna | openai | openai | – | 0.100000 | 0.500000 | 0.010000 | 0.125000 | – | – | 1050000 | o200k | – | – | high | openrouter |
| OpenAI: GPT-6 Luna | openai | azure | – | 0.100000 | 0.500000 | 0.010000 | 0.125000 | – | – | 1050000 | o200k | – | – | high | openrouter |
| Google: Gemini 2.5 Flash Lite | google-vertex | – | 0.100000 | 0.400000 | 0.010000 | 0.083333 | – | – | 1048576 | gemini | – | – | high | openrouter | |
| Google: Gemini 2.5 Flash Lite | google-vertex | – | 0.100000 | 0.400000 | 0.010000 | 0.083333 | – | – | 1048576 | gemini | – | – | high | openrouter | |
| Google: Gemini 2.5 Flash Lite | – | 0.100000 | 0.400000 | 0.010000 | 0.083333 | – | – | 1048576 | gemini | – | – | high | openrouter | ||
| OpenAI: GPT-6 Luna | openai | azure | unknown | 0.110000 | 0.550000 | 0.011000 | 0.137500 | – | – | 1050000 | o200k | – | – | high | openrouter |
| OpenAI: GPT-6 Luna | openai | azure | unknown | 0.110000 | 0.550000 | 0.011000 | 0.137500 | – | – | 1050000 | o200k | – | – | high | openrouter |
| OpenAI: GPT-6 Luna | openai | aws-bedrock | unknown | 0.110000 | 0.550000 | 0.011000 | 0.137500 | – | – | 1050000 | o200k | – | – | high | openrouter |
| OpenAI: GPT-6 Luna | openai | aws-bedrock | – | 0.110000 | 0.550000 | 0.011000 | 0.137500 | – | – | 1050000 | o200k | – | – | medium | modelsdev |
| OpenAI: GPT-6 Luna | openai | aws-bedrock | – | 0.110000 | 0.550000 | 0.011000 | 0.137500 | – | – | 1050000 | o200k | – | – | medium | modelsdev |
| OpenAI: GPT-6 Luna | openai | azure | – | 0.110000 | 0.550000 | 0.011000 | 0.137500 | – | – | 1050000 | o200k | – | – | high | openrouter |
| OpenAI: GPT-6 Luna | openai | azure | – | 0.110000 | 0.550000 | 0.011000 | 0.137500 | – | – | 1050000 | o200k | – | – | high | openrouter |
| OpenAI: GPT-6 Luna | openai | aws-bedrock | – | 0.110000 | 0.550000 | 0.011000 | 0.137500 | – | – | 1050000 | o200k | – | – | high | openrouter |
| Qwen2.5-72B-Instruct | alibaba | hyperbolic | – | 0.120000 | 0.300000 | – | – | – | – | 131072 | qwen2 | – | – | medium | litellm |
| Meta: Llama 3.3 70B Instruct | meta | hyperbolic | – | 0.120000 | 0.300000 | – | – | – | – | 128000 | llama3 | – | – | medium | litellm |
| OpenAI: GPT-5 Mini | openai | openai | unknown | 0.125000 | 1.000000 | 0.012500 | – | – | – | 400000 | o200k | – | – | high | openrouter |
| Microsoft: Phi 4 | microsoft | azure | – | 0.125000 | 0.500000 | – | – | – | – | 16384 | – | – | – | medium | litellm |
| Microsoft: Phi 4 | microsoft | azure | – | 0.125000 | 0.500000 | – | – | – | – | 16384 | – | – | – | medium | modelsdev |
| OpenAI: GPT-5 Mini | openai | openai | – | 0.125000 | 1.000000 | 0.012500 | – | – | – | 400000 | o200k | – | – | high | openrouter |
| Meta: Llama 3.3 70B Instruct | meta | nebius | – | 0.130000 | 0.400000 | – | – | – | – | 128000 | llama3 | – | – | medium | litellm |
| Qwen2.5-72B-Instruct | alibaba | nebius | – | 0.130000 | 0.400000 | – | – | – | – | 131072 | qwen2 | – | – | medium | litellm |
| Meta: Llama 3.3 70B Instruct | meta | novita | – | 0.135000 | 0.400000 | – | – | – | – | 128000 | llama3 | – | – | high | openrouter |
| Meta: Llama 3.3 70B Instruct | meta | novita | – | 0.135000 | 0.400000 | – | – | – | – | 128000 | llama3 | – | – | medium | litellm |
| Google: Gemini 2.5 Flash | unknown | 0.150000 | 1.250000 | 0.015000 | 0.041667 | – | – | 1048576 | gemini | – | – | high | openrouter | ||
| Mistral Small (latest) | mistral | mistral | – | 0.150000 | 0.600000 | 0.015000 | – | – | – | 256000 | mistral | – | – | medium | litellm |
| Google: Gemini 2.5 Flash | – | 0.150000 | 1.250000 | 0.015000 | 0.041667 | – | – | 1048576 | gemini | – | – | high | openrouter | ||
| Google: Gemini 2.5 Flash Lite | unknown | 0.180000 | 0.720000 | 0.018000 | 0.150000 | – | – | 1048576 | gemini | – | – | high | openrouter | ||
| Meta: Llama 4 Scout | meta | novita | – | 0.180000 | 0.590000 | – | – | – | – | 131072 | llama3 | – | – | high | openrouter |
| Qwen: Qwen3 235B A22B | alibaba | deepinfra | – | 0.180000 | 0.540000 | – | – | – | – | 128000 | qwen2 | – | – | medium | litellm |
| Meta: Llama 4 Scout | meta | novita | – | 0.180000 | 0.590000 | – | – | – | – | 131072 | llama3 | – | – | medium | litellm |
| Meta: Llama 4 Scout | meta | together | – | 0.180000 | 0.590000 | – | – | – | – | 131072 | llama3 | – | – | medium | litellm |
| Google: Gemini 2.5 Flash Lite | – | 0.180000 | 0.720000 | 0.018000 | 0.150000 | – | – | 1048576 | gemini | – | – | high | openrouter | ||
| Meta: Llama 4 Maverick | meta | openrouter | – | 0.187500 | 0.652500 | – | – | – | – | 1048576 | llama3 | – | – | high | openrouter |