Microsoft: Phi 4 vs Llama 3.3 Nemotron Super 49B v1.5

Llama 3.3 Nemotron Super 49B v1.5 is cheaper than Microsoft: Phi 4 by about 100% on a blended (75% input / 25% output) basis. Microsoft: Phi 4 charges $0.070000 / $0.140000 per 1M input/output tokens; Llama 3.3 Nemotron Super 49B v1.5 charges $0.000000 / $0.000000.

FieldMicrosoft: Phi 4Llama 3.3 Nemotron Super 49B v1.5
Input $/Mtok$0.070000$0.000000
Output $/Mtok$0.140000$0.000000
Cache read $/Mtok––
Cache write $/Mtok––
Input $/Mcharn/a0.000000
Output $/Mcharn/a0.000000
Context window16384131072
Tokenizer–llama3

Price history

Same workload, monthly cost

Illustrative monthly cost at 10M tokens/month, split per workload preset. Not a substitute for the calculator, which uses your own volume, cache and batch assumptions.

WorkloadMicrosoft: Phi 4Llama 3.3 Nemotron Super 49B v1.5
Coding agent$0.74$0.00
RAG chat$0.81$0.00
Summarisation (batch)$0.77$0.00
Classification$0.71$0.00
Chat assistant$0.98$0.00
Content generation$1.26$0.00

Cite this page

Microsoft: Phi 4 vs Llama 3.3 Nemotron Super 49B v1.5: $0.070000/$0.140000 vs $0.000000/$0.000000 per 1M tokens, as of 2026-10-04 (Token Barometer, https://tokenbarometer.com/compare/microsoft_phi-4-vs-nvidia_nemotron-super-49b-v1).

Microsoft: Phi 4 page | Llama 3.3 Nemotron Super 49B v1.5 page | Open in calculator