Live per-token prices for every hosted AI model on OpenRouter, ppq.ai, kie.ai, and Haven, grouped per model. Each row shows the cheapest marketplace for that model. Expand it to compare the others, then jump straight to the provider. Covers OpenAI, Anthropic, Google, Mistral, and 50+ other labs. Refreshed hourly.
One row per model with the cheapest current marketplace. Expand a model to compare OpenRouter, ppq.ai, kie.ai, and Haven side by side. Click a column header to sort.
| Buy | |||||||
|---|---|---|---|---|---|---|---|
Mistral: Mistral Nemo mistralai/mistral-nemo | OpenRouter | mistralai | 131K | $0.019 | $0.030 | $0.022 | Open |
inclusionAI: Ling 3.0 Flash inclusionai/ling-3.0-flash | OpenRouter | inclusionai | 262K | $0.021 | $0.063 | $0.034 | Open |
Sao10K: Llama 3 8B Lunaris sao10k/l3-lunaris-8b | OpenRouter | sao10k | 8K | $0.040 | $0.050 | $0.043 | Open |
IBM: Granite 4.0 Micro ibm-granite/granite-4.0-h-micro | OpenRouter | ibm-granite | 131K | $0.017 | $0.112 | $0.046 | Open |
Nex AGI: Nex-N2-Mini nex-agi/nex-n2-mini | OpenRouter | nex-agi | 262K | $0.025 | $0.100 | $0.048 | Open |
Upstage: Solar Pro 4 upstage/solar-pro4 | OpenRouter | upstage | 524K | $0.030 | $0.120 | $0.057 | Open |
Mistral: Mistral Small 3 mistralai/mistral-small-24b-instruct-2501 | OpenRouter | mistralai | 33K | $0.050 | $0.080 | $0.059 | Open |
Meta: Llama 3.1 8B Instruct meta-llama/llama-3.1-8b-instruct | OpenRouter | meta-llama | 131K | $0.050 | $0.080 | $0.059 | Open |
MythoMax 13B gryphe/mythomax-l2-13b | OpenRouter | gryphe | 8K | $0.060 | $0.060 | $0.060 | Open |
Qwen: Qwen3.7 Flash qwen/qwen3.7-flash | OpenRouter | qwen | 1.0M | $0.030 | $0.130 | $0.060 | Open |
OpenAI: gpt-oss-20b openai/gpt-oss-20b | OpenRouter | openai | 131K | $0.030 | $0.130 | $0.060 | Open |
DeepSeek V4 Flash Latest ~deepseek/deepseek-v4-flash-latest | OpenRouter | ~deepseek | 1.3M | $0.050 | $0.100 | $0.065 | Open |
Google: Gemma 3 4B google/gemma-3-4b-it | OpenRouter | 131K | $0.050 | $0.100 | $0.065 | Open | |
Amazon: Nova Micro 1.0 amazon/nova-micro-v1 | OpenRouter | amazon | 128K | $0.035 | $0.140 | $0.067 | Open |
Cohere: Command R7B (12-2024) cohere/command-r7b-12-2024 | OpenRouter | cohere | 128K | $0.037 | $0.150 | $0.071 | Open |
Inception: Mercury 2.5 Preview inception/mercury-2.5-preview | OpenRouter | inception | 260K | $0.040 | $0.150 | $0.073 | Open |
OpenAI: gpt-oss-120b openai/gpt-oss-120b | OpenRouter | openai | 131K | $0.037 | $0.170 | $0.077 | Open |
OpenAI: GPT-5 Nano (batch) openai/gpt-5-nano:batch | OpenRouter | openai | 400K | $0.025 | $0.200 | $0.077 | Open |
Poolside: Laguna XS 2.1 poolside/laguna-xs-2.1 | OpenRouter | poolside | 262K | $0.060 | $0.120 | $0.078 | Open |
Meta: Llama 3.2 1B Instruct meta-llama/llama-3.2-1b-instruct | OpenRouter | meta-llama | 60K | $0.027 | $0.201 | $0.079 | Open |
Google: Gemma 3 12B google/gemma-3-12b-it | OpenRouter | 131K | $0.050 | $0.150 | $0.080 | Open | |
Tencent: Hy-MT2-1.8B tencent/hy-mt2-1.8b | OpenRouter | tencent | 8K | $0.044 | $0.177 | $0.084 | Open |
Microsoft: Phi 4 microsoft/phi-4 | OpenRouter | microsoft | 16K | $0.070 | $0.140 | $0.091 | Open |
Qwen: Qwen3 30B A3B Instruct 2507 qwen/qwen3-30b-a3b-instruct-2507 | OpenRouter | qwen | 262K | $0.048 | $0.193 | $0.092 | Open |
NVIDIA: Nemotron 3 Nano 30B A3B nvidia/nemotron-3-nano-30b-a3b | OpenRouter | nvidia | 262K | $0.050 | $0.200 | $0.095 | Open |
OpenAI: gpt-oss-20b (batch) openai/gpt-oss-20b:batch | OpenRouter | openai | 131K | $0.050 | $0.200 | $0.095 | Open |
Gemini 2.5 Flash Lite google/gemini-2.5-flash-lite | ppq.ai | 1.0M | $0.050 | $0.200 | $0.095 | Open | |
Google: Gemini 2.5 Flash Lite (batch) google/gemini-2.5-flash-lite:batch | OpenRouter | 1.0M | $0.050 | $0.200 | $0.095 | Open | |
OpenAI: GPT-4.1 Nano (batch) openai/gpt-4.1-nano:batch | OpenRouter | openai | 1.0M | $0.050 | $0.200 | $0.095 | Open |
inclusionAI: Ling 3.0 Flash Fin inclusionai/ling-3.0-flash-fin | OpenRouter | inclusionai | 262K | $0.060 | $0.180 | $0.096 | Open |
DeepSeek: DeepSeek V4 Flash 0731 deepseek/deepseek-v4-flash-0731 | OpenRouter | deepseek | 1.3M | $0.065 | $0.180 | $0.100 | Open |
Reka Edge rekaai/reka-edge | OpenRouter | rekaai | 16K | $0.100 | $0.100 | $0.100 | Open |
Mistral: Ministral 3 3B 2512 mistralai/ministral-3b-2512 | OpenRouter | mistralai | 131K | $0.100 | $0.100 | $0.100 | Open |
Mistral: Mistral Small 3.2 24B mistralai/mistral-small-3.2-24b-instruct | OpenRouter | mistralai | 131K | $0.075 | $0.200 | $0.113 | Open |
Amazon: Nova Lite 1.0 amazon/nova-lite-v1 | OpenRouter | amazon | 300K | $0.060 | $0.240 | $0.114 | Open |
DeepSeek: DeepSeek V4 Flash 0423 deepseek/deepseek-v4-flash | OpenRouter | deepseek | 1.0M | $0.088 | $0.176 | $0.114 | Open |
IBM: Granite 4.2 8B ibm-granite/granite-4.2-8b | OpenRouter | ibm-granite | 131K | $0.100 | $0.150 | $0.115 | Open |
Qwen: Qwen3.5-9B qwen/qwen3.5-9b | OpenRouter | qwen | 262K | $0.100 | $0.150 | $0.115 | Open |
NVIDIA: Nemotron 3.5 Lightning nvidia/nemotron-3.5-lightning | OpenRouter | nvidia | 262K | $0.080 | $0.200 | $0.116 | Open |
Poolside: Laguna S 2.1 poolside/laguna-s-2.1 | OpenRouter | poolside | 1.0M | $0.090 | $0.180 | $0.117 | Open |
Qwen: Qwen3.5-Flash qwen/qwen3.5-flash-02-23 | OpenRouter | qwen | 1.0M | $0.065 | $0.260 | $0.124 | Open |
Z.ai: GLM Flash Latest ~z-ai/glm-flash-latest | OpenRouter | ~z-ai | 1.3M | $0.075 | $0.250 | $0.128 | Open |
Z.ai: GLM 5.3 Flash z-ai/glm-5.3-flash | OpenRouter | z-ai | 1.3M | $0.075 | $0.250 | $0.128 | Open |
Meta: Muse Spark 1.3 Contributor meta/muse-spark-1.3-contributor | OpenRouter | meta | 1.0M | $0.100 | $0.200 | $0.130 | Open |
Meta: Muse Spark 1.2 Contributor meta/muse-spark-1.2-contributor | OpenRouter | meta | 1.0M | $0.100 | $0.200 | $0.130 | Open |
ByteDance: UI-TARS 7B bytedance/ui-tars-1.5-7b | OpenRouter | bytedance | 128K | $0.100 | $0.200 | $0.130 | Open |
Reka Flash 3 rekaai/reka-flash-3 | OpenRouter | rekaai | 66K | $0.100 | $0.200 | $0.130 | Open |
Qwen: Qwen2.5 7B Instruct qwen/qwen-2.5-7b-instruct | OpenRouter | qwen | 33K | $0.100 | $0.200 | $0.130 | Open |
Qwen: Qwen3 Coder 30B A3B Instruct qwen/qwen3-coder-30b-a3b-instruct | OpenRouter | qwen | 262K | $0.070 | $0.280 | $0.133 | Open |
Meta: Llama 3.2 3B Instruct meta-llama/llama-3.2-3b-instruct | OpenRouter | meta-llama | 131K | $0.050 | $0.330 | $0.134 | Open |
Weighted at 70% input, 30% output tokens. Adjust the mix below.
AI API pricing is the per-token cost that hosted model providers charge for sending text to a model and receiving text back. Input tokens are the prompt; output tokens are the completion. Most labs publish two rates per model, quoted in US dollars per one million tokens, and update them several times a year.
This page reads the public OpenRouter, ppq.ai, kie.ai, and Haven model feeds live and shows the current input price, output price, and a blended price for every hosted AI text model the four marketplaces expose. The combined feed covers OpenAI (GPT-5, GPT-5 mini, o1, o3), Anthropic (Claude 4.6 Sonnet, Claude 4.6 Haiku, Claude Opus), Google (Gemini 3.1 Pro, Gemini 3.5 Flash), Meta Llama, Mistral, DeepSeek, Qwen, and roughly 50 other providers.
Prices update hourly. Listings are grouped per model. The row you see first is the cheapest marketplace for that model, and the "more offers" button opens the other marketplaces underneath it. Every row has a button that takes you to that marketplace. The blended column is a weighted average at a 70 percent input, 30 percent output token mix, which approximates typical chat and coding workloads. Use the blend selector to reweight for input-heavy retrieval pipelines or output-heavy generation tasks.

Run the Numbers
Take any model in the table and compare its API cost against buying a GPU and renting one in the cloud. The decision tool plugs the live API price in for you.
Open the Decision Tool
Find the break-even point between local hardware and cloud API spend.

Live cloud GPU rental rates across RunPod and the Vast.ai marketplace.

Every open and closed model we track, ranked by benchmark score and hardware fit.

Spec a full local AI rig matched to your budget, with curated parts and pre-built options.
See whether the APIs on this list are up before you send traffic.
Straight answers to the questions we hear most often.
Can't find what you're looking for? Send us a message
The live table refreshes hourly. We call the public OpenRouter, ppq.ai, kie.ai, and Haven APIs directly and cache the result for one hour so a traffic spike never overwhelms any of the upstreams. Each marketplace publishes its own rate per model, so the numbers here are what that marketplace charges, without scraping.
OpenRouter, ppq.ai, kie.ai, and Haven all resell the same underlying models, but each one applies its own markup or discount, so the per-token price for GPT-5, Claude, or Gemini can differ between them. We match their different model ids to one key and show one row per model: the cheapest marketplace. Click "more offers" to see the other marketplaces for that model, and use the Buy button to open the listing on the marketplace itself.
Every major AI lab charges separately for tokens you send in (the prompt) and tokens the model writes out (the completion). Output tokens are usually two to five times more expensive because generation is the compute-heavy step. The blended column shows a weighted average at the input/output mix you pick.
A single per-million-token rate that combines input and output cost at a chosen ratio. The default is 70 percent input, 30 percent output, which roughly matches typical chat and coding workloads where prompts are longer than answers. Switch the blend selector to 50/50 or 30/70 to reweight for your traffic.
Free-tier and promotional models on OpenRouter, ppq.ai, kie.ai, or Haven return $0 for both input and output. We surface them when the Free filter is on, but the sync job that copies prices into our reference model pages skips them on purpose so a temporary promotion never overwrites a real listed price.
A daily background job pulls the OpenRouter feed and writes the input and output prices onto every reference model in our directory that we have linked to an OpenRouter ID. That means model detail pages and the ROI and decision calculators read the same live numbers without manual edits. The ppq.ai, kie.ai, and Haven rows in this table are shown for marketplace comparison only and do not write back into the directory.
For internal modelling, yes: the figures are the same per-million-token rates the providers publish. For anything you bill a client on, treat this tracker as a "current as of the last refresh" view and confirm against the provider invoice or pricing page. Prices change without notice on the provider side.