KI‑Modell‑Intelligenz

Preise · 2026-09-29

LLM‑Preise

Tokenkosten bei großen Anbietern, umgerechnet in USD pro 1 Mio. Tokens.

Cheapest LLM APIs

AI APIs ranked by input + output token cost.

OpenAI API Pricing

All OpenAI model prices in one table — GPT-5, GPT-5 Mini, embeddings and more.

Anthropic Claude Pricing

All Anthropic Claude prices — Opus, Sonnet, Haiku and prompt caching costs.

Über diese Liste

Almost every commercial LLM provider charges separately for input tokens (the prompt you send) and output tokens (the response you get back). Output tokens typically cost 3–5× more than input tokens because generation is autoregressive — each token depends on the previous one and cannot be batched as efficiently.

The table below normalises all prices to USD per 1 million tokens. This is the industry-standard unit; per-1K or per-token rates are easy to misread by 1000×. The "Total" column sums input + output for a quick apples-to-apples comparison, but your real cost depends heavily on your input/output ratio.

What the table does NOT show

  • Prompt caching discounts — Anthropic and OpenAI offer cache_read rates 5–10× cheaper than standard input. If you reuse a long system prompt, caching dominates total cost.
  • Tiered pricing above 200K tokens — Google and Anthropic charge a premium for very long inputs. Each model detail page shows the >200K tier when applicable.
  • Reasoning tokens — thinking models (o-series, Claude Extended Thinking, DeepSeek R1) bill internal reasoning at output rates, often 2–3× the visible answer length.
  • Volume discounts and prepaid credits — most providers offer 10–30% off at scale. These are not reflected here.
  • Audio / image surcharges — multimodal inputs have separate per-image or per-second rates.

How to use this page

  1. Find the cheapest model that meets your context window and capability requirements (tool calling, structured output, vision).
  2. Open its detail page to check cache pricing, output limits and provider availability.
  3. Use the pricing calculator to estimate monthly cost for your actual token volume.
  4. Compare 2–3 finalists side-by-side on a comparison page.

Prices are refreshed daily. Models showing "Unknown" do not publish a public per-token rate — this usually means enterprise-only or invite-gated access. We deliberately do not show them as $0 to avoid misleading rankings.

#ModelVendorInput / 1MOutput / 1MTotalContext
1Voxtral Small 24B 2507Mistral$0.002$0.002$0.00533K
2All-MiniLM-L6-v2digitalocean$0.009Unknown$0.009256
3Multi-QA-mpnet-base-dot-v1digitalocean$0.009Unknown$0.009512
4Qwen3 Embedding 8BAlibaba (Qwen)$0.010Unknown$0.01033K
5Qwen3 Embedding 4BAlibaba (Qwen)$0.010Unknown$0.01033K
6BGE Reranker v2 M3digitalocean$0.010Unknown$0.0108K
7Llama 3.2 1B InstructMeta$0.010$0.010$0.02060K
8Qwen3 Embedding 0.6BAlibaba (Qwen)$0.010$0.010$0.02033K
9Prompt Guard 2 86MMeta$0.010$0.010$0.020512
10Llama Prompt Guard 2 22MMeta$0.010$0.010$0.020512
11text-embedding-3-smallOpenAI$0.020Unknown$0.0208K
12text-embedding-3-smallazure$0.020Unknown$0.0208K
13text-embedding-3-smallazure-cognitive-services$0.020Unknown$0.0208K
14text-embedding-3-smallsap-ai-core$0.020Unknown$0.0208K
15BGE M3digitalocean$0.020Unknown$0.0208K
16E5 Large v2digitalocean$0.020Unknown$0.020512
17GLiNER 2.5 Basepioneer$0.030Unknown$0.0304K
18Llama 3.2 3B InstructMeta$0.020$0.020$0.040131K
19PaddleOCR-VLnovita-ai$0.020$0.020$0.04016K
20Llama-3.1-8B-InstructMeta$0.020$0.030$0.050131K
21Mistral NemoMistral$0.020$0.030$0.05016K
22Nomic Embed Text v1.5tinfoil$0.050Unknown$0.0508K
23Meta Llama 3.1 8B Instruct TurboMeta$0.020$0.030$0.050128K
24DeepSeek OCR 2DeepSeek$0.030$0.030$0.0608K
25nvidia--llama-3.2-nv-embedqa-1bMeta$0.070Unknown$0.0708K
26Llama 3.2 3B Instruct (NovitaAI)novita$0.030$0.050$0.08033K
27Llama 3 8B InstructMeta$0.040$0.040$0.0808K
28Ministral 3Bazure$0.040$0.040$0.080128K
29Ministral 3Bazure-cognitive-services$0.040$0.040$0.080128K
30Ministral 3B (latest)Mistral$0.040$0.040$0.080128K
31Ling 3.0 Flash VLopenrouter$0.021$0.062$0.083262K
32Ling 3.0 Flashopenrouter$0.021$0.063$0.084262K
33Ling 3.0 Flashvercel$0.021$0.063$0.084256K
34Llama 3 8B LunarisMeta$0.040$0.050$0.0908K
35text-embedding-3-largesap-ai-core$0.090Unknown$0.0908K
36GTE Large (v1.5)digitalocean$0.090Unknown$0.0908K
37Mistral EmbedMistral$0.100Unknown$0.1008K
38text-embedding-ada-002OpenAI$0.100Unknown$0.1008K
39L3 8B Stheno V3.2novita-ai$0.050$0.050$0.1008K
40Sao10k L3 8B Lunaris novita-ai$0.050$0.050$0.1008K
41text-embedding-ada-002azure$0.100Unknown$0.1008K
42text-embedding-ada-002azure-cognitive-services$0.100Unknown$0.1008K
43DeepSeek V4 Flash 0731DeepSeek$0.035$0.070$0.1051.31M
44gpt-oss-20bOpenAI$0.018$0.090$0.108131K
45Llama 3.2 11B Vision InstructMeta$0.055$0.055$0.110128K
46Llama Guard 3 8BMeta$0.055$0.055$0.110131K
47Qwen3.5 4BAlibaba (Qwen)$0.040$0.070$0.110262K
48Gemma 3 4B ITGoogle$0.040$0.080$0.120131K
49GPT OSS 20B (FlexAI)edenai$0.020$0.100$0.120131K
50Sarvam 30Bfastrouter$0.020$0.100$0.120128K
51Nex AGI: Nex-N2.5-Minikilo$0.025$0.100$0.125262K
52Nex-N2.5-Miniopenrouter$0.025$0.100$0.125262K
53Granite 4.0 H Microcloudflare-workers-ai$0.017$0.112$0.129131K
54IBM: Granite 4.0 Microkilo$0.017$0.112$0.129131K
55Granite 4.0 Microopenrouter$0.017$0.112$0.129131K
56Mistral Small 3Mistral$0.050$0.080$0.13033K
57Llama 3.1 8BMeta$0.050$0.080$0.130131K
58text-embedding-3-largeOpenAI$0.130Unknown$0.1308K
59text-embedding-3-largeazure$0.130Unknown$0.1308K
60text-embedding-3-largeazure-cognitive-services$0.130Unknown$0.1308K
61amazon--nova-microsap-ai-core$0.030$0.100$0.130128K
62baichuan-m2-32bnovita-ai$0.070$0.070$0.140131K
63Model Routerazure$0.140Unknown$0.140200K
64Model Routerazure-cognitive-services$0.140Unknown$0.140200K
65amazon--titan-embed-textsap-ai-core$0.140Unknown$0.1408K
66Gemma 3 12B ITGoogle$0.050$0.100$0.150131K
67Gemini Embedding 001Google$0.150Unknown$0.1502K
68LFM2 24B A2Bpioneer$0.030$0.120$0.15033K
69LFM2-24B-A2Btogetherai$0.030$0.120$0.15033K
70Mellum2 12B A2.5Bwandb$0.050$0.100$0.150131K
71Granite 4.1 8Bwandb$0.050$0.100$0.150131K
72Qwen3.7 FlashAlibaba (Qwen)$0.030$0.130$0.1601M
73DeepSeek R1 Distill Llama 70BMeta$0.030$0.130$0.160128K
74GPT OSS 120Bllmgateway$0.032$0.140$0.172131K
75Amazon: Nova Micro 1.0kilo$0.035$0.140$0.175128K
76Nova Micro 1.0openrouter$0.035$0.140$0.175128K
77Nova Microvercel$0.035$0.140$0.175128K
78Nova Microedenai$0.035$0.140$0.175128K
79Nova Micro (US)edenai$0.035$0.140$0.175128K
80Nova Microamazon-bedrock$0.035$0.140$0.175128K
81Nova Micro (US)amazon-bedrock$0.035$0.140$0.175128K
82Phi 4 Multimodalnano-gpt$0.070$0.110$0.180128K
83Manta Mini 1.0nano-gpt$0.020$0.160$0.1808K
84Manta Flash 1.0nano-gpt$0.020$0.160$0.18016K
85Schematron V2 Turbonano-gpt$0.030$0.150$0.180128K
86Inference.net: Schematron V2 Turbokilo$0.030$0.150$0.180128K
87Mythomax L2 13Bnovita-ai$0.090$0.090$0.1804K
88Laguna XS 2.1openrouter$0.060$0.120$0.180262K
89Schematron V2 Turboopenrouter$0.030$0.150$0.180128K
90Schematron V2 Turbovercel$0.030$0.150$0.180128K
91Nova Micro (APAC)amazon-bedrock$0.037$0.148$0.185128K
92Command R7BCohere$0.037$0.150$0.188128K
93Command R7B ArabicCohere$0.037$0.150$0.188128K
94Qwen3.5 9BAlibaba (Qwen)$0.040$0.150$0.190262K
95Mercury 2.5inception$0.040$0.150$0.190260K
96Mercury 2.5 Previewinception$0.040$0.150$0.190260K
97MythoMax 13Bkilo$0.080$0.110$0.1904K
98MythoMax 13Bopenrouter$0.080$0.110$0.1908K
99Trinity Miniclarifai$0.045$0.150$0.195131K
100nova-micro-v1cortecs$0.040$0.159$0.199128K

Showing top 100 of 1464. Use the full directory to see the rest.

Frequently asked questions

Why are output tokens more expensive than input tokens?

Generation is autoregressive — each output token requires a full forward pass conditioned on all previous tokens, which cannot be batched as efficiently as reading input tokens in parallel. The 3–5× premium most providers charge reflects this compute asymmetry.

What does 'per 1M tokens' mean in practice?

One million tokens is roughly 750,000 English words or 3,000 pages of standard text. A typical chatbot request uses 1,000–5,000 input tokens and 200–1,000 output tokens, so 1M tokens represents hundreds to thousands of requests depending on your workload.

Why are some models showing 'Unknown' instead of a price?

We deliberately do not coerce missing data to $0. 'Unknown' means the provider does not publish a public per-token rate — often models behind enterprise sales or invite-only access. Treating Unknown as free would push paid-but-unpriced models to the top of every cheap list.

How often do these prices change?

Vendor list-price moves are typically picked up within hours of an announcement, and our pipeline re-syncs daily. Each change is written to /changelog so you can audit historical pricing over time.

Does this include prompt caching or batch discounts?

No. The table shows standard headline rates only. Prompt caching (Anthropic, OpenAI) can reduce input cost by 50–90%. Batch API discounts (OpenAI) offer ~50% off for non-real-time workloads. Both are shown on each model's detail page.

Last updated:

Prices in USD per 1M tokens. Unknown means the provider does not publish per-token pricing.

Pricing and capabilities are refreshed daily and reconciled against each provider's official documentation. Always verify critical production decisions with the provider directly.