AI 模型情报

能力 · 2026-08-14

支持超长上下文的 AI 模型

对比支持 200K tokens 及以上上下文窗口的 AI 模型 —— 长文档与大规模代码场景。

这是什么?

  • 长上下文 LLM 可一次接受 200K tokens 或更长的输入 —— 足以装下整本书、多文件代码库或数小时转写稿。
  • 部分模型可扩展到 1M、2M 甚至 10M tokens 的上下文。

为什么重要

  • 长上下文是 RAG 的替代或补充 —— 你可以直接粘贴全部内容,而不只检索片段。
  • 注意:有效召回会随输入变长而下降,且按百万 token 计价会让长提示很贵。
  • 部分厂商在 200K 以上有阶梯价 —— 详见各模型详情页的 >200K 费率。

693 个模型支持此能力

模型厂商输入 / 1M输出 / 1M上下文服务商
Ling-2.6-flashopenrouter$0.010$0.030262K1
Ling-3.0-flashopenrouter$0.021$0.063262K1
Nex N2 Mininano-gpt$0.025$0.100262K1
Nex-N2-Miniopenrouter$0.025$0.100262K1
Nex AGI: Nex-N2-Minikilo$0.025$0.100262K1
Model Routerazure-cognitive-services$0.140Unknown200K1
Model Routerazure$0.140Unknown200K1
Qwen3.7 FlashAlibaba (Qwen)$0.030$0.1181M9
Laguna XS 2.1openrouter$0.060$0.120262K1
Qwen3.5 9BAlibaba (Qwen)$0.040$0.150262K22
Greg 1 Minicrof$0.070$0.150229K1
ministral-3b-2512cortecs$0.111$0.111256K1
Google Gemma 3 27B InstructGoogle$0.080$0.160203K11
Ling 3.0 Flashvercel$0.060$0.180256K1
InclusionAI Ling 3.0 Flashllmgateway$0.060$0.180262K1
Ling-3.0-flashkilo$0.060$0.180262K1
Qwen3 30B A3B Instruct 2507Alibaba (Qwen)$0.048$0.193262K12
Qwen TurboAlibaba (Qwen)$0.050$0.2001M5
Nemotron 3.5 Lightning 30B A3BNVIDIA$0.050$0.2001M3
DeepSeek V4 Flash 0731DeepSeek$0.080$0.1801.05M27
Gemini 2.0 Flash LiteGoogle$0.052$0.2101.05M3
Laguna S 2.1openrouter$0.090$0.1801.05M1
Hy3 previewopenrouter$0.063$0.210262K1
Ling 3.0 Flashnano-gpt$0.075$0.220262K1
Ling 3.0 Flash Thinkingnano-gpt$0.075$0.220262K1
nvidia-nemotron-3-nano-omniNVIDIA$0.059$0.237300K4
Amazon Nova Lite 1.0nano-gpt$0.059$0.238300K1
Ministral 3 8B 2512Mistral$0.150$0.150262K3
Nemotron 3 Nano 30B A3BNVIDIA$0.060$0.240262K3
Nova Lite 1.0openrouter$0.060$0.240300K1
Nova Liteamazon-bedrock$0.060$0.240300K1
Nova Litevercel$0.060$0.240300K1
Ministral 8Bllmgateway$0.150$0.150262K1
Amazon: Nova Lite 1.0kilo$0.060$0.240300K1
Muse Spark 1.2 ContributorMeta$0.100$0.2001.05M3
Laguna S 2.1 Thinkingnano-gpt$0.100$0.2001.05M1
Laguna S 2.1nano-gpt$0.100$0.2001.05M1
Laguna S 2.1vercel$0.100$0.2001M1
Poolside: Laguna S 2.1kilo$0.100$0.2001.05M1
Poolside: Laguna XS 2.1kilo$0.100$0.200262K1
Laguna S 2.1pioneer$0.100$0.2001M1
Qwen3.5 FlashAlibaba (Qwen)$0.029$0.2871M7
Tencent Hy3nano-gpt$0.066$0.260262K1
Hy3 previewsiliconflow$0.066$0.260262K1
DeepSeek V4 Flash LatestDeepSeek$0.080$0.2521.05M3
ministral-8b-2512cortecs$0.167$0.167256K1
GLM-4.7-FlashZ.AI / Zhipu$0.040$0.300200K20
Mistral Small 3.2 24BMistral$0.094$0.250256K3
nova-lite-v1cortecs$0.069$0.275300K1
Nemotron 3.5 Lightning 30B A3BNVIDIA$0.100$0.250262K4
Qwen LongAlibaba (Qwen)$0.072$0.28710M2
Llama Guard 4 12BMeta$0.180$0.1801.05M3
Seed 1.6 Flash (250715)llmgateway$0.070$0.300256K1
Seed 1.6 Flashopenrouter$0.075$0.300262K1
ByteDance Seed: Seed 1.6 Flashkilo$0.075$0.300262K1
Llama 4 ScoutMeta$0.080$0.3001.31M5
Gemma 4 26B A4B ITGoogle$0.060$0.330262K18
cosmos3-super-reasonercortecs$0.099$0.296256K1
Gemma 4 31B ITGoogle$0.102$0.297262K33
Step 3.5 FlashStepFun$0.100$0.300256K11

显示前 60 / 共 693 项。 完整目录 进一步筛选。

Frequently asked questions

How many AI models support 200K+ 上下文窗口?

693 canonical models in our database currently support 200K+ 上下文窗口. The list is regenerated on every data refresh, so it always reflects the latest releases tracked in our catalogue.

What is the cheapest model with 200K+ 上下文窗口?

Ling-2.6-flash from openrouter is currently the lowest-priced option, at $0.010 per 1M input tokens and $0.030 per 1M output tokens. The full table above is sorted price-ascending.

Which model with 200K+ 上下文窗口 has the largest context window?

Qwen Long (Alibaba (Qwen)) leads on context at 10M tokens. This may matter if you also need long-document understanding alongside 200K+ 上下文窗口.

Which models are available on the most providers?

Production-readiness usually correlates with how many independent providers host the same weights. The top three by provider count are: GLM-5.2 (75), Kimi K2.6 (63), DeepSeek V4 Pro (55).

How is 200K+ 上下文窗口 different from a regular LLM?

Long-context models accept ≥ 200K input tokens — enough for entire books, codebases or hours of transcripts in one prompt. Effective recall and per-token pricing both degrade with input length, so 'big context' is not always the right choice over RAG.

How often is this list updated?

Daily. Our data pipeline syncs once a day, regenerates the canonical model list, and rebuilds these pages so newly released models appear within 24 hours.

最近更新:

Prices in USD per 1M tokens. Unknown means the provider does not publish per-token pricing.

Pricing and capabilities are refreshed daily and reconciled against each provider's official documentation. Always verify critical production decisions with the provider directly.