Возможность · 2026-09-28
ИИ-модели с длинным контекстом
Модели с окном контекста 200K токенов и более.
Что это?
- LLM с длинным контекстом принимают 200K токенов и более в одном промпте — целые книги, многофайловые репозитории или часы транскрипции.
- Некоторые модели масштабируются до 1M, 2M и более токенов контекста.
Зачем это важно
- Длинный контекст дополняет или заменяет RAG — можно вставить весь контент, а не только извлечённые фрагменты.
- Эффективный recall может снижаться с длиной, а длинные промпты дорожают при тарификации за миллион токенов.
- Некоторые провайдеры применяют ступенчатые тарифы выше 200K — см. страницу модели для деталей.
1027 моделей с этой возможностью
Показаны первые 60 из 1027. Перейдите к полному каталогу для дополнительной фильтрации.
Frequently asked questions
How many AI models support контекст 200K+?
1027 canonical models in our database currently support контекст 200K+. The list is regenerated on every data refresh, so it always reflects the latest releases tracked in our catalogue.
What is the cheapest model with контекст 200K+?
Ling 3.0 Flash VL from openrouter is currently the lowest-priced option, at $0.021 per 1M input tokens and $0.062 per 1M output tokens. The full table above is sorted price-ascending.
Which model with контекст 200K+ has the largest context window?
Qwen Long (Alibaba (Qwen)) leads on context at 10M tokens. This may matter if you also need long-document understanding alongside контекст 200K+.
Which models are available on the most providers?
Production-readiness usually correlates with how many independent providers host the same weights. The top three by provider count are: GLM-5.2 (89), Kimi K3 (75), DeepSeek V4 Pro (72).
How is контекст 200K+ different from a regular LLM?
Long-context models accept ≥ 200K input tokens — enough for entire books, codebases or hours of transcripts in one prompt. Effective recall and per-token pricing both degrade with input length, so 'big context' is not always the right choice over RAG.
How often is this list updated?
Daily. Our data pipeline syncs once a day, regenerates the canonical model list, and rebuilds these pages so newly released models appear within 24 hours.
Explore more
Top models with this capability
- Ling 3.0 Flash VL$0.02 in / $0.06 out
- Ling 3.0 Flash$0.02 in / $0.06 out
- Ling 3.0 Flash$0.02 in / $0.06 out
- DeepSeek V4 Flash 0731$0.04 in / $0.07 out
- Qwen3.5 4B$0.04 in / $0.07 out
Other capabilities
Best-of lists you might also want
Pricing comparisons
Vendors in this list
Последнее обновление:
Prices in USD per 1M tokens. Unknown means the provider does not publish per-token pricing.
Pricing and capabilities are refreshed daily and reconciled against each provider's official documentation. Always verify critical production decisions with the provider directly.