Intelligence des modèles d'IA

Capacité · 2026-09-29

Modèles d'IA avec entrée visuelle

Modèles acceptant des images en plus du texte — compréhension multimodale.

Qu'est-ce que c'est ?

  • Les modèles vision-langage acceptent des images en plus de (ou à la place du) texte.
  • La plupart répondent en texte — ce sont des LLMs multimodaux, pas des générateurs d'images.

Pourquoi c'est important

  • Cas d'usage : compréhension de documents (scans, PDF, captures d'écran), revue UI/code depuis des captures, Q&A sur photos produit, accessibilité (alt text), imagerie médicale/satellite.
  • La facturation inclut souvent un coût par image en plus du coût par token — consultez les tableaux d'offering de chaque fournisseur.

897 modèles avec cette capacité

ModèleÉditeurEntrée / 1MSortie / 1MContexteFournisseurs
PaddleOCR-VLnovita-ai$0.020$0.02016K1
DeepSeek OCR 2DeepSeek$0.030$0.0308K3
Ling 3.0 Flash VLopenrouter$0.021$0.062262K1
Llama 3.2 11B Vision InstructMeta$0.055$0.055128K4
Qwen3.5 4BAlibaba (Qwen)$0.040$0.070262K3
Gemma 3 4B ITGoogle$0.040$0.080131K10
Nex AGI: Nex-N2.5-Minikilo$0.025$0.100262K1
Nex-N2.5-Miniopenrouter$0.025$0.100262K1
Model Routerazure$0.140Unknown200K1
Model Routerazure-cognitive-services$0.140Unknown200K1
Gemma 3 12B ITGoogle$0.050$0.100131K12
Qwen3.7 FlashAlibaba (Qwen)$0.030$0.1301M13
Qwen3.5 9BAlibaba (Qwen)$0.040$0.150262K23
nova-micro-v1cortecs$0.040$0.159128K1
Ministral 3 3B 2512Mistral$0.100$0.100131K4
Nemotron 3.5 Content SafetyNVIDIA$0.050$0.150131K4
Gemini Embedding 2Google$0.200Unknown8K4
Agnes 3.0 Flashnano-gpt$0.050$0.150524K1
Space Bunny Alphanano-gpt$0.050$0.1501M1
Reka Edgekilo$0.100$0.10016K1
Reka Edgeopenrouter$0.100$0.10016K1
Ministral 3Bllmgateway$0.100$0.100131K1
Greg 1 Minicrof$0.070$0.150229K1
Gemma 3 27B ITGoogle$0.080$0.160131K17
Ling 3.0 Flash VLnano-gpt$0.060$0.180262K1
Ling 3.0 Flash VL (DeepInfra)llmgateway-providers$0.060$0.180131K1
Ling 3.0 Flash VLllmgateway$0.060$0.180131K1
GLM-4.6V-FlashZ.AI / Zhipu$0.022$0.218128K6
ministral-3b-2512cortecs$0.123$0.123256K1
Gemma 4 26B A4B ITGoogle$0.042$0.220262K22
Gemini-2.0-Flash-LiteGoogle$0.052$0.210990K3
inclusionAI: Ling 3.0 Flash VLkilo$0.075$0.220262K1
Ling 3.0 Flash VLvercel$0.075$0.220256K1
Ministral 3 8B 2512Mistral$0.150$0.150262K4
Gemma 4 12B InstructGoogle$0.050$0.250131K3
Amazon: Nova Lite 1.0kilo$0.060$0.240300K1
Nova Lite 1.0openrouter$0.060$0.240300K1
Nova Litevercel$0.060$0.240300K1
Nova Lite (US)edenai$0.060$0.240300K1
Nova Liteedenai$0.060$0.240300K1
Pixtral 12BMistral$0.150$0.150128K1
Nova Lite (US)amazon-bedrock$0.060$0.240300K1
Nova Liteamazon-bedrock$0.060$0.240300K1
Ministral 8Bllmgateway$0.150$0.150262K1
Muse Spark 1.2 ContributorMeta$0.100$0.2001.05M6
Muse Spark 1.3 ContributorMeta$0.100$0.2001.05M5
Muse Spark 1.3 Contributorbothub$0.100$0.2001.05M1
Muse Spark 1.2 Contributor (Meta Contributor)llmgateway-providers$0.100$0.2001.05M1
Muse Spark 1.3 Contributor (Meta Contributor)llmgateway-providers$0.100$0.2001.05M1
Muse Spark 1.2 Contributoropencode-go$0.100$0.2001.05M1
Muse Spark 1.3 Contributoropencode-go$0.100$0.2001.05M1
Muse Spark 1.2 Contributorllmgateway$0.100$0.2001.05M1
Muse Spark 1.3 Contributorllmgateway$0.100$0.2001.05M1
Seed 2.0 MiniByteDance (Doubao)$0.030$0.280256K5
Seed 2.0 Mini (260428) (ByteDance)ByteDance (Doubao)$0.030$0.280262K3
Nova Lite (APAC)amazon-bedrock$0.063$0.252300K1
GLM Flash LatestZ.AI / Zhipu$0.020$0.3001.31M4
Nova Lite (CA)amazon-bedrock$0.064$0.256300K1
Nex AGI: Nex-N2.5-Prokilo$0.075$0.250262K1
Nex-N2.5-Proopenrouter$0.075$0.250262K1

Top 60 sur 897 affichés. Utilisez le répertoire complet pour filtrer davantage.

Frequently asked questions

How many AI models support entrée d'image?

897 canonical models in our database currently support entrée d'image. The list is regenerated on every data refresh, so it always reflects the latest releases tracked in our catalogue.

What is the cheapest model with entrée d'image?

PaddleOCR-VL from novita-ai is currently the lowest-priced option, at $0.020 per 1M input tokens and $0.020 per 1M output tokens. The full table above is sorted price-ascending.

Which model with entrée d'image has the largest context window?

Llama 4 Scout 17B Instruct (Meta) leads on context at 10M tokens. This may matter if you also need long-document understanding alongside entrée d'image.

Which models are available on the most providers?

Production-readiness usually correlates with how many independent providers host the same weights. The top three by provider count are: Kimi K3 (75), GLM-5.3-Flash (71), DeepSeek V4 Flash (70).

How is entrée d'image different from a regular LLM?

Vision-language models accept image input alongside text. They are multimodal LLMs, not image generators — most reply in text after looking at the image.

How often is this list updated?

Daily. Our data pipeline syncs once a day, regenerates the canonical model list, and rebuilds these pages so newly released models appear within 24 hours.

Dernière mise à jour :

Prices in USD per 1M tokens. Unknown means the provider does not publish per-token pricing.

Pricing and capabilities are refreshed daily and reconciled against each provider's official documentation. Always verify critical production decisions with the provider directly.