Inteligencia de modelos de IA

Capacidad · 2026-08-14

Modelos de IA con entrada visual

Modelos que aceptan imágenes junto con texto — comprensión multimodal.

¿Qué es esto?

  • Los modelos visión-lenguaje aceptan imágenes además de (o en lugar de) texto.
  • La mayoría responde con texto — son LLMs multimodales, no generadores de imágenes.

Por qué importa

  • Casos de uso: comprensión de documentos (escaneos, PDF, capturas), revisión de UI/código desde capturas, Q&A de fotos de producto, accesibilidad (alt text), imágenes médicas/satelitales.
  • La facturación suele incluir un coste por imagen además del coste por token — consulta las tablas de offering de cada proveedor.

584 modelos con esta capacidad

ModeloEditorEntrada / 1MSalida / 1MContextoProveedores
PaddleOCR-VLnovita-ai$0.020$0.02016K1
DeepSeek OCR 2DeepSeek$0.030$0.0304K3
Llama 3.2 11B Vision InstructMeta$0.055$0.055128K4
Gemma 3 4B ITGoogle$0.040$0.080128K5
Gemma 4 E4B ITGoogle$0.020$0.100131K3
Nex N2 Mininano-gpt$0.025$0.100262K1
Nex-N2-Miniopenrouter$0.025$0.100262K1
Nex AGI: Nex-N2-Minikilo$0.025$0.100262K1
Model Routerazure-cognitive-services$0.140Unknown200K1
Model Routerazure$0.140Unknown200K1
Qwen3.7 FlashAlibaba (Qwen)$0.030$0.1181M9
Google Gemma 3 12BGoogle$0.050$0.100131K8
Qwen3.5 9BAlibaba (Qwen)$0.040$0.150262K22
nova-micro-v1cortecs$0.040$0.159128K1
Ministral 3 3B 2512Mistral$0.100$0.100131K3
Gemini Embedding 2Google$0.200Unknown8K2
Reka Edgeopenrouter$0.100$0.10016K1
Ministral 3Bllmgateway$0.100$0.100131K1
Reka Edgekilo$0.100$0.10016K1
Greg 1 Minicrof$0.070$0.150229K1
ministral-3b-2512cortecs$0.111$0.111256K1
Google Gemma 3 27B InstructGoogle$0.080$0.160203K11
Gemini 2.0 Flash LiteGoogle$0.052$0.2101.05M3
Ministral 3 8B 2512Mistral$0.150$0.150262K3
Pixtral 12BMistral$0.150$0.150128K2
Nova Lite 1.0openrouter$0.060$0.240300K1
Nova Liteamazon-bedrock$0.060$0.240300K1
Nova Litevercel$0.060$0.240300K1
Ministral 8Bllmgateway$0.150$0.150262K1
Amazon: Nova Lite 1.0kilo$0.060$0.240300K1
Muse Spark 1.2 ContributorMeta$0.100$0.2001.05M3
Qwen3.5 FlashAlibaba (Qwen)$0.029$0.2871M7
ministral-8b-2512cortecs$0.167$0.167256K1
Qwen-Omni TurboAlibaba (Qwen)$0.070$0.27033K3
Mistral Small 3.2 24BMistral$0.094$0.250256K3
nova-lite-v1cortecs$0.069$0.275300K1
Llama Guard 4 12BMeta$0.180$0.1801.05M3
Seed 1.6 Flash (250715)llmgateway$0.070$0.300256K1
Seed 1.6 Flashopenrouter$0.075$0.300262K1
ByteDance Seed: Seed 1.6 Flashkilo$0.075$0.300262K1
Llama 4 ScoutMeta$0.080$0.3001.31M5
Gemma 4 26B A4B ITGoogle$0.060$0.330262K18
Gemma 4 31B ITGoogle$0.102$0.297262K33
Llama 4 Scout 17B 16E InstructMeta$0.100$0.300128K8
Mistral Small 3.1Mistral$0.100$0.300128K3
Ministral 3 14B 2512Mistral$0.200$0.200262K3
Phi-4-multimodalMicrosoft$0.080$0.320128K2
Mistral Small 3.2Mistral$0.100$0.300128K2
Pixtral 12B 2409scaleway$0.200$0.200128K1
Cosmos 3 Super Reasonerllmgateway$0.100$0.300262K1
Ministral 14Bllmgateway$0.200$0.200262K1
Mistral Small 3.2 24B InstructMistral$0.100$0.31033K7
Meta Llama Guard 4 12BMeta$0.210$0.210131K1
MiMo-V2.5xiaomi$0.140$0.2801.05M11
MiMo-V2-Omnixiaomi$0.140$0.280262K2
MiMo V2.5 Thinkingxiaomi$0.140$0.2801.05M1
MiMo V2.5opencode-go$0.140$0.2801M1
MiMo-V2.5cline-pass$0.140$0.2801.05M1
MiMo-V2.5llmgateway$0.140$0.2801M1
MiMo-V2.5pioneer$0.140$0.2801.05M1

Mostrando los 60 primeros de 584. Usa el directorio completo para filtrar más.

Frequently asked questions

How many AI models support entrada de imagen?

584 canonical models in our database currently support entrada de imagen. The list is regenerated on every data refresh, so it always reflects the latest releases tracked in our catalogue.

What is the cheapest model with entrada de imagen?

PaddleOCR-VL from novita-ai is currently the lowest-priced option, at $0.020 per 1M input tokens and $0.020 per 1M output tokens. The full table above is sorted price-ascending.

Which model with entrada de imagen has the largest context window?

Llama 4 Scout 17B Instruct (US) (Meta) leads on context at 3.50M tokens. This may matter if you also need long-document understanding alongside entrada de imagen.

Which models are available on the most providers?

Production-readiness usually correlates with how many independent providers host the same weights. The top three by provider count are: Kimi K2.6 (63), Kimi K2.5 (50), Kimi K3 (50).

How is entrada de imagen different from a regular LLM?

Vision-language models accept image input alongside text. They are multimodal LLMs, not image generators — most reply in text after looking at the image.

How often is this list updated?

Daily. Our data pipeline syncs once a day, regenerates the canonical model list, and rebuilds these pages so newly released models appear within 24 hours.

Última actualización:

Prices in USD per 1M tokens. Unknown means the provider does not publish per-token pricing.

Pricing and capabilities are refreshed daily and reconciled against each provider's official documentation. Always verify critical production decisions with the provider directly.