AI MODEL ATLAS · DECISION PAGE
Vision-language models
Models whose architecture declares image input plus text output.
OBJECT ID
asca:openrouter-model-page-vision-language-models@2026-08-27.v1FETCHED
STATUSREPRODUCIBLE SNAPSHOT
QUESTION · METHOD
Vision-language models
Filter input_modalities for image and output_modalities for text; derived_cost_usd = prompt + completion USD per 1M text tokens; request price is reported separately and excluded.
OBSERVED MODELS · 10 SHOWN
Comparable fields from one dated catalog snapshot
| Rank / order | Model | Input / 1M | Output / 1M | Request | Context | Tools | Structured | Derived cost / basis |
|---|---|---|---|---|---|---|---|---|
| 140 | Nex AGI: Nex-N2-Mini | $0.025 | $0.10 | $0.00 | 262.14K | Yes | Yes | $0.125 · $ per 1M input + 1M output text tokens |
| 204 | Google: Gemma 3 4B | $0.05 | $0.10 | $0.00 | 131.07K | No | Yes | $0.15 · $ per 1M input + 1M output text tokens |
| 50 | Qwen: Qwen3.7 Flash | $0.03 | $0.13 | $0.00 | 1M | Yes | No | $0.16 · $ per 1M input + 1M output text tokens |
| 131 | Google: Gemma 3 12B | $0.05 | $0.15 | $0.00 | 131.07K | Yes | Yes | $0.20 · $ per 1M input + 1M output text tokens |
| 183 | Mistral: Ministral 3 3B 2512 | $0.10 | $0.10 | $0.00 | 131.07K | Yes | Yes | $0.20 · $ per 1M input + 1M output text tokens |
| 341 | Reka Edge | $0.10 | $0.10 | $0.00 | 16.38K | Yes | Yes | $0.20 · $ per 1M input + 1M output text tokens |
| 202 | OpenAI: GPT-5 Nano (batch) | $0.025 | $0.20 | $0.00 | 400K | Yes | Yes | $0.225 · $ per 1M input + 1M output text tokens |
| 219 | Google: Gemini 2.5 Flash Lite (batch) | $0.05 | $0.20 | $0.00 | 1.05M | Yes | Yes | $0.25 · $ per 1M input + 1M output text tokens |
| 373 | OpenAI: GPT-4.1 Nano (batch) | $0.05 | $0.20 | $0.00 | 1.05M | Yes | Yes | $0.25 · $ per 1M input + 1M output text tokens |
| 100 | Qwen: Qwen3.5-9B | $0.10 | $0.15 | $0.00 | 262.14K | Yes | Yes | $0.25 · $ per 1M input + 1M output text tokens |
Prices are listed USD per token converted to USD per 1M tokens. Request fee is a separate listed field; it is not included in derived text-token cost. A free listing means zero listed token price in this snapshot, not unlimited access.
RANGE · BIAS · LIMIT · CONFIDENCE
- Observed range: 238 matching catalog records; 10 shown.
- Bias: Architecture declaration and catalog bias.
- Limitation: Modalities do not specify image resolution, video support, visual quality, or successful task performance.
- Confidence: high for the recorded response and stated transformation.
DOWNLOAD · VERIFY · CITE
Keep the snapshot and object ID with the decision.
“Filter input_modalities for image and output_modalities for text; derived_cost_usd = prompt + completion USD per 1M text tokens; request price is reported separately and excluded.”— asca:openrouter-model-page-vision-language-models@2026-08-27.v1
RELATED DECISION PAGES
- Cheapest paid LLM API models
- Long-context LLMs for a 128K + 4K workload
- Tool-calling compatible models
- Structured-output compatible models
- Vision-language models
- Free LLM API models
- Most popular AI models by OpenRouter weekly order
- Fastest LLM API models by OpenRouter order
- Coding AI models by OpenRouter order
- Agentic AI models by OpenRouter order