AI MODEL ATLAS · DECISION PAGE
Long-context LLMs for a 128K + 4K workload
Models whose listed context can fit 128,000 input plus 4,000 output tokens.
OBJECT ID
asca:openrouter-model-page-long-context-llm@2026-08-27.v1FETCHED
STATUSREPRODUCIBLE SNAPSHOT
QUESTION · METHOD
Long-context LLMs for a 128K + 4K workload
Keep context >= 132,000 and positive prompt/completion prices; derived_cost_usd = 128,000 × prompt + 4,000 × completion; request price is reported separately and excluded.
OBSERVED MODELS · 10 SHOWN
Comparable fields from one dated catalog snapshot
| Rank / order | Model | Input / 1M | Output / 1M | Request | Context | Tools | Structured | Derived cost / basis |
|---|---|---|---|---|---|---|---|---|
| 73 | Ling-3.0-flash | $0.021 | $0.063 | $0.00 | 262.14K | Yes | No | $0.0029 · $ per 128K input + 4K output text tokens |
| 140 | Nex AGI: Nex-N2-Mini | $0.025 | $0.10 | $0.00 | 262.14K | Yes | Yes | $0.0036 · $ per 128K input + 4K output text tokens |
| 202 | OpenAI: GPT-5 Nano (batch) | $0.025 | $0.20 | $0.00 | 400K | Yes | Yes | $0.004 · $ per 128K input + 4K output text tokens |
| 388 | DeepSeek V4 Flash Latest | $0.03 | $0.075 | $0.00 | 1.31M | Yes | Yes | $0.0041 · $ per 128K input + 4K output text tokens |
| 23 | Upstage: Solar Pro 4 | $0.03 | $0.12 | $0.00 | 524.29K | Yes | Yes | $0.0043 · $ per 128K input + 4K output text tokens |
| 50 | Qwen: Qwen3.7 Flash | $0.03 | $0.13 | $0.00 | 1M | Yes | No | $0.0044 · $ per 128K input + 4K output text tokens |
| 96 | Qwen: Qwen3 30B A3B Instruct 2507 | $0.0482 | $0.1931 | $0.00 | 262.14K | Yes | Yes | $0.0069 · $ per 128K input + 4K output text tokens |
| 219 | Google: Gemini 2.5 Flash Lite (batch) | $0.05 | $0.20 | $0.00 | 1.05M | Yes | Yes | $0.0072 · $ per 128K input + 4K output text tokens |
| 177 | NVIDIA: Nemotron 3 Nano 30B A3B | $0.05 | $0.20 | $0.00 | 262.14K | Yes | Yes | $0.0072 · $ per 128K input + 4K output text tokens |
| 373 | OpenAI: GPT-4.1 Nano (batch) | $0.05 | $0.20 | $0.00 | 1.05M | Yes | Yes | $0.0072 · $ per 128K input + 4K output text tokens |
Prices are listed USD per token converted to USD per 1M tokens. Request fee is a separate listed field; it is not included in derived text-token cost. A free listing means zero listed token price in this snapshot, not unlimited access.
RANGE · BIAS · LIMIT · CONFIDENCE
- Observed range: 275 matching catalog records; 10 shown.
- Bias: Context metadata and price selection bias.
- Limitation: A listed context limit does not prove usable quality, latency, or successful long-document completion.
- Confidence: high for the recorded response and stated transformation.
DOWNLOAD · VERIFY · CITE
Keep the snapshot and object ID with the decision.
“Keep context >= 132,000 and positive prompt/completion prices; derived_cost_usd = 128,000 × prompt + 4,000 × completion; request price is reported separately and excluded.”— asca:openrouter-model-page-long-context-llm@2026-08-27.v1
RELATED DECISION PAGES
- Cheapest paid LLM API models
- Long-context LLMs for a 128K + 4K workload
- Tool-calling compatible models
- Structured-output compatible models
- Vision-language models
- Free LLM API models
- Most popular AI models by OpenRouter weekly order
- Fastest LLM API models by OpenRouter order
- Coding AI models by OpenRouter order
- Agentic AI models by OpenRouter order