Loading...
Loading...
Vision-language models can interpret images, extract text via OCR, understand charts and diagrams, and answer questions about visual content. Frontier models like GPT-4o, Claude, and Gemini now match or exceed specialised vision systems on most real-world tasks.
| # | Model | Input $/1M | |
|---|---|---|---|
| #1 | Qwen2.5-Coder-7B-Instruct is a 7B parameter instruction-tuned language model optimized for code-related tasks such as code generation, reasoning, and bug fixing. Based on the Qwen2.5 architecture, it incorporates enhancements like RoPE,... | $0.03 | Details |
| #2 | Amazon Nova Micro is the fastest and most cost-effective text-only model in the Nova family, optimized for speed and low latency. Ideal for customer service, summarization, and translation at scale. | $0.04 | Use |
Not sure which model fits your budget and latency requirements?
Use the comparison tool to evaluate any two models side-by-side, or the cost calculator to project your monthly spend.
| #3 |
Amazon Nova Lite is a very low-cost multimodal model that can process image, video, and text inputs. Fast and accurate for a wide range of tasks requiring visual and language understanding. |
| $0.06 |
| Use |
| #4 | Open-source Qwen3-VL-8B-Instruct model from qwen — available for download and self-hosting on Hugging Face. | $0.08 | Details |
| #5 | Open-source Qwen3-VL-32B-Instruct model from qwen — available for download and self-hosting on Hugging Face. | $0.10 | Details |
| #6 | Open-source Qwen3-VL-30B-A3B-Instruct model from qwen — available for download and self-hosting on Hugging Face. | $0.13 | Details |
| #7 | Qwen's Enhanced Large Visual Language Model. Significantly upgraded for detailed recognition capabilities and text recognition abilities, supporting ultra-high pixel resolutions up to millions of pixels and extreme aspect ratios for... | $0.14 | Details |
| #8 | A powerful multimodal Mixture-of-Experts chat model featuring 28B total parameters with 3B activated per token, delivering exceptional text and vision understanding through its innovative heterogeneous MoE structure with modality-isolated routing.... | $0.14 | Details |
| #9 | Amazon Titan Text Express is a generative LLM for summarization, text generation, classification, open-ended Q&A, and information extraction. Optimized for enterprise workloads via AWS Bedrock. | $0.20 | Use |
| #10 | AI21 Jamba 1.6 Mini is a lightweight Mamba-Transformer hybrid optimized for cost-effective, high-throughput inference with an impressive 256K context window. An excellent choice for document-heavy workloads on a budget. | $0.20 | Use |