Loading...
Loading...
Document analysis — extracting tables, clauses, dates, and entities from unstructured text — is one of the most common enterprise AI use cases. The best models combine strong OCR-equivalent accuracy, long context to process full documents, and reliable structured output for downstream automation.
| # | Model | Input $/1M | |
|---|---|---|---|
| #1 | Gemini 2.0 Flash Lite offers a significantly faster time to first token (TTFT) compared to [Gemini Flash 1.5](/google/gemini-flash-1.5), while maintaining quality on par with larger models like [Gemini Pro 1.5](/google/gemini-pro-1.5),... | $0.07 | Details |
| #2 | Granite 4.0 instruct models deliver strong performance across benchmarks, achieving industry-leading results in key agentic tasks like instruction following and function calling. These efficiencies make the models well-suited for a wide range of use cases like retrieval-augmented generation (RAG), multi-agent workflows, and edge deployments. | $0.02 | Details |
Not sure which model fits your budget and latency requirements?
Use the comparison tool to evaluate any two models side-by-side, or the cost calculator to project your monthly spend.
| #3 | Qwen2.5-Coder-7B-Instruct is a 7B parameter instruction-tuned language model optimized for code-related tasks such as code generation, reasoning, and bug fixing. Based on the Qwen2.5 architecture, it incorporates enhancements like RoPE,... | $0.03 | Details |
| #4 | Qwen-Turbo, based on Qwen2.5, is a 1M context model that provides fast speed and low cost, suitable for simple tasks. | $0.03 | Details |
| #5 | Amazon Nova Micro is the fastest and most cost-effective text-only model in the Nova family, optimized for speed and low latency. Ideal for customer service, summarization, and translation at scale. | $0.04 | Use |
| #6 | OLMo-2 32B Instruct is a supervised instruction-finetuned variant of the OLMo-2 32B March 2025 base model. It excels in complex reasoning and instruction-following tasks across diverse benchmarks such as GSM8K,... | $0.05 | Details |
| #7 | Meta's Llama 3.1 8B served on Groq's LPU for ultra-low latency — ideal for fast, lightweight text tasks. | $0.05 | Use |
| #8 | Open-source Mistral-Small-24B-Instruct-2501 model from mistralai — available for download and self-hosting on Hugging Face. | $0.05 | Details |
| #9 | The Llama 3.2 instruction-tuned text only models are optimized for multilingual dialogue use cases, including agentic retrieval and summarization tasks. | $0.05 | Details |
| #10 | Amazon Nova Lite is a very low-cost multimodal model that can process image, video, and text inputs. Fast and accurate for a wide range of tasks requiring visual and language understanding. | $0.06 | Use |