Multimodal, Multitask
One catalog spanning multi-modal inputs and multi-task outputs: OCR, detection, segmentation, pose, keypoints, and more.
Document OCR, captioning, and multi-modal chat: every visual capability behind one MCP server.
| Model | Capabilities | Inputs | Context | Input / 1M | Cached Input / 1M | Output / 1M |
|---|---|---|---|---|---|---|
Muse Glimmer 30B | chat | 128K | $0.35 | — | $1.50 | |
Qwen3.8 27B | chat | 256K | $0.35 | $0.085 | $2.55 | |
Unlimited-OCR | ocrmarkdown | 32K | $0.25 | $0.03 | $0.55 | |
Kimi K3 | chat | 256K | $3 | — | $15 | |
PP-OCRv6 | ocrdetection | — | $0.01 | $0.01 | $0.20 | |
MiniMax M3 | chat | 1M | $0.30 | — | $1.20 | |
PaddleOCR-VL 1.6 | ocr | 16K | $0.15 | $0.15 | $0.35 | |
dots.mocr | ocrmarkdown | 32K | $0.20 | $0.03 | $0.40 | |
Gemma 4 26B-A4B Instruct | chat | 128K | $0.10 | — | $0.30 | |
Gemma 4 31B Instruct | chat | 128K | $0.15 | — | $0.40 | |
Qwen3.5 0.8B | detectionchat | 256K | $0.08 | $0.02 | $0.15 | |
Infinity-Parser2-Flash | markdown | 32K | $0.25 | $0.03 | $0.45 |
Showing 12 of 25 models
Pricing calculator
Pick a document type, page volume, and OCR model. See cost savings versus closed vision APIs.
Page volume / month
Gateway OCR tokens (in / out)
120M - 250M / 100M
Frontier billed tokens (in / out)
250M / 200M
2.5K image tokens per page; output includes reasoning at 2× OCR text.
Total costwhat the gateway bills for this workload
$23 - $90
Total cost savingsvs the cheapest frontier model
$848 - $915
Documents / month
100Kpages
Estimates vs typical OCR, Document AI, and frontier VLM pricing.Source: llm-prices.com
Footage volume / month
Gateway billed tokens (in / out)
2.4B - 3.5B / 300M
Gemini 3.5 Flash Lite billed tokens (in / out)
10.4B / 600M
Total costwhat the gateway bills for this workload
$335 - $2K
Total cost savingsvs the cheapest frontier model
$2.7K - $4.3K
Video / month
10Khours
LLM routers and gateways route to 100s of LLMs, yet only a handful of VLMs, and often no OCR or classical CV models. Visual AI deserves its own stack.
One catalog spanning multi-modal inputs and multi-task outputs: OCR, detection, segmentation, pose, keypoints, and more.
OCR, captioning, and multimodal chat run through the same OpenAI-compatible chat completions API you already use.
Send a 500-page PDF or a 2-hour video in a single call. The gateway chunks, batches, and reassembles for you. No pipelines to build.
JSON-schema enforcement on every call. CV wrappers emit fixed schemas; VLMs honor response_format.
Every model is open-weight, served through the OpenAI-compatible API you already use. Swap models or providers freely, your client code never changes.
Give any MCP client instant access to the full visual model catalog. Agents see, read, and reason over images out of the box.
Optimized for production workloads and cost-efficiency. Every model deployed gets its own performance tune-up.
Models Supported
21
p50 Latency
<100ms
Uptime SLA
99.9%
Type II · HIPAA · BAA
SOC 2
Point the OpenAI SDK at the gateway and swap the model, or hand the same catalog to an agent over MCP. Same signature, exhaustive visual model catalog.
Use cases
Compose vision models like building blocks. One SDK. One key. One bill.
Layout, OCR, and markdown extraction across contracts, statements, and forms, at sub-cent per-page economics.
Build powerful multi-modal chatbots and agents with VLMs that pack native support for multi-modal inputs such as images, PDFs, or videos in a single chat completion.
Auto-caption, tag, and index large image and video libraries for search, dedup, and recommendations.
Stand up the whole visual stack inside your own VPC or on-prem, with SSO, RBAC, and audit log export. Regulated data stays inside your boundary, under a BAA, on dedicated capacity you control.