Gemma 4 26B-A4B Instruct
About
Gemma 4 26B-A4B Instruct is an open-weight model from Google. Google Gemma 4 26B MoE (4B active), instruction-tuned. Text, up to 64 images, or one video per request. It accepts text, images, and video. Context window is 80K. Released May 1, 2026.
Context window
80K
Tokens of context on a request.
License
—
License on the weights.
Pricing
Token rates are US dollars per 1M tokens. Image rates are per 1K images. Audio rates are per 1M audio seconds.
| Rate | Price |
|---|---|
| Input / 1M | $0.10 |
| Cached input / 1M | $0.05 |
| Output / 1M | $0.30 |
Compare
Gemma 4 26B-A4B Instruct is one of the catalog's detection models. The table is each model's published price, size, and context.
| Model | What it does | Price | Parameters | Context |
|---|---|---|---|---|
| Gemma 4 26B-A4B Instruct | Google Gemma 4 26B MoE (4B active), instruction-tuned. Text, up to 64 images, or one video per request. | $0.10 / 1M input | — | 80K |
| PP-OCRv6 | PaddleOCR PP-OCRv6 medium text detection and recognition; scene OCR JSONL on image chat; text returns plain text; document_url fan-out via text. | $0.02 / 1M input | 20M | — |
| DiffusionGemma 26B-A4B | Block-diffusion Gemma 4 26B-A4B: chat over text and up to 8 images, and System One typed decisions on POST /typesafe/v1/systemone. | $0.35 / 1M input | 26B | 16K |
| Gemini 3.5 Flash Lite | Fastest and cheapest Gemini tier. Multimodal chat; emits no reasoning tokens. | $0.30 / 1M input | — | 1M |
Benchmarks
Published scores for Gemma 4 26B-A4B Instruct, from the Gemma 4 model card.
- MMMU Pro73.8
- MATH-Vision82.4
- MedXpertQA MM58.1
| Reading | Example |
|---|---|
| Qualitative | Clear structure, grounded in the input |
| Quantitative | MMMU Pro: 73.8. |
| Cost and performance | Lower listed rate, mid-pack latency |
Performance
Latency is the end-to-end round trip. Throughput is completion tokens divided by that time. It is a request rate, not decode speed.
Latency
154 ms
Throughput
6,797 tok/s
Requests
89
Methods
Methods this model serves.
| Method | Returns |
|---|---|
| detection | Boxes around objects |
| chat | The model's reply |
Estimate cost
Estimate. An hour of video is 15 frames a minute at 144 tokens a frame, plus the audio in that request, and 500 output tokens a minute.
$33.48
Quick start
Call this model on the OpenAI-compatible gateway. The model id is already filled in.
