DiffusionGemma 26B-A4B
About
DiffusionGemma 26B-A4B is a 26B open-weight model from Google. Block-diffusion Gemma 4 26B-A4B: chat over text and up to 8 images, and System One typed decisions on POST /typesafe/v1/systemone. It accepts text and images. Context window is 16K. Released June 10, 2026.
Served on two routes. On chat completions it behaves like the Gateway's other chat models. On POST /typesafe/v1/systemone it powers System One, answering typed questions with a probability per label, read in one denoise step rather than generated, so an answer can never be off-schema.
Parameters
26B
Published parameter count.
Context window
16K
Tokens of context on a request.
License
Apache-2.0
License on the weights.
Pricing
Token rates are US dollars per 1M tokens. Image rates are per 1K images. Audio rates are per 1M audio seconds.
| Rate | Price |
|---|---|
| Input / 1M | $0.35 |
| Cached input / 1M | $0.05 |
| Output / 1M | $1.50 |
Compare
DiffusionGemma 26B-A4B is one of the catalog's detection models. The table is each model's published price, size, and context.
| Model | What it does | Price | Parameters | Context |
|---|---|---|---|---|
| DiffusionGemma 26B-A4B | Block-diffusion Gemma 4 26B-A4B: chat over text and up to 8 images, and System One typed decisions on POST /typesafe/v1/systemone. | $0.35 / 1M input | 26B | 16K |
| PP-OCRv6 | PaddleOCR PP-OCRv6 medium text detection and recognition; scene OCR JSONL on image chat; text returns plain text; document_url fan-out via text. | $0.02 / 1M input | 20M | — |
| Gemma 4 26B-A4B Instruct | Google Gemma 4 26B MoE (4B active), instruction-tuned. Text, up to 64 images, or one video per request. | $0.10 / 1M input | — | 80K |
| Florence-2 | Vision foundation model for captioning, OCR, detection, and region tasks. | $0.10 / 1M input | 232M | — |
Benchmarks
Published scores for DiffusionGemma 26B-A4B, from the DiffusionGemma NVFP4 model card. NVFP4 checkpoint, thinking enabled.
- GPQA Diamond68.6
- AIME 202567.33
- MMLU Pro80.7
| Reading | Example |
|---|---|
| Qualitative | Clear structure, grounded in the input |
| Quantitative | GPQA Diamond: 68.6. |
| Cost and performance | Lower listed rate, mid-pack latency |
Performance
Latency is the end-to-end round trip. Throughput is completion tokens divided by that time. It is a request rate, not decode speed.
Latency
240 ms
Throughput
1,453 tok/s
Requests
208
Methods
Methods this model serves.
| Method | Returns |
|---|---|
| detection | Boxes around objects |
| chat | The model's reply |
Capabilities
| Routes | POST /v1/openai/chat/completions, POST /typesafe/v1/systemone, and WS /typesafe/ws |
| Method | chat |
| Accepted inputs | text, image_url. On the decisions route also a file or document_url part (one PDF), carried in content. |
| Max images | 8 per request, 5 MB each, JPEG / PNG / WebP / GIF |
| Video | Not supported |
| Streaming | Token streaming on chat completions. A read is a single response; WS /typesafe/ws answers per frame. |
On the decisions route, detail sets the vision budget per image or page:
detail | Vision tokens |
|---|---|
auto (default) | 280, the same as high |
high | 280 |
low | 70 |
Warning
This model rejects temperature and seed with a 400 rather than ignoring them, which differs from the other chat models on the Gateway. Omit both.
Typed decisions
On POST /typesafe/v1/systemone the same model answers typed questions with a probability per label, read in one denoise step rather than generated, so an answer can never be off-schema. It is the default engine there, and the most sharply calibrated.
| Question types | noul (yes/no), choice (one of 2 to 128 labels), score (a 2 to 10 level rubric) |
| Media | image_url parts, or one file / document_url PDF, on content |
reasoning_effort | Not supported on a diffusion engine; a request carrying it is a 422 |
| Output tokens | Always 0. Nothing is generated |
See System One for the request shape and a worked example in four languages, and Models for how this engine compares with the generative ones.
Chat completions
The same model on the OpenAI-compatible route, for when you want generated text rather than a typed decision.
Sampling fields are not interchangeable with the other chat models: temperature and seed return a 400 here instead of being ignored.
Estimate cost
$1.85
Quick start
Call this model on the OpenAI-compatible gateway. The model id is already filled in.
