All models

Google

DiffusionGemma 26B-A4B

About

DiffusionGemma 26B-A4B is a 26B open-weight model from Google. Block-diffusion Gemma 4 26B-A4B: chat over text and up to 8 images, and System One typed decisions on POST /typesafe/v1/systemone. It accepts text and images. Context window is 16K. Released June 10, 2026.

Served on two routes. On chat completions it behaves like the Gateway's other chat models. On POST /typesafe/v1/systemone it powers System One, answering typed questions with a probability per label, read in one denoise step rather than generated, so an answer can never be off-schema.

  • Parameters

    26B

    Published parameter count.

  • Context window

    16K

    Tokens of context on a request.

  • License

    Apache-2.0

    License on the weights.

Pricing

Token rates are US dollars per 1M tokens. Image rates are per 1K images. Audio rates are per 1M audio seconds.

RatePrice
Input / 1M$0.35
Cached input / 1M$0.05
Output / 1M$1.50

Compare

DiffusionGemma 26B-A4B is one of the catalog's detection models. The table is each model's published price, size, and context.

ModelWhat it doesPriceParametersContext
DiffusionGemma 26B-A4BBlock-diffusion Gemma 4 26B-A4B: chat over text and up to 8 images, and System One typed decisions on POST /typesafe/v1/systemone.$0.35 / 1M input26B16K
PP-OCRv6PaddleOCR PP-OCRv6 medium text detection and recognition; scene OCR JSONL on image chat; text returns plain text; document_url fan-out via text.$0.02 / 1M input20M—
Gemma 4 26B-A4B InstructGoogle Gemma 4 26B MoE (4B active), instruction-tuned. Text, up to 64 images, or one video per request.$0.10 / 1M input—80K
Florence-2Vision foundation model for captioning, OCR, detection, and region tasks.$0.10 / 1M input232M—

Benchmarks

Published scores for DiffusionGemma 26B-A4B, from the DiffusionGemma NVFP4 model card. NVFP4 checkpoint, thinking enabled.

Evals
ReadingExample
QualitativeClear structure, grounded in the input
QuantitativeGPQA Diamond: 68.6.
Cost and performanceLower listed rate, mid-pack latency

Performance

Latency is the end-to-end round trip. Throughput is completion tokens divided by that time. It is a request rate, not decode speed.

  • Latency

    240 ms

  • Throughput

    1,453 tok/s

  • Requests

    208

Methods

Methods this model serves.

MethodReturns
detectionBoxes around objects
chatThe model's reply

Capabilities

RoutesPOST /v1/openai/chat/completions, POST /typesafe/v1/systemone, and WS /typesafe/ws
Methodchat
Accepted inputstext, image_url. On the decisions route also a file or document_url part (one PDF), carried in content.
Max images8 per request, 5 MB each, JPEG / PNG / WebP / GIF
VideoNot supported
StreamingToken streaming on chat completions. A read is a single response; WS /typesafe/ws answers per frame.

On the decisions route, detail sets the vision budget per image or page:

detailVision tokens
auto (default)280, the same as high
high280
low70

Warning

This model rejects temperature and seed with a 400 rather than ignoring them, which differs from the other chat models on the Gateway. Omit both.

Typed decisions

On POST /typesafe/v1/systemone the same model answers typed questions with a probability per label, read in one denoise step rather than generated, so an answer can never be off-schema. It is the default engine there, and the most sharply calibrated.

Question typesnoul (yes/no), choice (one of 2 to 128 labels), score (a 2 to 10 level rubric)
Mediaimage_url parts, or one file / document_url PDF, on content
reasoning_effortNot supported on a diffusion engine; a request carrying it is a 422
Output tokensAlways 0. Nothing is generated

See System One for the request shape and a worked example in four languages, and Models for how this engine compares with the generative ones.

Chat completions

The same model on the OpenAI-compatible route, for when you want generated text rather than a typed decision.

from openai import OpenAI

client = OpenAI(
    base_url="https://gateway.vlm.run/v1/openai",
    api_key="<VLMRUN_API_KEY>",
)

response = client.chat.completions.create(
    model="google/diffusiongemma-26b-a4b-it",
    messages=[{"role": "user", "content": [
        {"type": "text", "text": "What is this page and what is the amount due?"},
        {"type": "image_url", "image_url": {"url": f"data:image/jpeg;base64,{page}"}},
    ]}],
    extra_body={"method": "chat"},
)

print(response.choices[0].message.content)

Sampling fields are not interchangeable with the other chat models: temperature and seed return a 400 here instead of being ignored.

Estimate cost

$1.85

Quick start

Call this model on the OpenAI-compatible gateway. The model id is already filled in.

from openai import OpenAI

client = OpenAI(
    base_url="https://gateway.vlm.run/v1/openai",
    api_key="<VLMRUN_API_KEY>",
)

response = client.chat.completions.create(
    model="google/diffusiongemma-26b-a4b-it",
    messages=[{"role": "user", "content": "What is in this image?"}],
)

print(response.choices[0].message.content)

Get an API key. Ship vision today.