All models

Google

Gemma 4 26B-A4B Instruct

About

Gemma 4 26B-A4B Instruct is an open-weight model from Google. Google Gemma 4 26B MoE (4B active), instruction-tuned. Text, up to 64 images, or one video per request. It accepts text, images, and video. Context window is 80K. Released May 1, 2026.

  • Context window

    80K

    Tokens of context on a request.

  • License

    —

    License on the weights.

Pricing

Token rates are US dollars per 1M tokens. Image rates are per 1K images. Audio rates are per 1M audio seconds.

RatePrice
Input / 1M$0.10
Cached input / 1M$0.05
Output / 1M$0.30

Compare

Gemma 4 26B-A4B Instruct is one of the catalog's detection models. The table is each model's published price, size, and context.

ModelWhat it doesPriceParametersContext
Gemma 4 26B-A4B InstructGoogle Gemma 4 26B MoE (4B active), instruction-tuned. Text, up to 64 images, or one video per request.$0.10 / 1M input—80K
PP-OCRv6PaddleOCR PP-OCRv6 medium text detection and recognition; scene OCR JSONL on image chat; text returns plain text; document_url fan-out via text.$0.02 / 1M input20M—
DiffusionGemma 26B-A4BBlock-diffusion Gemma 4 26B-A4B: chat over text and up to 8 images, and System One typed decisions on POST /typesafe/v1/systemone.$0.35 / 1M input26B16K
Gemini 3.5 Flash LiteFastest and cheapest Gemini tier. Multimodal chat; emits no reasoning tokens.$0.30 / 1M input—1M

Benchmarks

Published scores for Gemma 4 26B-A4B Instruct, from the Gemma 4 model card.

Evals
ReadingExample
QualitativeClear structure, grounded in the input
QuantitativeMMMU Pro: 73.8.
Cost and performanceLower listed rate, mid-pack latency

Performance

Latency is the end-to-end round trip. Throughput is completion tokens divided by that time. It is a request rate, not decode speed.

  • Latency

    154 ms

  • Throughput

    6,797 tok/s

  • Requests

    89

Methods

Methods this model serves.

MethodReturns
detectionBoxes around objects
chatThe model's reply

Estimate cost

Estimate. An hour of video is 15 frames a minute at 144 tokens a frame, plus the audio in that request, and 500 output tokens a minute.

$33.48

Quick start

Call this model on the OpenAI-compatible gateway. The model id is already filled in.

from openai import OpenAI

client = OpenAI(
    base_url="https://gateway.vlm.run/v1/openai",
    api_key="<VLMRUN_API_KEY>",
)

response = client.chat.completions.create(
    model="google/gemma-4-26b-a4b-it",
    messages=[{"role": "user", "content": "What is in this image?"}],
)

print(response.choices[0].message.content)

Get an API key. Ship vision today.