Gemini Robotics-ER 2 (preview)
About
Gemini Robotics-ER 2 (preview) is a closed-weight model from Google. Embodied-reasoning VLM for robotics. Text, multi-image, and native video. It accepts text, images, documents, and video. Context window is 1M. License is Gemini API Terms. Released July 30, 2026.
Context window
1M
Tokens of context on a request.
License
Gemini API Terms
Terms for the API.
Pricing
Token rates are US dollars per 1M tokens. Image rates are per 1K images. Audio rates are per 1M audio seconds.
| Rate | Price |
|---|---|
| Input / 1M | $0.30 |
| Cached input / 1M | $0.03 |
| Output / 1M | $2.50 |
Compare
Gemini Robotics-ER 2 (preview) is one of the catalog's chat models. The table is each model's published price, size, and context.
| Model | What it does | Price | Parameters | Context |
|---|---|---|---|---|
| Gemini Robotics-ER 2 (preview) | Embodied-reasoning VLM for robotics. Text, multi-image, and native video. | $0.30 / 1M input | — | 1M |
| Qwen3.5 0.8B | Multimodal chat model with up to 64 images or one video per request. | $0.08 / 1M input | 0.8B | 256K |
| Gemini 3.5 Flash Lite | Fastest and cheapest Gemini tier. Multimodal chat; emits no reasoning tokens. | $0.30 / 1M input | — | 1M |
| Gemini 3.7 Flash | Newest Gemini Flash tier with extended reasoning. Multimodal chat + native video. | $0.75 / 1M input | — | 1M |
Benchmarks
Published scores for Gemini Robotics-ER 2 (preview), from the Gemini Robotics ER 2 post.
- Moment finding91.3
| Reading | Example |
|---|---|
| Qualitative | Clear structure, grounded in the input |
| Quantitative | Progress classification: 57.4. |
| Cost and performance | Lower listed rate, mid-pack latency |
Methods
Methods this model serves. Payload shapes are in the docs.
| Method | Returns |
|---|---|
| chat | The model's reply |
Estimate cost
Estimate. A page is 2,500 input tokens and 1,000 output tokens. An hour of video is the whole timeline: 258 tokens a second of picture and 32 a second of audio, plus 500 output tokens a minute.
$391.45
Quick start
Call this model on the OpenAI-compatible gateway. The model id is already filled in.
