Meta
Muse Spark 1.2
About
Muse Spark 1.2 is a closed-weight model from Meta. Meta Muse Spark 1.2. Text, multi-image, and native video. Reasoning model. It accepts text, images, and video. Context window is 1M. Released August 5, 2026.
Context window
1M
Tokens of context on a request.
License
Meta API Terms
Terms for the API.
Pricing
Token rates are US dollars per 1M tokens. Image rates are per 1K images. Audio rates are per 1M audio seconds.
| Rate | Price |
|---|---|
| Input / 1M | $1.25 |
| Cached input / 1M | $0.15 |
| Output / 1M | $4.25 |
Compare
Muse Spark 1.2 is one of the catalog's detection models. The table is each model's published price, size, and context.
| Model | What it does | Price | Parameters | Context |
|---|---|---|---|---|
| Muse Spark 1.2 | Meta Muse Spark 1.2. Text, multi-image, and native video. Reasoning model. | $1.25 / 1M input | — | 1M |
| PP-OCRv6 | PaddleOCR PP-OCRv6 medium text detection and recognition; scene OCR JSONL on image chat; text returns plain text; document_url fan-out via text. | $0.02 / 1M input | 20M | — |
| Muse Glimmer 30B | Meta Muse Glimmer 30B. Text, multi-image, and native video. | $0.35 / 1M input | 30B | 128K |
| Gemini 3.5 Flash Lite | Fastest and cheapest Gemini tier. Multimodal chat; emits no reasoning tokens. | $0.30 / 1M input | — | 1M |
Benchmarks
Published scores for Muse Spark 1.2, from the Gemini 3.7 Flash model card. Scores Google published for Muse Spark 1.2.
- DeepSWE v1.154.9
- GDP.pdf16
| Reading | Example |
|---|---|
| Qualitative | Clear structure, grounded in the input |
| Quantitative | Terminal-bench 2.1: 82.9. |
| Cost and performance | Lower listed rate, mid-pack latency |
Performance
Latency is the end-to-end round trip. Throughput is completion tokens divided by that time. It is a request rate, not decode speed.
Latency
7,296 ms
Throughput
242 tok/s
Requests
77
Methods
Methods this model serves.
| Method | Returns |
|---|---|
| detection | Boxes around objects |
| chat | The model's reply |
Estimate cost
Estimate. An hour of video is 15 frames a minute at 256 tokens a frame, plus the audio in that request, and 500 output tokens a minute.
$559.50
Quick start
Call this model on the OpenAI-compatible gateway. The model id is already filled in.
