Meta
SAM 3.1
About
SAM 3.1 is a 848M open-weight model from Meta. SAM 3.1 Object Multiplex: promptable segmentation and faster multi-object video tracking. It accepts text, images, and video. Released March 26, 2026.
Parameters
848M
Published parameter count.
License
SAM License
License on the weights.
Pricing
Token rates are US dollars per 1M tokens. Image rates are per 1K images. Audio rates are per 1M audio seconds.
| Rate | Price |
|---|---|
| Per 1K images | $1 |
Compare
SAM 3.1 is the only segmentation model in the catalog. The other rows take similar inputs and do a different job.
| Model | What it does | Price | Parameters | Context |
|---|---|---|---|---|
| SAM 3.1 | SAM 3.1 Object Multiplex: promptable segmentation and faster multi-object video tracking. | $1 / 1K images | 848M | — |
| PP-OCRv6 | PaddleOCR PP-OCRv6 medium text detection and recognition; scene OCR JSONL on image chat; text returns plain text; document_url fan-out via text. | $0.02 / 1M input | 20M | — |
| Muse Spark 1.2 | Meta Muse Spark 1.2. Text, multi-image, and native video. Reasoning model. | $1.25 / 1M input | — | 1M |
| Gemini 3.5 Flash Lite | Fastest and cheapest Gemini tier. Multimodal chat; emits no reasoning tokens. | $0.30 / 1M input | — | 1M |
Benchmarks
Published scores for SAM 3.1, from the SAM 3.1 release notes.
- SA-V cgF130.5
- MOSEv2 J&F62.3
| Reading | Example |
|---|---|
| Qualitative | Clear structure, grounded in the input |
| Quantitative | YT-Temporal-1B cgF1: 52.9. |
| Cost and performance | Lower listed rate, mid-pack latency |
Performance
Latency is the end-to-end round trip. Throughput is completion tokens divided by that time. It is a request rate, not decode speed.
Latency
68 ms
Throughput
—
Requests
4,766
Examples

Donut segmentation
Promptable masks via SAM 3.1 with the prompt `donut`.
{"masks":[{"label":"donut","score":0.94,"box_xyxy":[42,210,332,494]}, ...]}
Supported inputs
- textBeside media
A text part is the segmentation target when
method_params.promptis absent. A text-only request is a400. - image_urlYes
One per request, for
segmentandsegment_box. - video_urlYes
One clip, for
track. Sampling is set inmethod_params(video_fps,video_skip_frames,video_max_frames), not at the top level. - document_urlNo
Not accepted.
- Default method
segment.segmenton an image andtrackon a video both takemethod_params.prompt. - Each instance returns a normalized bounding box, a
score, itsarea, and aninstance_id. The pixels are not on the item: every instance shares one label map on the container.
Methods
| Method | Returns |
|---|---|
| segment (default) | One item per instance matching method_params.prompt, plus the container's label map |
| segment_box | The same shape, seeded from method_params.bbox_xywh instead of a text prompt. Items carry no label. |
| track (video) | One item per instance per sampled frame, carrying frame_id and track_id, with a label map on each entry of frames |
Method parameters
| Parameter | Default | Description |
|---|---|---|
prompt | none | What to segment or track, in plain words. Returned verbatim as each item's label. |
bbox_xywh | none | segment_box only: the seed box as [x, y, w, h], normalized 0-1. |
mask_format | png | none drops the label map and keeps area. |
polygons | false | true adds polys_xy outline rings to each item. |
video_fps | none | track only: sample the clip at this rate, in frames per second. Must be greater than 0. |
video_skip_frames | 1 | track only: sample every Nth frame. Cannot be combined with video_fps. |
video_max_frames | 128 | track only: cap on sampled frames, 1 to 128. Above 128 is a 400; lower video_fps to cover a longer clip. |
Unlike the models that decode a video_url natively, facebook/sam3.1 reads its sampling out of method_params, not from a top-level video_fps.
Estimate cost
$1.00
Quick start
Call this model on the OpenAI-compatible gateway. The model id is already filled in.
Response
A single image returns the json block alone, with no wrapper:
{
"object": "img.segment.masks",
"items": [
{
"bbox_xywh": [0.4391, 0.1035, 0.3344, 0.8118],
"label": "person",
"score": 0.9688,
"area": 0.0992,
"instance_id": 1
}
],
"mask": {
"format": "png",
"height": 425,
"width": 640,
"data": "data:image/png;base64,iVBORw0KGgo…"
}
}
bbox_xywh and area are normalized against the source frame. The pixels live in the label map: one 8-bit grayscale PNG where a pixel's value is the item's instance_id on an image, or its track_id on a video.
0 is background, so ids start at 1 and run to 255, which is also the most instances one map can carry. Test pixel == n to cut out instance n, or pixel != 0 for the whole foreground. One decode gives every instance, and area and score sit on the items so a filter never has to decode at all.
Where two instances overlap the pixel goes to the higher score, so each item's area, box and outline describe what is actually visible in the map.
video_fps, video_nframes and video_duration describe the source clip. frames is what was actually sampled, listed in frame_id order and including frames where nothing was found.
A prompt that matches nothing is not an error. The call succeeds with an empty items array.
