mlx-community/SAM2.1-hiera-small-fp16

SAM 2.1 Hiera-Small, image mode (point / box → mask), converted to Apple MLX (fp16) from Meta's official sam2.1_hiera_small.pt for the mlx-edgetam-swift Swift package (SAM21Package, an MLXEngine promptSegment ModelPackage).

  • Contents: the image path only — Hiera trunk + FPN neck + SAM prompt encoder + mask decoder + no_mem_embed (359 tensors, 38.5 M params). The video memory stack is not included.
  • Keys: upstream's names, unchanged. Convolutions are NHWC (O,kH,kW,I). The trunk's pos_embed and pos_embed_window are (1,H,W,C). Converter: oracle/convert_sam21.py.
  • Parity: the Swift port matches PyTorch SAM2ImagePredictor through set_image on the CPU fp32 stream. On a 1024² flat-shaded image, a 1024² captioned copy and a 1200×800 synthetic page:
    • image_embed relative error ≤ 5.3e-6;
    • every click and box mask matches at IoU 1.0000, with predicted-IoU Δ ≤ 1e-5. With these fp16 weights and fp16 activations on the GPU, mask IoU vs fp32 PyTorch is ≥ 0.9996 on object-sized prompts.
  • Footprint (M5 Max): 1024² encode ~0.1 s warm; each further prompt on the same image 11–18 ms; process phys_footprint 2.4–2.8 GB.

Why use it next to EdgeTAM

EdgeTAM (13.9 M params) is the fast default and handles video. On single clicks over flat-shaded or synthetic images, SAM 2.1-S is much better:

  • a click on a cartoon fox's belly selects the whole fox (IoU 0.968 vs 0.008 for EdgeTAM);
  • a click on one small square of a grey page selects only that square, where EdgeTAM also adds a corner square every time.

Use

// .package(url: "https://github.com/xocialize/mlx-edgetam-swift", from: "0.6.0")
import EdgeTAM
let p = try EdgeTAMPredictor.fromPretrained(weightsPath, dtype: .float16)   // Hiera detected from the weights
p.setImage(sourceCGImage)
let r = p.predict(point: (471, 635))            // r.mask, r.soft (anti-aliased), r.score
let boxed = p.predict(points: [], labels: [], box: [40, 245, 745, 910])

Or through MLXEngine: SAM21Package (MLXEdgeTAM), with mode: softMatte for an anti-aliased matte.

Weights: Apache-2.0 (facebookresearch/sam2). Port code: MIT.

Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
38.5M params
Tensor type
F16
·
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mlx-community/SAM2.1-hiera-small-fp16

Finetuned
(20)
this model