Instructions to use mlx-community/SAM2.1-hiera-small-fp16 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use mlx-community/SAM2.1-hiera-small-fp16 with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir SAM2.1-hiera-small-fp16 mlx-community/SAM2.1-hiera-small-fp16
- sam2
How to use mlx-community/SAM2.1-hiera-small-fp16 with sam2:
# Use SAM2 with images import torch from sam2.sam2_image_predictor import SAM2ImagePredictor predictor = SAM2ImagePredictor.from_pretrained(mlx-community/SAM2.1-hiera-small-fp16) with torch.inference_mode(), torch.autocast("cuda", dtype=torch.bfloat16): predictor.set_image(<your_image>) masks, _, _ = predictor.predict(<input_prompts>)# Use SAM2 with videos import torch from sam2.sam2_video_predictor import SAM2VideoPredictor predictor = SAM2VideoPredictor.from_pretrained(mlx-community/SAM2.1-hiera-small-fp16) with torch.inference_mode(), torch.autocast("cuda", dtype=torch.bfloat16): state = predictor.init_state(<your_video>) # add new prompts and instantly get the output on the same frame frame_idx, object_ids, masks = predictor.add_new_points(state, <your_prompts>): # propagate the prompts to get masklets throughout the video for frame_idx, object_ids, masks in predictor.propagate_in_video(state): ... - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
mlx-community/SAM2.1-hiera-small-fp16
SAM 2.1 Hiera-Small, image mode (point / box → mask),
converted to Apple MLX (fp16) from Meta's official sam2.1_hiera_small.pt for the
mlx-edgetam-swift Swift package (SAM21Package, an MLXEngine
promptSegment ModelPackage).
- Contents: the image path only — Hiera trunk + FPN neck + SAM prompt encoder + mask decoder +
no_mem_embed(359 tensors, 38.5 M params). The video memory stack is not included. - Keys: upstream's names, unchanged. Convolutions are NHWC
(O,kH,kW,I). The trunk'spos_embedandpos_embed_windoware(1,H,W,C). Converter:oracle/convert_sam21.py. - Parity: the Swift port matches PyTorch
SAM2ImagePredictorthroughset_imageon the CPU fp32 stream. On a 1024² flat-shaded image, a 1024² captioned copy and a 1200×800 synthetic page:- image_embed relative error ≤ 5.3e-6;
- every click and box mask matches at IoU 1.0000, with predicted-IoU Δ ≤ 1e-5. With these fp16 weights and fp16 activations on the GPU, mask IoU vs fp32 PyTorch is ≥ 0.9996 on object-sized prompts.
- Footprint (M5 Max): 1024² encode ~0.1 s warm; each further prompt on the same image 11–18 ms; process
phys_footprint2.4–2.8 GB.
Why use it next to EdgeTAM
EdgeTAM (13.9 M params) is the fast default and handles video. On single clicks over flat-shaded or synthetic images, SAM 2.1-S is much better:
- a click on a cartoon fox's belly selects the whole fox (IoU 0.968 vs 0.008 for EdgeTAM);
- a click on one small square of a grey page selects only that square, where EdgeTAM also adds a corner square every time.
Use
// .package(url: "https://github.com/xocialize/mlx-edgetam-swift", from: "0.6.0")
import EdgeTAM
let p = try EdgeTAMPredictor.fromPretrained(weightsPath, dtype: .float16) // Hiera detected from the weights
p.setImage(sourceCGImage)
let r = p.predict(point: (471, 635)) // r.mask, r.soft (anti-aliased), r.score
let boxed = p.predict(points: [], labels: [], box: [40, 245, 745, 910])
Or through MLXEngine: SAM21Package (MLXEdgeTAM), with mode: softMatte for an anti-aliased matte.
Weights: Apache-2.0 (facebookresearch/sam2). Port code: MIT.
Model size
38.5M params
Tensor type
F16
·
Hardware compatibility
Log In to add your hardware
Quantized
Model tree for mlx-community/SAM2.1-hiera-small-fp16
Base model
facebook/sam2.1-hiera-small