Technical

On-Device Segmentation and Tracking: What SAM 3.1 on Apple Silicon Unlocks for an Editor

ยท RenderBob team

mlx-cv runs text-prompted grounding, depth, SAM 3.1 segmentation, and video tracking on Apple Silicon through MLX. The project is pre-alpha, and the published video check is two frames.

A local workstation creates subject masks, depth layers, and tracking paths across verified frames with an experimental continuation.

An earlier note in this series covered SAM 3.1 arriving as native ComfyUI segmentation. A parallel Apple Silicon library, mlx-cv from the same appautomaton account as the LTX MLX port, targets the editor's machine. It is an inference-only Python package on MLX: text-prompted grounding with LocateAnything-3B, COCO detection with RF-DETR, depth and camera geometry with Depth Anything V3, panoptic segmentation, and SAM 3.1 for text-prompted image masks plus video propagation. Weights stay outside the package. The README marks the project pre-alpha.

Masking is an interactive loop, so the round trip hurts

A one-shot cloud request can mask an object. Scrubbing a timeline, nudging a selection, and checking the next frame wants the result on the machine you are editing on. Unified memory is the relevant Apple Silicon property here: the frame does not have to be copied to a separate GPU and back for every tweak. mlx-cv does not publish an interactive latency number. What it publishes is that the models run through MLX on Apple Silicon.

A text prompt is a different tool from a fixed object list

RF-DETR in this library detects COCO categories. LocateAnything-3B and SAM 3.1 image mode take a text prompt. The README's own example is "find every traffic sign" and a SAM prompt of "traffic sign". That is open-vocabulary grounding: the category does not have to be one of a fixed detector's labels. For an editor, the useful version is describing the thing in the shot, then getting a mask, instead of hunting a preset list that never included it.

Video tracking is implemented. The published check is two frames

SAM 3.1 video mode in mlx-cv propagates a point, box, or mask with Object Multiplex tracking. The parity table records a real two-frame propagation, with a multiplex mask IoU of 0.99215, and an image mask IoU of 0.999618 on Metal. A selection that holds for a whole shot is the editing feature you want. Two verified frames are evidence the path runs, not a guarantee across a camera move. Depth Anything V3 adds multi-view depth, confidence, and camera intrinsics and extrinsics, which is the geometry you would use to keep a matte or a generated insert consistent. Treat both as early, local building blocks.

Finishing work that still lives in a separate app is another tab in the six-app stack: rotoscope, object isolation, a matte for a composite. On-device segmentation is how that work can sit on the same timeline as the generation. RenderBob's case for Apple Silicon is that local operations like this, and local generation, share the machine with the cut.

More from the blog

All posts