News
LTX-2.3 on MLX: A Real Text-to-Video Model Built to Run on Apple Silicon
ยท RenderBob team
A community MLX port runs LTX-2.3, a 22-billion-parameter text-to-video and image-to-video model with synchronised audio, at 8-bit or 4-bit on Apple Silicon, and can train a LoRA on the Mac.

For most of this year, local video on a Mac meant a small model or a full-size one dragging through an awkward port. The ltx-video-mlx project, from the appautomaton GitHub account, runs Lightricks' LTX-2.3 on Apple's MLX framework: a 22-billion-parameter text-to-video and image-to-video model that also writes synchronised audio. The README describes clips of about 5 to 10 seconds, generated on the machine, with no PyTorch in the loop.
Why this port is MLX
MLX is Apple's open-source array framework for Apple Silicon. Arrays live in unified memory, operations can run in parallel while MLX tracks dependencies, and computation is lazy: the graph is built first and evaluated when a result is needed. Apple walked through that design at WWDC25. A model written for MLX inherits the shared memory pool instead of copying tensors between a CPU and a discrete GPU.
8-bit and 4-bit are what make 22 billion parameters fit
The project quantises inference to 8-bit or 4-bit through MLX quantized matmul. Its requirements list about 14 GB of unified memory at 8-bit and about 10 GB at 4-bit, on an M1 or later. The download is larger than the runtime figure: an FP8 checkpoint of about 29 GB, a distilled LoRA, a spatial upscaler, and a Gemma 3 12B text encoder of about 24 GB on disk. Four-bit is the smaller, faster path. Eight-bit spends memory to keep more of the model's fidelity. That is the same trade this series has already described for ComfyUI quantisation, pointed at a Mac.
The LoRA trains on the same machine
The repository also ships on-device LoRA fine-tuning for a video style or subject. You precompute embeddings, train an adapter, and pass it back into generate.py. The example command uses 1,000 steps. Reference clips for that run stay on the Mac. The code is marked for research use, and the weights remain under Lightricks' LTX-Video licence, so a client job still starts by reading that licence.
A timeline that lets you pick a local model or a cloud model per clip needs a local option with real picture, audio, and a training path. This port is one. RenderBob is built for that choice: a generated clip lands on its own lane, local when the Mac can carry the preview, cloud when the shot needs a heavier model.
More from the blog
- Visual Dubbing Goes Mainstream: Prime Video Changes the Mouth, Not the Voice
On 9 September 2026, Prime Video launched AI lip-sync for the English dub of Maxton Hall. Human actors record the dialogue. The actors' mouths are regenerated to match.
- Suggestive Editing Arrives: Story-Aware Rough Cuts Move Into the Mainstream
From Eddie AI's story-structured assemblies at NAB 2026 to Premiere's AI Assistant and Resolve 21 search, editing tools are proposing the cut, not only cleaning the footage.