Technical
The ComfyUI Speed Stack: SageAttention, TeaCache and torch.compile
· RenderBob team
SageAttention, TeaCache and torch.compile can stack to a 3–5x ComfyUI speedup if you can get them installed and keep them stable. Here is what each does and where they bite.

There is a well-known trio of optimisations that can turn a painful ComfyUI video render into something you can actually iterate on. Used together they routinely deliver a multiple-times speedup, and every one of them has a catch that can cost you an afternoon.
SageAttention
SageAttention replaces standard attention with an ultra-fast 8-bit implementation, accelerating computation and cutting memory. When it is not available, tools fall back to PyTorch's default attention (SDPA), so it is additive. Flash Attention plays a similar role. Practitioners report meaningful gains over SDPA, one dropping from ~2.6 to ~1.7 seconds per iteration just by switching.
TeaCache
TeaCache (Timestep Embedding Aware Cache) is a training-free caching method that skips redundant computation across diffusion timesteps. It works across image, video and audio diffusion models and supports the ones studios actually use: Flux, HunyuanVideo, LTX, Wan, Mochi. Typical results are around a 1.5x lossless speedup, or roughly 2x if you accept minor quality degradation.
torch.compile
torch.compile JIT-compiles model blocks for a further boost, commonly cited at 20–40% on its own. Stack all three and reported acceleration climbs steeply. One optimisation suite (ANIMA_BOOSTER, measured on its Anima workflow) quotes fp16 at 1.4x, plus SageAttention 1.8–2.5x, plus TeaCache 1.5–2x, both together 2.5–3.5x, and adding torch.compile 3.5–5x, after a short warm-up.
The catches
- Installation is fragile. SageAttention and torch.compile lean on Triton, and "no module named sageattention" and "DLL load failed importing libtriton" are rites of passage. Triton and PyTorch versions must match, and corrupted Triton caches need clearing.
- Hardware matters. 8-bit and fp8 paths want modern cards; older GPUs (a V100, say) cannot use them and need fp16 model variants.
- Quality is a dial, not free. Aggressive caching smooths fine detail, exactly what a high-fidelity model is prized for. The right cache threshold is per-model, and it is a speed-versus-quality decision, not a default.
- torch.compile has a warm-up. The first couple of steps pay a compilation cost, so it helps sustained runs more than one-off generations.
These optimisations are worth a lot. A 3–5x speedup changes what you can deliver, but they are finicky, version-sensitive, and easy to get subtly wrong in ways that quietly degrade output. That is the kind of thing that should be pinned and standardised once, in a tested workflow, not re-discovered by each artist on each machine. A studio's speed stack belongs in a reproducible, version-locked environment shared across every node, owned and cloud alike, so the whole team gets the 3–5x and nobody ships a smeared render because their TeaCache threshold was too aggressive.
More from the blog
- Human Authorship Is Becoming the Unit of Value
The Oscars' 2026 AI rules, Cannes' unpublished competition stance, a patchwork of festival policies, and the Thaler copyright denial all ask how much of a film a human authored.
- Ideation Is the Green Zone: Where Studios Allow AI and Where They Stop
Netflix's generative AI guidance for production partners treats concept work as the open case and final deliverables as the restricted one. That split is spreading as a working template.