Technical

The ComfyUI Speed Stack: SageAttention, TeaCache and torch.compile

· RenderBob team

SageAttention, TeaCache and torch.compile can stack to a 3–5x ComfyUI speedup if you can get them installed and keep them stable. Here is what each does and where they bite.

Attention optimisation, reusable cache blocks, and compiled execution paths combine into one accelerated rendering engine.

There is a well-known trio of optimisations that can turn a painful ComfyUI video render into something you can actually iterate on. Used together they routinely deliver a multiple-times speedup, and every one of them has a catch that can cost you an afternoon.

SageAttention

SageAttention replaces standard attention with an ultra-fast 8-bit implementation, accelerating computation and cutting memory. When it is not available, tools fall back to PyTorch's default attention (SDPA), so it is additive. Flash Attention plays a similar role. Practitioners report meaningful gains over SDPA, one dropping from ~2.6 to ~1.7 seconds per iteration just by switching.

TeaCache

TeaCache (Timestep Embedding Aware Cache) is a training-free caching method that skips redundant computation across diffusion timesteps. It works across image, video and audio diffusion models and supports the ones studios actually use: Flux, HunyuanVideo, LTX, Wan, Mochi. Typical results are around a 1.5x lossless speedup, or roughly 2x if you accept minor quality degradation.

torch.compile

torch.compile JIT-compiles model blocks for a further boost, commonly cited at 20–40% on its own. Stack all three and reported acceleration climbs steeply. One optimisation suite (ANIMA_BOOSTER, measured on its Anima workflow) quotes fp16 at 1.4x, plus SageAttention 1.8–2.5x, plus TeaCache 1.5–2x, both together 2.5–3.5x, and adding torch.compile 3.5–5x, after a short warm-up.

The catches

  • Installation is fragile. SageAttention and torch.compile lean on Triton, and "no module named sageattention" and "DLL load failed importing libtriton" are rites of passage. Triton and PyTorch versions must match, and corrupted Triton caches need clearing.
  • Hardware matters. 8-bit and fp8 paths want modern cards; older GPUs (a V100, say) cannot use them and need fp16 model variants.
  • Quality is a dial, not free. Aggressive caching smooths fine detail, exactly what a high-fidelity model is prized for. The right cache threshold is per-model, and it is a speed-versus-quality decision, not a default.
  • torch.compile has a warm-up. The first couple of steps pay a compilation cost, so it helps sustained runs more than one-off generations.

These optimisations are worth a lot. A 3–5x speedup changes what you can deliver, but they are finicky, version-sensitive, and easy to get subtly wrong in ways that quietly degrade output. That is the kind of thing that should be pinned and standardised once, in a tested workflow, not re-discovered by each artist on each machine. A studio's speed stack belongs in a reproducible, version-locked environment shared across every node, owned and cloud alike, so the whole team gets the 3–5x and nobody ships a smeared render because their TeaCache threshold was too aggressive.

More from the blog

All posts