Technical

Why Scheduling ComfyUI Nodes Across Local and Cloud GPUs Is So Hard

· RenderBob team

Sending each ComfyUI node to whatever GPU is free sounds simple. It is not. The execution model, model-switch cost, data locality and mixed VRAM make node scheduling across local and cloud genuinely hard.

Mixed GPUs, model cargo, data transfers, memory limits, and queue conflicts form a difficult routing maze with one optimised path.

"Just send each node to whatever GPU is free" sounds obvious and turns out to be hard. If you have tried to spread a ComfyUI workload across two local cards, or a local card plus a cloud worker, you have probably felt the friction without quite being able to name it. Here is the problem, drawn from how ComfyUI actually executes and what practitioners keep running into.

The execution model is one-workflow-at-a-time

ComfyUI accepts jobs into a server-side queue, and a single ComfyUI process executes one workflow at a time. One process is one runner. Within a graph, nodes execute sequentially. There is no built-in fine-grained scheduler handing individual nodes to whichever GPU is idle. True concurrency means running multiple processes, each pinned to a different GPU, or reaching for community multi-GPU extensions. So schedule nodes across GPUs is fighting the grain of the tool from the start.

Switching models has a large, often hidden cost

A GPU can only hold so much, so serving different models means loading and unloading weights, and that switch cost is brutal. Users have reported checkpoints unloading after every generation, with subsequent runs around 155 seconds. A 2026 Dynamic VRAM regression made models reload on every prompt after a mid-session workflow or model switch, turning ~25-second runs into 120–190-second ones. Any scheduler that moves work around a fleet pays this cost every time it lands a node on a GPU that does not already hold the right weights.

This is a known, formally hard trade-off

The general version (one GPU, several model classes, a switch cost to change which model is loaded) has an established tension: high utilisation wants long uninterrupted runs on one model to amortise the switch cost, while avoiding starvation wants responsiveness to whatever is queued for the other models. You cannot maximise both. Every real scheduler is choosing a point on that trade-off, and the obvious greedy approach thrashes.

Data locality pins nodes in place

A node can only run where its weights already are. In a heterogeneous cluster, that means every machine effectively needs every checkpoint and VAE the workflow might use, and modern model sets run to hundreds of gigabytes. You cannot freely place a node on any idle GPU if that GPU does not have the multi-gigabyte model the node needs. Staging it first can cost more than the render.

Heterogeneous VRAM means not every node fits every GPU

Mixed fleets are the norm: a 24GB card here, a 12GB card there, a unified-memory box, a cloud worker. A node that needs 22GB simply cannot run on the 12GB card. A scheduler has to reason about per-node VRAM against per-GPU capacity, not just who is free. Even device selection is fiddly: ComfyUI defaults to cuda:0 unless told otherwise.

The memory system itself is dynamic and sometimes unpredictable

Offloading trades VRAM for system RAM and latency, so a node's real cost depends on RAM pressure and paging, not just the model size. Behaviour shifts with driver versions. One 2026 report had a driver update turn normal model loading into slow cache-streaming. And note that the popular memory-splitting extensions (ComfyUI-MultiGPU, DisTorch) improve memory management, not parallelism. Steps still run sequentially. A scheduler cannot assume a stable, knowable cost per node.

Remote adds a whole layer on top

Send a node to a cloud worker and you inherit network latency, data egress, cold-start time on the remote GPU, and the security overhead of not exposing anything to the internet. Batch-parallel tools like ComfyUI-Distributed help by fanning independent work across workers, but they do not pool VRAM, and static distribution leaves you with straggler problems when one GPU is slower.

What would actually solve it

A scheduler that models the things that make this hard: which weights are already resident on which GPU, each node's VRAM footprint against each device's capacity, the switch cost of loading a new model, the local-versus-remote boundary with its latency and egress, and the utilisation-versus-responsiveness trade-off. That is a control plane, not a frontend toggle. Until someone builds it, spread it across the GPUs stays a good idea that thrashes in practice.

More from the blog

All posts