News

Wan Animate 2: Drop Any Character Into Any Performance, No Skeleton Required

ยท RenderBob team

Wan Animate 2 transfers motion from a driving video onto a reference character with no skeleton extraction, and can swap a character into footage while keeping lighting and camera.

A performer's motion passes directly through a rig-free transfer core into a new character while camera, lighting, and environment remain intact.

Wan Animate 2, natively supported in ComfyUI, does two things a motion studio has historically needed a rig, a mocap pipeline, or hours of manual rotoscoping to achieve. Hand it a still character reference and a driving video, and it either animates the still to copy the performer's motion and expression frame for frame, or swaps the character already in existing footage for your reference character while preserving the original lighting, colour tone and camera. Both jobs run through a redesigned Diffusion Transformer that consumes driving footage directly, with no intermediate pose extractor or skeleton sitting between the source video and the model.

Prior motion-transfer approaches typically ran a driving video through a pose-extraction step first, OpenPose or similar, converting it into a skeleton the animation model then followed. Wan Animate 2 removes that step and consumes raw driving frames directly. The model's authors report this improves both motion fidelity and identity preservation, since a skeleton extractor is itself a place where subtle motion information, weight shifts, secondary movement, facial nuance, gets lost. Text-driven viewpoint control further decouples the output camera from the driving video's camera, so a studio can restage a performance from a different angle than it was shot at.

Two nodes ship with it: WanAnimate2ToVideo, the conditioning node taking the reference character and driving video, and WanAnimate2Cache, which caches the pose branch to roughly halve generation time, another entry in the caching techniques this blog has covered as part of the speed stack. A Lite variant, aimed at real-time streaming character animation, has been announced. Independent testing notes the officially released weights so far are Base and Distillation checkpoints, so treat Lite's real-time claim as directional until a published weight or a reproducible benchmark confirms it.

Character-replacement work that used to mean either reshooting with an actor or a slow rotoscope-and-composite process becomes a single conditioning pass. Independent testing has flagged fast action and severe motion blur, occlusion, hand detail, multi-character scenes, and large viewpoint changes that require inventing unseen detail. Those are the areas worth stress-testing against your own production shots before committing a client deliverable to the technique, and the kind of thing a fixed benchmarking protocol is built to catch before a client does.

More from the blog

All posts