Workflow

Train Your Own Look: In-House LoRAs for Brand and Character Consistency

· RenderBob team

A generic model cannot hold your client's character or house style across a campaign. In-house LoRA training can. Here is the workflow, the costs, and where the compute goes.

Curated campaign imagery trains a compact LoRA capsule that preserves one character and visual identity across many scenes.

The single hardest thing to get from a base generative model is consistency: the same character across scenes and poses, or a client's exact house style across dozens of assets. Prompting alone will not hold it. Base models only know the styles in their training data, and on-brand at scale needs more than describing a look and hoping. The production answer is a trained LoRA: a small adaptation that teaches a model your specific character or style and activates with a trigger word.

Why LoRA, not a full fine-tune

A full fine-tune of a diffusion model for one character can consume 24–48GB of VRAM and weeks of GPU time, and can flatten the very stylistic nuance you are trying to capture. A LoRA gets most of the benefit for a tiny fraction of the cost: a well-curated dataset of roughly 15–50 images and 1,000–3,000 training steps, with per-project costs on rented compute reported around a few dollars to ten.

The workflow options are converging on ComfyUI

Historically most studios trained externally (Kohya_ss) and loaded the resulting LoRA into ComfyUI. Increasingly you can train in-place: realtime LoRA-training nodes now support SDXL, SD 1.5, FLUX, Z-Image, Qwen Image and Wan 2.2, training on the fly without leaving ComfyUI. SDXL trains in a few minutes and SD 1.5 in under two on decent hardware, which makes it practical to lock a subject for consistency and immediately use it in the same workflow, including training on a video's first and last frames for Wan temporal consistency.

The craft is in the data

The failures are almost always dataset failures: too few images, or twenty photos from one sitting under one light that teach the model the lighting instead of the subject; over-training that overfits; generic captions instead of precise, consistent trigger words. Curation is most of the job, and it is where a studio's judgement shows.

Two things make this a pipeline concern rather than a solo trick. First, training is a periodic, VRAM-heavy workload, exactly the spiky, occasional demand that suits burst capacity rather than a card sitting idle most of the month. You do not buy a training GPU; you rent one when a campaign needs a new style LoRA, then give it back. Second, a trained LoRA is studio IP, a client's character or house style, encoded. It belongs in a governed, versioned model registry with clear licensing and access control, not scattered across artists' drives where it can drift, leak, or get used on the wrong client's work.

Trained models are how a generative studio stops being interchangeable. The look you can produce that nobody else can is the product. Build the capability to train, store, version and govern those LoRAs, and to summon training compute only when you need it.

More from the blog

All posts