
Train the Identity, Then Edit the Scene
A practical LoRA dataset and captioning strategy for keeping AI characters, brand mascots, and visual assets consistent across generated images.
AI-assisted draft. Reviewed and edited by the Phosphene team before publication.
The fastest way to break a recurring AI character is to keep treating every image as a new prompt.
A prompt can describe the same person, mascot, product, or illustration style ten different times and still drift on the eleventh. The face becomes softer. The hair shape changes. The logo moves. The outfit keeps the color but loses the silhouette. At that point, the problem is not wording. The model does not have a stable identity to return to.
The fix is not a longer prompt. The fix is to train the identity first, then use editing and reference workflows to change the scene around it.
LoRA is useful when repetition matters
LoRA training is overkill for a one-off concept image. It starts to make sense when the same visual identity must survive across many outputs:
- a brand mascot used in ads, stickers, packaging, and short videos
- an AI influencer or fictional character shown in different locations
- a product line that needs consistent material, proportions, and logo placement
- an illustration style that should remain recognizable across a campaign
- a game or story character that needs expressions, costumes, and key art
For those jobs, prompt-only control is fragile. A trained LoRA gives the model a compressed memory of the subject. You can then ask for new outfits, lighting, poses, or environments without re-explaining the entire identity every time.
But the LoRA only learns what the dataset teaches it. Bad inputs create bad memory.
The dataset is the product, not a setup chore
Creators often treat dataset prep as admin work: collect images, upload them, auto-caption, train, hope. That is where consistency problems begin.
A good identity dataset has a clear contract. It tells the model which traits define the subject and which traits are incidental.
For a character, the locked traits might be:
- face shape and facial proportions
- hair silhouette and color placement
- eye color or asymmetry
- signature accessory
- recurring palette
- outfit silhouette or emblem
The variable traits might be:
- background
- pose
- camera distance
- expression
- lighting
- temporary props
If the training images mix those groups carelessly, the LoRA learns the wrong lesson. For example, if every source image uses the same red studio background, the model may treat that background as part of the identity. If every image shows the same jacket, the jacket may become impossible to remove without face drift. If captions describe every background detail but never describe the stable facial markers, the training signal points at the wrong things.
The goal is not to collect the prettiest images. The goal is to collect images that separate identity from context.
A practical minimum dataset
You do not always need hundreds of images. Modern image models can learn a surprising amount from a small, clean set, especially when the subject is visually distinctive.
A practical starting set for a character LoRA:
- Neutral reference — front or three-quarter view, readable face, simple light.
- Expression variants — smile, serious, surprised, tired, angry, but same core design.
- Angle variants — front, three-quarter, profile, slight top-down, slight low angle.
- Framing variants — headshot, bust, half-body, full-body if relevant.
- Controlled outfit or context variants — only if the identity survives clearly.
For a product or brand asset, use the same logic:
- clean hero view
- side and angle views
- close-up material/detail views
- one or two lifestyle contexts
- examples where the logo or key geometry remains readable
Small datasets are less forgiving. Every image has more influence. Remove anything that teaches the wrong identity: blurry faces, heavy filters, accidental props, warped logos, inconsistent anatomy, or style experiments you do not want repeated.
Captions decide what the model is allowed to ignore
Captions are not just labels. They are instructions for what should be associated with the trigger word and what should remain editable.
Weak caption:
a beautiful anime girl standing in a room
Better caption:
trigger_character, short silver bob hair, one amber eye and one blue eye, compact black jacket with cyan trim, small fox emblem on sleeve, neutral expression, simple studio background
The better caption names the actual identity markers. It also describes the background as a background, not as the essence of the character.
For images where the background changes, caption that change explicitly:
trigger_character, same silver bob hair, heterochromia, cyan-trim jacket, sitting at a cafe table, warm window light, background cafe interior
That helps the model learn: the cafe is optional; the hair, eyes, jacket logic, and emblem are not.
For product work, the caption should protect geometry:
trigger_product, matte black cylindrical bottle, centered white label, copper cap, embossed circular logo, photographed on gray stone surface, soft side light
If the label text or logo must stay stable, say so consistently. If the surface should be variable, describe it as context.
Auto-captioning is useful, but not neutral
Vision-language models can speed up captioning, especially inside ComfyUI-style preprocessing flows. They can inspect images, write captions, normalize sizes, and generate matching .txt files for each image.
That is useful. It is not a replacement for art direction.
Auto-captioners often describe everything they see. That can create noisy captions:
- random background objects become part of the training signal
- temporary clothing details get over-emphasized
- camera artifacts or lighting accidents get repeated
- the caption misses subtle identity markers a human would protect
Use auto-captioning as a first pass, then edit with a ruthless question:
Would I want this detail to appear again when I use the LoRA?
If yes, keep it. If no, either remove it or label it as context.
A clean captioning pass usually does three things:
- Repeats the trigger word and core identity markers across every image.
- Names variable traits only when they appear in that specific image.
- Avoids over-describing irrelevant props, backgrounds, or accidents.
Train on the flexible model, generate on the fast model
Some modern workflows split the training model and the generation model. The source workflow uses the idea of training on a more flexible base model, then deploying the learned identity through a faster generation variant.
The principle matters more than the brand name:
- use the model variant that learns cleanly for training
- use the model variant that generates quickly and predictably for production
- test the trained identity with conservative edits before pushing complex scenes
Do not jump from training straight into a crowded cinematic poster. First run identity checks:
- neutral portrait
- three expressions
- one outfit change
- one background change
- one lighting change
If the face or product geometry fails on those simple tests, the dataset or captions need work. More dramatic prompts will only hide the problem until later.
Edit the scene after the identity is stable
Once the LoRA can hold the subject, the workflow changes. You stop asking the model to invent the person and start asking it to edit around a known identity.
A strong editing prompt separates the locked identity from the desired change:
Use trigger_character as the same person. Preserve face shape, hair silhouette, eye colors, jacket proportions, and fox sleeve emblem. Change only the environment to a rainy neon street at night. Match wet pavement reflections, cool blue-magenta light, shallow depth of field, and realistic contact shadows.
For a product:
Use trigger_product as the same bottle. Preserve geometry, cap shape, logo placement, label proportions, and matte material. Place it on a bathroom counter with soft morning window light, subtle reflection, and realistic contact shadow. Do not warp the typography.
This structure is boring on purpose. It gives the model a job: protect identity, change context, reconcile the image physically.
Where Phosphene fits
Phosphene is not a LoRA trainer. The useful connection is earlier in the workflow: building the character bible before training.
Use the same tag discipline you would use for a Phosphene generation session:
- Build a core identity group: face, hair, silhouette, palette, signature accessory.
- Build variable groups: expression, outfit, camera, lighting, background.
- Generate or collect a clean reference set where identity is visible and context changes deliberately.
- Use that reference set as the basis for captions and LoRA training.
- Bring the trained identity back into an editing workflow for campaign images, product scenes, or character variants.
The value is not that every step happens inside one app. The value is that your visual decisions stay structured. A good LoRA workflow starts before training, with a clear idea of what the identity actually is.
The consistency test
Before treating a trained LoRA as production-ready, run a small acceptance test.
Ask:
- Does the subject remain recognizable without reading the prompt?
- Do the locked traits survive expression changes?
- Can the background change without dragging identity with it?
- Can the outfit change without changing the face?
- Does the style stay stable when camera distance changes?
- Does the model preserve logos, emblems, or product geometry where required?
If two or more fail, do not fix it by writing bigger prompts. Fix the training set. Remove noisy images, improve captions, add missing angles, or separate identity from context more clearly.
The takeaway
Character consistency is not a single feature. It is a pipeline.
Define the identity. Build a small, clean dataset. Caption for what must stay and what can change. Train the LoRA on the right model. Test simple edits before complex scenes. Only then use image editing to place the subject into new worlds.
That is how AI visual work moves from lucky one-off generations to repeatable creative production.