
AI Compositing That Does Not Look Pasted: A Practical Image-Editing Pipeline
A practical workflow for using AI image-editing models, VLM prompt refinement, and reference images to make subject-background composites look physically believable.
AI-assisted draft. Reviewed and edited by the Phosphene team before publication.
Most AI composites fail before anyone studies the details. The subject may be sharp, the background may be beautiful, and the prompt may contain all the right words, but the final image still feels pasted together. The light direction does not match. The floor contact is weak. The skin texture belongs to one camera while the background belongs to another. The result is not obviously broken, but it is not believable either.
The fix is not to write a longer prompt.
The fix is to treat compositing as a pipeline: separate the subject, background, editing model, prompt-refinement layer, and final realism check. Modern image-editing models can do much of the manual masking and relighting work that used to require Photoshop skill, but they still need a clean structure around them.
The useful shift: edit with references, not only text
Text-to-image is good for invention. Compositing is different. You are not asking the model to imagine a scene from scratch; you are asking it to preserve parts of existing visual evidence while changing their relationship.
That means the input stack matters:
- Subject reference — the character, product, person, prop, or object that must survive the edit.
- Background reference — the environment, surface, room, landscape, or plate image that sets the physical world.
- Editing instruction — what should change and what must not change.
- Style constraint — whether the final image should read as editorial photography, cinematic still, anime compositing, product mockup, or something else.
- Quality pass — a second check for light, scale, shadows, reflections, and contact edges.
When creators skip that separation, they usually get one of two failures: the model ignores the source image, or it preserves the source so literally that the subject looks glued onto the scene.
A better workflow gives each model a smaller job.
A three-step AI compositing pipeline
A practical version looks like this.
1. Prepare the visual contract
Before running an edit model, define what must stay stable. This is the part many people leave vague, then blame the model when the output drifts.
Write a short preservation contract:
Preserve the subject identity, outfit silhouette, face shape, color palette, and camera angle. Replace only the surrounding environment. Match the new scene's light direction, focal length, grain, and shadow behavior. The result should look like a single photograph, not a cutout.
For product work, the contract changes:
Preserve the product geometry, logo placement, material finish, and brand color. Place it into the new environment with realistic contact shadows and reflections. Do not distort the label or change the product proportions.
This is not just prompt decoration. It tells the system which pixels are negotiable and which are not.
2. Use the edit model for physical integration
The source article focuses on a Krea 2 + Boogu Image style workflow inside ComfyUI: Krea 2 provides strong photoreal generation and texture behavior, while a fast image-editing model handles the compositing operation. The exact toolchain will change, but the principle is stable: use a model that is good at realism for the image language, and a model that is good at editing for the transformation.
In practice, the edit pass should solve four physical problems:
- Lighting — the subject needs to inherit the scene's key light, fill, contrast, and color temperature.
- Scale — the subject must sit at a believable size relative to doors, furniture, hands, horizon lines, or surface texture.
- Contact — feet, product bases, wheels, fabric edges, and shadows need to touch the environment convincingly.
- Camera language — focal length, depth of field, grain, blur, and perspective should feel like one capture system.
If the model only pastes the subject into the background, the result is still collage. A good edit pass makes the background influence the subject and the subject influence the background.
3. Run a realism pass, not a style pass
The last pass should not ask for "more detail" or "higher quality." Those words often push the image into the glossy AI look you were trying to avoid.
Use a realism pass that targets specific defects:
Refine the composite so the lighting, shadows, reflections, surface contact, lens behavior, and texture grain are consistent across the whole image. Do not change the subject identity. Do not redesign the outfit or product. Remove cutout edges, mismatched sharpness, and artificial smoothing.
This final pass is where a Krea 2-style photoreal model is useful. The goal is not to invent a new image. The goal is to reconcile the image into one believable visual system.
Why VLM prompt refinement helps
A vision-language model is useful here because compositing is partly visual diagnosis. A human may write "make it natural," but a VLM can inspect the input and generate more precise instructions: the background is warm backlit dusk, the subject is too front-lit, the floor contact needs a longer shadow, the edges are too crisp, the lens compression does not match.
The strongest workflows do not use one giant model for every step. They use a hybrid strategy:
- a faster model for setup, captions, routing, and simple prompt expansion
- a stronger VLM for the important edit instruction
- an image-editing model for the actual transformation
- a photoreal model or refinement pass for final integration
That matters because creative work has rhythm. If every tiny decision waits on the heaviest model, the workflow feels dead. If every important decision is delegated to a lightweight model, the output becomes sloppy. The trick is to spend intelligence where it changes the image.
The compositing checklist
Before accepting an AI composite, check it like an art director, not like a prompt engineer.
Ask these questions:
- Does the subject cast the right kind of shadow? A soft studio shadow does not belong on a harsh noon street.
- Do highlights come from the same direction? Look at cheekbones, glass, metal, glossy fabric, and product edges.
- Is the sharpness consistent? A razor-sharp subject on a soft background is one of the fastest ways to reveal a fake.
- Does the contact point work? Feet, chair legs, product bases, tires, and hands need believable pressure.
- Is the scale anchored? Use known objects in the background to judge whether the subject is too large or too small.
- Did the model preserve identity? If the face, logo, silhouette, or product geometry changed, the edit is not successful.
- Does the whole image share one camera? Grain, blur, focal length, contrast, and color should feel like the same capture.
If one of those fails, do not regenerate from scratch. Isolate the failure and ask for a targeted correction.
Good prompts are boring and specific
Weak compositing prompt:
Put this character in Tokyo and make it realistic.
Better compositing prompt:
Place the subject into a rainy Tokyo street at night. Preserve the character identity, hairstyle, jacket, and camera angle. Match the background's neon reflections, wet pavement, soft lens bloom, and cool color temperature. Add realistic contact shadows and reflected light from the street signs. Do not change the face or outfit design.
For product imagery:
Place the product on a matte stone bathroom counter beside soft morning window light. Preserve the bottle shape, logo, cap, label text, and brand color. Match the counter perspective and add realistic contact shadows and subtle reflection. Do not warp the typography.
The point is not poetry. The point is constraint.
Where this fits in a production workflow
AI compositing is strongest when you need fast visual exploration without losing control:
- product mockups in multiple environments
- character key art with different locations
- social campaign variants
- e-commerce lifestyle images
- storyboard frames for AI video
- poster concepts that combine generated subjects and designed backgrounds
It is weakest when used as a one-click deception machine. The more photoreal these tools become, the more important disclosure and context become. If a composite could be mistaken for documentary evidence, label it clearly and follow platform rules. For commercial concept art, mockups, and visual development, the value is speed and iteration — not pretending the generated scene is real.
How to use the idea inside Phosphene
Phosphene already encourages the same kind of separation: subject, style, lighting, camera, environment, and model choice can be treated as separate parts of the creative direction instead of one fragile paragraph.
A useful loop is:
- Build the subject and style direction first.
- Generate a clean baseline image.
- Change only the environment or lighting layer.
- Compare variants side by side.
- Keep the prompt structure that preserved identity best.
- Use that structure as the input contract for external edit/refinement tools when needed.
That is the real lesson from the new compositing workflows: good AI image editing is not magic. It is disciplined decomposition. The more clearly you separate what must stay, what should change, and what physical rules must be reconciled, the less your final image looks like an AI cutout — and the more it looks like a deliberate visual direction.