ChatGPT Images 2.5 Makes Editing the Main Workflow
OpenAI’s new image model keeps reference subjects recognizable, follows local edits without redrawing everything, and holds changes stable across a long conversation. Here is what changed for people who actually iterate on images.
AI-assisted draft. Reviewed and edited by the Phosphene team before publication.
On this page
- Reference fidelity: the subject survives the edit
- Local edits: change one thing, leave the rest alone
- Multi-turn editing: the conversation remembers
- Sketch: draw the layout, type the style
- Templates and prompt sharing: reproducibility as a feature
- Two API models: Flare for speed, Sunburst for finish
- What the update does not fix
- What actually changes for your workflow

Everyone who makes AI images has been here. You generate something really good, then want one small fix: the dress should be red, the background should be a street at dusk, the hand is wrong. You ask, and the model changes the face, moves the whole composition, or slowly erases the qualities you liked. Regenerate too many times and you end up comparing drafts, hoping one of them still has the original charm.
OpenAI announced ChatGPT Images 2.5 on September 8, 2026 with a response to exactly that problem. The model is aimed at editing as the main workflow, not one-shot generation: keep the subject recognizable, change only what was asked, and hold earlier edits stable while the conversation continues. The Japanese creative community picked up on the same point immediately, because it is the pain they feel daily on character art and key visuals. This article is an adaptation of that report, checked against OpenAI's own announcement.
Reference fidelity: the subject survives the edit
The first headline change is fidelity to reference images. If you feed the model a photo of a person or a place, Images 2.5 is supposed to carry that identity into new scenes, styles, and compositions. The official examples show a child in a printed studio portrait being reposed and restyled while the face, the pose logic, and the blue backdrop stay intact.
That sounds like a quality tweak, but it changes how you work. Previously, keeping a character consistent across variations meant prompt gymnastics: describe the face, the hair, the outfit, the pose, the lighting, every time, and pray. With reliable reference anchoring, the reference image becomes the source of truth and the prompt becomes the delta: what is different this time. Teams building on the API get the same benefit, because reference-led variations stay anchored to the original instead of drifting a little further on every generation.
Local edits: change one thing, leave the rest alone
The bigger shift is precision editing. Images 2.5 can modify only the instructed elements while preserving the person, composition, background, and design treatment around them. OpenAI's examples include editing a single ticket in a complex layout and full-body outfit changes on a person without rebuilding the scene.
This is the update anime and commercial illustrators have been asking for. The classic workflow loop goes: generate, spot a flaw, regenerate everything, spot a new flaw. With local editing, the loop becomes: generate, point at the flaw, fix just that. For API users it means updating one element, like a product, a background, or a piece of copy, without rebuilding the whole asset.
There is a nuance here. "Edit only what was asked" is a claim about reliability, not a guarantee. The same instruction can still touch more than you intended, especially with complex scenes. Treat local editing as a strong default and still check every diff, the way you check a regenerate.
Multi-turn editing: the conversation remembers
The third improvement is the one most people overlook. Images 2.5 holds changes across multiple turns of the same conversation. You can ask for an evening background, then a yellow jacket, then remove the object in the character's hand, and the model should carry all of that forward instead of silently un-doing earlier requests.
Japanese creators described this as the difference between a vending machine and a workshop. Old image tools treated every generation as a fresh roll of the dice. A multi-turn session behaves more like a revision history: each new edit builds on the state you already approved. For character design, where you juggle hairstyle, expression, costume, pose, and background across one session, that removes a huge amount of re-explaining. You describe the world once, then talk about what changes next.
Sketch: draw the layout, type the style
Images 2.5 also ships a feature called Sketch. Inside ChatGPT you can draw roughly on the canvas, then ask the model to turn the scribble into a finished image. In ChatGPT, @Sketch opens the tool; you draw the layout, the pose, the placement of elements, then add text for style, lighting, and mood.
The Japanese review captured why this matters: words are terrible at describing spatial arrangement. Try writing a prompt that says exactly where the title goes, how three product shots are ordered, and what the bottom strip does. A thirty-second sketch does it faster and more precisely. The sketch is the skeleton, the prompt is the skin. Do not describe geometry in text, and do not trust the crayon with your art direction. Draw the structure, write the style, let the model reconcile the two.
Templates and prompt sharing: reproducibility as a feature
Two features target the blank-page problem. Templates give starting structures for common formats like posters, flyers, product photos, and merch: choose a format, answer a few questions, and the model builds within an expected layout instead of inventing one. Prompt sharing lets you publish a generated image together with the prompt behind it, so someone else can run their own content through the same visual rules.
Both are worth taking seriously. A template is not a guarantee of taste, and shared prompts can spread a mediocre look as fast as they spread a good one. But as a primitive, prompt sharing is a small design system: one person establishes the reference, everyone else pulls from the same source. For teams and client work, that is a missing primitive that has been awkward to do with screenshots.
Two API models: Flare for speed, Sunburst for finish
Developers get two new API models. GPT-Image-2.5 Flare brings the quality, editing, and speed improvements to the API and is the obvious default for most pipelines. GPT-Image-2.5 Sunburst trades longer generation time for extra precision in detailed creative work. OpenAI reports latency up to 50% lower than GPT-Image-2 for the main path.
The practical split is iteration versus finish. When you are exploring, you want cheap fast rounds: generate, glance, adjust. That is Flare's job. When you are locking a hero asset, a poster, a product shot, a final illustration, slower and more precise beats fast and approximate. Sunburst is the "it is almost there" button, not the daily driver. Pricing on both can change, so check the official pricing page before building a pipeline on assumptions.
What the update does not fix
OpenAI claims sharper details, more natural lighting, richer textures, and better handling of complex layouts, transparent backgrounds, and visual styles. Fine. What the announcement does not claim is that text rendering is now flawless. Posters, infographics, and tickets with text still need a human pass: check the letters, the numbers, and the proper nouns before shipping. The Japanese report is explicit about this, and it is the right instinct. Generator examples are marketing, not a certification.
Multi-turn consistency has the same caveat. It is improved, not absolute. Long conversations still accumulate small drifts, so keep approving each step and treat the session as an evolving draft, not a finished file.
What actually changes for your workflow
Put the improvements together and three workflow shifts appear.
First, generation stops being a lottery. Reference anchoring plus local editing means you can protect what you like and change what you do not. The base image becomes the thing you defend, not a draft you gamble on.
Second, iteration becomes a session instead of a series of fresh starts. Multi-turn editing changes the rhythm: describe the world once, then converse about the changes. That is a much more natural interface for character work and commercial art.
Third, structure beats prompting. Sketch for layout, templates for format, shared prompts for repeatability. The freeform prompt was the minimum viable interface, not the end state, and 2.5 layers structure on top of it. The creators who build their pipeline around structure will get more consistent output than the ones typing longer paragraphs at a blank box.
The through line, in the Japanese report's words, is moving from "landing a great image in one shot" to "growing an image together with the model." That is exactly the same reasoning behind template-first generation that we build at Phosphene: exploration belongs in the freeform box, repeatable results come from structure a human reviews before pressing generate.