GPT Image 2.5: Flare vs Sunburst, Credits, and the Spatial Consistency Audit
GPT Image 2.5 ships as two models with different jobs, and the credit math reshuffles which flagship models make sense. A Japanese studio test shows how to audit multi-panel spatial consistency with a critic model.
AI-assisted draft. Reviewed and edited by the Phosphene team before publication.
On this page

GPT Image 2.5 arrived on September 8, 2026 as two models under one name: Flare and Sunburst. OpenAI positions Flare as a small model optimized for speed and Sunburst as the base model, its most capable image generation and editing model. Two tiers, different jobs.
Flare for exploration, Sunburst for decisions
CreativeEdge, a Japanese studio that tests generative tools on real productions, came away from its sessions with a clean operating rule: explore with Flare, finalize with Sunburst. Use Flare for the work it is good enough for, all the way through. Reach for Sunburst when the edit is precise: don't change the face, keep the product shape, change only the clothes, change only the background, run many edits in a row. Those are exactly the tasks where text-to-image models usually drift.
The corollary matters as much: Sunburst is wasted on ideation. Generating twenty to fifty composition ideas, mass-comparing rough designs, or developing prompts are Flare jobs. There is no reason to spend a stronger model's capacity on outputs you will discard.
One detail that trips people up: Flare is not a cheaper model. It consumes the same credits as Sunburst. Speed, not price, is the differentiator, so the strategy is matching difficulty to capability rather than saving money.
Resolution, quality, and where credits go
Five quality levels (Low through Max) and three resolutions (1K, 2K, 4K) multiply into a lot of combinations, and aggregators do not agree on which ones they expose. Higgsfield and Magnific offer twelve aspect ratios; Adobe Firefly offers three, with five quality levels on Firefly Board and three on the web version, where the resolution setting is missing entirely, as of September 9, 2026. If a specific resolution matters for a client deliverable, check what the tool you pay for actually exposes before promising it.
The credit math is where the economics get weird. In Adobe Firefly Boards, as tested on September 9, 2026, GPT Image 2.5 at High quality costs 5 credits and produces 1536x1024 pixels. Nano Banana Pro costs 40 credits for a comparable output and Nano Banana 2 costs 20. Even 4K at Max, which is 3840x2160, lands at 40 credits, the same amount Nano Banana Pro charges for 1K at 1264x848. CreativeEdge's conclusion is blunt: within Firefly Boards, Nano Banana Pro usage is likely to drop. At four to eight times the credit cost, the flagship would have to deliver an enormous quality gap to justify itself, and across the sessions it did not.
The spatial consistency audit
The most reusable technique in the test log has nothing to do with pricing. The studio wanted to know whether GPT Image 2.5 could render one room from four camera positions in a single 2x2 image. They wrote an incomplete, human-style prompt: one lab with a fixed layout, four angles, focal lengths included. The model returned a plausible-looking 2x2 grid, and it was wrong in a specific way: the four panels were not four views of one lab, they were four similar-looking labs with furniture placed ad hoc. Same taste, different rooms.
Then came the interesting move. Instead of regenerating blindly, they brought in GPT-6 Astra as an auditor. Its first diagnosis was brutal: the image is full of contradictions. The panels share no 3D coordinates; each one is a collage of lab symbols — dark table, windows, bottles, equipment — arranged independently. The premise of a single space collapsed.
They rewrote the brief as a verification prompt with explicit scene rules: keep the four-panel structure, treat the space as one fixed architectural room with four virtual cameras, cameras may move and rotate while furniture may not, objects may be occluded by an angle but must never reappear elsewhere, relative distances stay constant across panels, and keep the already-correct top-right angle locked.
Round two fixed the cameras and failed on a different axis: equipment positions still floated between panels. Astra's second verdict was the operational one — stop regenerating the whole image, lock the composition that works, and pinpoint-edit only the two remaining contradictions, the top-right equipment and the bottom-left fridge. The studio switched to a local-consistency prompt: preserve the entire composition, treat the top-left and the bottom-right overhead panels as the spatial ground truth, and change only what is left rather than redrawing anything. That pass converged.
Three lessons transfer directly to any multi-panel or multi-view work:
- Use a second model to find the failure, not to redo the image. The critic caught that the failure moved: first cameras, then object permanence. Blind regeneration would have paid for the same two mistakes repeatedly.
- When a composition is mostly right, stop treating the image as a draft. Switch from full regeneration to local edits and pin the parts that already work. This holds for spatial grids and for character work.
- Write spatial prompts as scene constraints, not vibes. "One lab filmed by four cameras" fails because the model resolves it as four pictures of labs. Explicit rules — one fixed layout, cameras move, furniture does not, occlusion yes, repositioning no — turn a wish into a spec.
The same discipline shows in their text-layout test: a risograph-style card brief that pinned the speech balloon to four centered lines at roughly seven percent of image height, with exact colors (indigo #2B3A67 text, brick red #B4442E for a single emphasized word, cream #F2E8D5 background) and a single tail, reproduced the layout faithfully. Give the model numbers and it follows them; leave the layout vague and you get approximate placement.
The combination play: Suno v6 and Seedance 2.5
The same session includes a quick Suno v6 run against a Japanese enka and folk benchmark, with Pro, wild, and mini variants at 15 seconds, plus a pairing where a generated 1970s idol-enka track fed a music-reference lip-sync video through Seedance 2.5. That pipeline belongs to the video side; we covered Seedance 2.5's reference surface in the Seedance 2.5 guide and its prompting rules. What the test shows for image work is the direction of the stack: when a mid-tier image model undercuts flagship credit costs, studios start pairing it with music and motion models instead of paying flagship rates for everything.
If you are calibrating which capabilities you actually need before paying for them, the AI video model comparison is a useful baseline for that decision.