
From Phone Snapshot to ID Photo: Spec-First Prompting for Nano Banana Pro
A Japanese writer needed an ID photo the night before a deadline and tried to rescue a casual phone snapshot with Nano Banana Pro. Four failed rewrites and forty lost minutes later, the lesson is clear: give an image editor measurable specs, not genre labels. A practical adaptation, with the boundaries you should not cross.
AI-assisted draft. Reviewed and edited by the Phosphene team before publication.
The entry sheet is due tomorrow and the photo field on your resume is still empty. The photo booth at the station has a line, a studio booking costs thousands of yen and keeps business hours, and a random selfie will not do. That is the moment you start searching "can I fix this tonight, at home, for free."
A Japanese writer who publishes as planetdive hit exactly this wall in August 2026, with one asset: a phone photo a family member had snapped against the living room wall a few weeks earlier. The reasonable idea was to feed it to Nano Banana Pro and ask for an ID photo. What happened next is a better lesson than most prompting guides, because the failure had nothing to do with the model's capability.
Four rewrites that fixed the wrong thing
The first instruction was the obvious one: "make this an ID photo." The background turned white, but the head stayed off center, which is an instant rejection for any formal spec. The next attempt added "center the face," and the headroom above the hair collapsed. Two more rewrites followed. About forty minutes gone, and the reason is worth spelling out: every correction was aimed at a symptom, while the actual spec, where the head sits, how tall it is, how much empty space belongs above it, was never stated once.
This is the genre-label trap. "ID photo" is a category name, and a category name averages over thousands of variants. Every country, every document type, and every institution slices the geometry differently, so the model guesses, and you iterate against its guess. Instruction-following editors respond to geometry, proportions, and explicit constraints far more reliably than to labels, because the edit-training data rewards doing exactly what was asked, and "look formal" is not a thing anyone can do exactly.
State the spec, not the genre
The fix is to write the specification into the prompt as measurable statements. Four dimensions cover nearly every ID photo format in use:
- Head geometry. Head height as a share of frame height, vertical centering, and top margin. A US passport photo, for instance, wants the head between 1 and 1⅜ inches tall inside a 2×2 inch frame, which works out to roughly 50 to 70 percent of frame height with clear space above the hair. Japanese resume photos and visa formats land near the same zone. Pick the numbers for your target document and say them.
- Background. Plain, evenly lit, white or light gray, no gradient, no shadows, no seam. This is the part everyone gets for free, and the part that fails subtly when the model "studio-ifies" the whole scene instead of replacing the backdrop.
- Lighting. Soft, frontal, even. The goal is not flattering light, it is shadowless light, because shadows on the face and background are the most common reason a formal photo reads as fake.
- Preservation. The face itself does not get edited. No beautification, no jaw adjustments, no skin smoothing, no "professional look" upgrades. You are reformatting a portrait, not generating a person.
A skeleton that carries all four looks like this:
Convert this photo into a formal ID photo.
Keep the person's facial features exactly as they are.
Do not beautify or alter the face in any way.
Head vertically centered, head height about 65 percent of frame height,
with empty space above the hair about 10 percent of frame height.
Replace the background with a plain evenly lit white backdrop.
Soft frontal lighting, no shadows on the face or background.
Neutral expression, shoulders square to the camera.
Change one dimension at a time when you iterate. If the head drifts, fix the head with a coordinate statement, not with "make it more centered," which invites the model to re-guess everything else. The planetdive author's forty minutes came from fixing composition words while never naming composition numbers.
Where this method ends
The original article is unusually honest about who should not use it, and those limits deserve to travel with the technique.
Official documents with strict likeness checks, passports, national ID cards, driver's licenses, are off limits. These exist to confirm you are you, and if an edit shifts facial structure even slightly, the document fails at the counter or worse, succeeds in a way that creates problems later. Some employers and programs also explicitly require unretouched photos with a known capture date. Respect that, it is a stated contract, not a style preference.
The third limit is the source photo itself. The model normalizes a portrait into a spec, it does not recover information that was never captured. A blurry, backlit face produces a confidently blurry, backlit ID photo. Garbage in, formal garbage out.
The economics that make it worth learning
A booth session costs a few hundred yen and queues. A studio costs thousands and closes at six. A single targeted edit costs a fraction of both, takes one minute, and can be rerun for every application cycle from the same snapshot. For anyone job hunting or switching roles, who needs the same compliant photo submitted repeatedly, the spec-first edit stops being a hack and becomes the sensible default, exactly the kind of everyday editing task conversational editors are built for. If you want to see the broader pattern, we covered where this interface style is going in conversational image editing.
One provider-side check first: an ID-style portrait is identifying data about you, so glance at how the tool handles uploaded images, retention and training use, before sending one, and get consent before running anyone else's face through it. With that checked, Phosphene is a fair test bed: run the same spec-first prompt against two models, Nano Banana Pro and GPT Image, and keep the one that nails the headroom on the first pass. Spec compliance is measurable, so it is the rare editing task you can genuinely A/B.