Amazon's AI Dubbing Redraws the Actor's Lips, Frame by Frame

Prime Video is testing a tool that rewrites mouth movement to match a translated voice, starting with the German series Maxton Hall. The lip sync is the easy claim. The open question is who agreed to the voice, and why an almost-right face can read worse than an obvious mismatch.

AI-assisted draft. Reviewed and edited by the Phosphene team before publication.

On this page
Amazon's AI Dubbing Redraws the Actor's Lips, Frame by Frame

Dubbing has always been a compromise. A face and a voice belong to different people, and viewers have spent a century agreeing to ignore that as long as the timing roughly holds. Amazon is now testing a version where the face moves to match the new voice, and it is selling the result as immersion rather than correction.

What Amazon actually described

Prime Video, together with Amazon MGM Studios, has described an AI tool that takes the original footage as reference and alters the actor's lips and the lower part of the face frame by frame, so the mouth tracks the dubbed audio track. Raf Soltanovich, vice president of technology for Prime Video and Amazon MGM Studios, framed the goal as a "more fluid and immersive way to enjoy global content."

Two things are worth keeping apart. First, no demo footage has been published, so every claim about quality is currently a claim about a product description rather than a visible result. Second, the pipeline being described rewrites an existing performance instead of generating one: the original timing, head position, and expression stay in place while a narrow band of the face is replaced.

The pilot is the German series Maxton Hall, for its English dub, with the company saying it plans to extend the approach to more titles. That matters less as a single launch than as a direction, because Amazon has already put adjacent systems into production: AI-assisted dubbing in some markets, fully AI-generated dubbing for some anime, and automated episode recaps, some of which had to be pulled for errors. Each step was announced the same way, as a convenience, and each one drew a version of the same objection.

The technique is not new

Frame-accurate face retiming for dubbing has been shipping for years. Flawless AI sells exactly this job under the name TrueSync, and the technique is visible in its work on the film "Fall". Games arrived earlier for purely practical reasons: Cyberpunk 2077 used AI lip sync to make dialogue look correct across ten languages, because hand-animating every cutscene per language was never viable. The translated deepfake demos that circulated before the current model wave were the same idea with weaker tooling.

So there is no new capability in the announcement, which is the useful part. What changed is the placement: a first-party studio, a large streaming catalog, and a stated intent to keep going. When a technique stops being a vendor demo and becomes a line item in a distribution plan, the interesting questions move from possibility to policy and from output quality to who is allowed to be in the frame.

Why the arithmetic of the uncanny valley is worse than the pitch

The pitch assumes better lip matching equals a better viewing experience. The failure mode suggests otherwise. Audiences have decades of practice ignoring imperfect sync, because dubbing is a known convention. A slightly wrong mouth movement is an error you can see and dismiss. A face that moves almost correctly is a different category of problem, because reading faces is fast and largely involuntary. Small errors in how a mouth moves recruit attention instead of tolerance.

That asymmetry has a commercial edge. The tool repairs an error viewers had already learned to forgive and introduces artifacts they are built to detect. On a wide shot it will look like progress. In a close-up of a long line, where mouth shapes carry a lot of motion, is where these systems are at their least convincing. The realistic outcome is not "dubbing solved" but "dubbing becomes a per-shot decision": run the face pass where the framing is forgiving, skip it where it would pull the eye.

That is a production judgment, which is a different thing from a toggle in a release note. It also means the useful skill is not operating the tool. It is knowing which shots will survive it, which is the same instinct that decides when a generative pass helps an image and when it damages one you already have.

The dispute is about the voice, not the lips

The real argument is running one layer down, and mouth shapes have nothing to do with it. Spain's largest dubbing performers' union, Adoma, publicly attacked Amazon for shipping dubbing work without an adequate AI clause, and said companies train models "capable of generating voices imitating them, without permission and without compensation." The reported sticking point in those negotiations is not whether dubbing will involve AI. It is that the studios want performers to hand over rights to their recorded voices for model training, which the unions have refused. The consequence is concrete rather than theoretical: where no agreement exists, some titles simply do not get dubbed into some languages at all.

That pattern repeats well outside streaming, and it is worth stating plainly. The technology for a face pass is available to anyone with a license and a GPU. The scarce input is a performer's agreement, and it is scarce on purpose. Some jurisdictions are turning the same argument into law; Mexico's new federal cinema law (Article 29) requires dubbing of foreign audiovisual works in national languages to be performed by human performer-interpreters, which the same outlet's Mexican edition covered as a direct hit to the cheapest localization tool streaming platforms had.

There is a similar lesson two weeks old on this blog: a German broadcaster aired five episodes of a court show rewritten with image editing, voice synthesis, and lip-sync adjustment, with rights settled and legal review completed before any of it went out. Same toolkit, different use, and the same conclusion. The technology was never the constraint. The television version of that story is here, and the likeness question underneath it, where studios recast actors with AI versions of themselves, is here.

What changes if you publish video in more than one language

None of this touches your model choice. It does reshuffle a few decisions that used to sit quietly inside the localization budget:

  • Plan the source performance with dubbing in mind. Clear articulation, moderate mouth movement, and framing that does not force extreme close-ups all make a later face pass cheaper and more convincing.
  • Log which voice and which face pass produced each language variant. Consent and auditability moved upstream, and a version log is what lets you answer a question about a specific cut six months later.
  • Treat the face pass as a per-shot threshold, not a project setting. Wide shots and short lines are the safe zone; sustained dialogue in close-up is where the artifacts live.
  • Keep your own masters. The value of a catalog is that the original frames remain the reference everything else is measured against.

Buying this as a single setting is how projects get an artifact they cannot remove later. Buying it as a documented step, applied where it holds up, is how the same tool stops being a liability.

What to watch

The tool will ship, or it will not, and either way the announcement is not the signal. Two things are. One is whether the first public demo is a wide shot or a close-up, because that choice tells you what the company believes about its own output. The other is what the dubbing contracts say, since the voice clause, not the video model, decides which languages exist and what they cost. Anyone who sells video work should plan for the version where a face pass is one decision among many, priced like a step in the pipeline rather than a headline.

Sources

Share this guide
X LinkedIn