All articles
Build the World Before the Music Video: A Three-Week AI Production Post-Mortem

Build the World Before the Music Video: A Three-Week AI Production Post-Mortem

A Japanese solo creator spent three weeks on one AI music video and calls the result "60 points." The production log shows where the time went: a week of world design before any generation, a hand-built storyboard player, and a stack of five narrow AI tools instead of one. An honest adaptation, including the parts that did not work.

AI-assisted draft. Reviewed and edited by the Phosphene team before publication.

An AI music video takes an afternoon now. Write a song with one model, generate stills with another, animate them with a third, cut it together. The pipeline is real and it works.

So why did a Japanese solo creator spend three weeks on a single four-minute music video, and why is the production log worth reading?

The author, publishing as Beta, titled the piece "Before making an MV with AI, I built a world." The video is "Anemos Traveler," a JRPG-flavored travel piece with an airship, a sea of clouds, and two protagonists crossing a map that was drawn before a single frame existed. The final self-grade is 60 out of 100. The three weeks get 100. That gap is the interesting part.

The world came first, and it was not inefficient

The project started from a casual X theme event: "travel" or "otaku." Asked what "otaku" meant to him, the author landed on JRPGs, Final Fantasy airships, the feeling of a world map slowly filling in. Instead of jumping to a song, he spent close to a week on setting design alone:

  • Sky ruins and a cloud-sea harbor
  • A wind-riding airship and an enemy called the "wind-eater"
  • A travel route: starting grasslands, a mirror-water forest, a sand-bone sea, the cloud harbor, a star-crossing corridor, floating islands, and finally the sky ruins
  • Two protagonists defined by how they travel that route

Only after the route existed did the video's structure fall out of it: the MV is sequenced along the journey, stop by stop. The map was not decoration. It was the storyboard's spine.

There is a known failure mode in AI production where you generate a pile of beautiful clips and then hunt for a story to glue them together. It rarely holds. World-first inverts that: coherence is decided before generation, so every later prompt inherits it. The author's own summary is the cleanest statement of the method: AI does not create the world, it gives form to a world. The imagining stays a human job.

A tool stack with narrow, honest roles

Five tools, each doing one thing:

StageToolRole
MusicSunoCompose the song
SettingChatGPTOrganize and pressure-test the world notes
StillsGPT ImagePaint the world into key images
MotionVeo 3 FastAnimate the landscapes
EditCapCutAssemble the final cut

Two decisions in this table deserve attention. First, the author explicitly refused the "which AI is best" question and asked "which AI fits this stage" instead. That is the right question once your bottleneck moves from generation to coherence. Second, look at what is absent: no single end-to-end "make me a video" tool. Assembly stayed manual, in CapCut, where a human controls rhythm.

The stack principle travels: anchor your production to a pipeline of narrow tools with a human at the edit, and swap any stage when something better ships, without redesigning the whole production. We have argued this before from the material side in the modular AI anime pipeline, where characters, backgrounds, and audio are interchangeable kit parts.

He built his own tooling mid-production

The detail that separates this log from most AI video write-ups: halfway through, the author got tired of bouncing between lyrics, audio, cuts, and prompts, so he built an "MV Storyboard Player," a single screen to review the storyboard, manage cuts, and check the audio source together.

That is a production decision, not a hobbyist detour. When your iteration loop crosses four or five tools, the loop itself becomes the bottleneck, and investing a day in tooling pays back every day after. Studios do this with pipeline TDs. A solo creator does it with a web player and stubbornness. Same physics at different budgets.

The 60-point parts, kept honest

The post does not hide the failures, and they are the standard suspects of 2026 video generation:

  • The airship would not stay consistent across shots and changed shape throughout
  • Character unity drifted
  • Battle scenes lacked impact
  • Title card presentation was uneven

The author grades the finished video 60 and the experience 100, and plans to make another. Two lessons sit inside that split. Consistency across shots remains the hardest unsolved problem in AI video, no amount of world design fully fixes it at the shot level, even though it fixes coherence at the structure level. And process satisfaction is a legitimate production goal: the author's point is that time spent on the world was not overhead to be optimized away, it was the part that felt like the RPG he was homage-ing in the first place.

His closing argument is worth keeping intact: AI is fundamentally an efficiency tool, and making things quickly is a fine way to use it. But spending as long as you want imagining a world is a luxury this era affords, and "keeping up with AI" is the wrong frame. Find the best that is possible right now, make the thing, enjoy it.

What to steal

If you are planning an AI music video or short:

  1. Design the world and the route before generating anything. Let structure precede shots.
  2. Assign each tool a narrow role and keep a human on final assembly.
  3. When the tool-juggling hurts, stop and build (or template) the missing tool for a day.
  4. Grade the output and the process separately. Sixty-point videos with hundred-point processes get followed by better videos. The reverse does not.

The unfixed problems, airship geometry, character drift, are exactly where the next round of tools will compete, which is a good reason to keep your pipeline modular rather than married to any single vendor's answer.

Sources