How to Make a Singing Ad With AI in Under 25 Minutes
Ori Silver
·
Co-founder, Maxfusion AI
·

Singing ads usually take a ridiculous amount of back and forth to make properly. You script the song, get the visuals right, keep the character consistent, and then stitch the whole thing together.
I wanted to see how much of that I could cut with Opus 5.5 and the Maxfusion MCP, so I built a skill that takes me from an idea to a finished singing ad in under 25 minutes.
This guide walks through the exact process: why the script and the song matter more than most people think, the mistake that kills retention on these ads, and how most of the production runs on autopilot.
The finished ad first
The ad above is a commercial for Maxfusion AI. It was made with the exact process in this article.
The process at a glance
1. Script: write a normal ad script, then turn it into lyrics.
2. Song: generate the song with the vocals starting at second zero.
3. Timing map: find the exact second every word is sung.
4. Visuals: a character sheet, then a shot sheet for every shot.
5. Videos: turn each shot sheet into a clip with Seedance 2.0.
6. Edit: cut the clips to the timing map and lay the song on top.
Every part has a few specific things you have to get right. If you miss them, you go in circles and waste a lot of money on generations you throw away.
Stage 1: the script
If you ask Claude or Codex to “give me a singing ad for my product,” you get a jingle back with verses, a chorus and rhymes that don’t sell anything.
Think of the singing as the format and nothing more. The script you need is the same script you’d write for a spoken ad: a hook, the problem, why other options failed, your product, and a call to action. Write that first, then turn it into a song.
The more specific the script is, the better the ad gets. Name the exact problem your customer has and exactly what your product does about it.
This assumes you’ve already done your research. Know which sub-avatar you’re talking to and what awareness level they’re at, because an unaware viewer and a problem-aware viewer need different scripts.

Stage 2: the song
If you don’t know much about music genres, expect to experiment. You’ll generate a lot of versions before one works.
Every time you listen to a version, check two things:
1. You can understand every word on the first listen.
2. The pacing holds. Songs stretch and squeeze words to fit the melody, and in an ad that can mean the line that sells your product gets rushed past.
Stay away from genres that are hard to listen to, like heavy metal. For the Maxfusion ad, R&B at around 120 BPM worked best. Your script and your product change what fits, so test it.
If you don’t know where to start, reverse engineer ads that already work. Find a singing ad you like, figure out its genre, and use that as your starting point. Most singing ads are made with Suno, and Suno lets you upload an audio file and remix it, so you can upload an existing ad’s audio and use it to reverse prompt the style you want.
I made the song through the Maxfusion MCP without leaving the chat, and that step is built into the skill.
The mistake that kills retention
Songs usually open with a few seconds of music before anyone sings. In an ad, those are the seconds your viewer scrolls away.
When you write the music prompt:
1. Tell it the vocals start at the very first second, from 00:00.
2. Keep anything that suggests an opener, like a build-up, out of the style prompt. It should sound like you dropped in halfway through the song.
However good a song sounds, if the singing doesn’t start at second zero, the ad will fail.
The song is the backbone of the ad, because it is the script. You can regenerate any shot and recut any clip later, but if you change the song, you redo everything.
Stage 3: the timing map
Before you make a single image, let Claude listen to the song and write down every word with the exact second it’s sung.
Turn that into a timing map: a table with every shot, what happens in it, its start and end time, and the words sung during it.
That table decides what each shot shows and how many seconds each clip needs. Skip it, and your shots won’t line up with the song, and you’ll fix it by hand in the edit.

Stage 4: the visuals
Make these ads with image-to-video. Text-to-video can’t keep a character consistent between shots.
Start with a character sheet: the same character four times on a plain background, a close-up of the face, a full body front view, and two three-quarter views.

Here’s my full character consistency video and prompt.
There are two ways to build the shots.
1. One image per shot. You animate every image separately. It works, but you need more generations, and every new image is another chance for your character to drift.
2. Shot sheets. One image holds four or six frames that describe everything that happens in the shot, including the cuts and camera moves. Every frame comes from the same image, so you get the best consistency possible.
The skill uses shot sheets.

Here’s the full storyboard method for the shot sheets.
Stage 5: the videos
The video model makes a big difference here. When you give a lot of models a shot sheet, they animate the sheet itself, and you get a video of a grid of images.
Seedance 2.0 can do that too, but when it happens it usually lasts half a second to a second at the start, and you can cut it out. Prompted correctly, it follows the frames in order and builds the scene from them, so you get a lot more control.
Set the clip lengths from your timing map. For the Maxfusion ad, most clips were eight seconds and some were shorter.
On the Maxfusion MCP, you can generate every clip at once, in parallel. Send one or two first to check that the prompts and shot sheets turn into good videos, then send the rest.
The clip above was made from the shot sheet in Stage 4.
Stage 6: the edit
If you’re editing by hand, it’s simple:
1. Open CapCut.
2. Drop in all the clips.
3. Mute every clip.
4. Put the song on top.
5. Check that every shot matches the words.
I let Opus 5.5 do the edit instead, using the same timing map. It already knew which words are sung on which second and which shot goes where, so when it stitched the clips and laid the song over them, the ad was 99% done and everything lined up the first time.
A more complex ad might take a bit more back and forth, but it’s still much faster than editing by hand.
Run it as a skill
You don’t have to run this whole process yourself. Connect Claude to the Maxfusion MCP, activate the singing ad skill, and it asks you for your script. Every step after that happens in the same chat, so you never switch to another tool.
