Where do backing vocals actually go: style box or lyrics box?
Both, but for different jobs. The style or prompt box tells the model there should be backing vocals at all, and roughly what kind (stacked harmonies, a choir, shouted ad-libs). The lyrics box is where you write the actual words the backing voices sing, using parentheses so the model can tell them apart from the main line, for example (oh, oh) or (hold on) under a lead line.
If you only describe backing vocals in the style box and leave the lyrics box empty for them, the model tends to either skip them or invent generic filler. Writing the parenthetical lines yourself gives more control over timing.
What words reliably describe how backing vocals sit in the mix?
In our measurements on ACE-Step 1.5, words that describe loudness and vocal character moved the sound in the direction their name suggests. Useful ones for backing vocals:
- whisper — pulls a vocal quieter, useful for backing lines you want low and intimate rather than competing with the lead
- belting — pushes a vocal louder, useful if you want the backing choir to be prominent
- high voice / deep voice — sets register, useful for separating lead and backing so they do not blur together
Words like "stacked", "layered" and "choir" describe the number and arrangement of voices rather than loudness, so pair one of those with a loudness or register word for a fuller instruction.
How does this differ for funk, soul and disco-soul?
Funk backing vocals are usually short, shouted, rhythmic ad-libs that answer the lead line by line, sitting around 90 to 120 BPM with punchy bass and tight drums underneath. Soul ballads favour slower, sustained harmonies stacked behind the lead, often 60 to 80 BPM in a major key, with the backing voices arriving gently rather than shouting.
Disco-soul sits faster, 110 to 125 BPM, and backing harmonies here often sit above the lead rather than under it, closer to a bright chorus wash than a call-and-response. Naming the sub-genre in the prompt, alongside the tempo, steers the backing style along with everything else.
How many backing voices should you ask for?
Keep the description to one clear idea: a duo answering the lead, a small stacked harmony, or a full choir. Asking for several different backing textures at once ("choir, duet, ad-libs, harmonies") in a single short prompt tends to blur together into something generic rather than distinct layers.
If you want more than one type of backing across a song, for example ad-libs in the verse and a choir in the chorus, say where each one happens: "backing choir on chorus only" is more specific than "backing choir" alone.
Lyrics skeleton
Original, for structure
[Verse] I walk this road you used to know (oh, you used to know) The porch light on, the wind is slow (so slow, so slow) [Chorus] Hold on, hold on to me (hold on, hold on) This house still hums your melody (still hums, still hums)
Common mistakes
- Leaving backing vocal lines out of the lyrics box and expecting the style box alone to generate them — write the actual (parenthetical) lines yourself.
- Asking for choir, duet and ad-libs all in one short prompt — pick one backing style per prompt for a cleaner result.
- Forgetting to say where the backing vocals happen (verse, chorus only, whole song) — the model may apply them everywhere or nowhere.
- Using loudness or reverb percentages to control how backing vocals sit — these numbers made no measurable difference in our tests, so use words like whisper or belting instead.
- Naming a real singer or group for the harmony style — describe the voice and mix instead, since artist names are excluded from prompts.