What does the [Spoken] tag actually do?
The [Spoken] tag is a section marker, similar to [Verse] or [Chorus], that tells the model to read that block of text as talked rather than sung. It works often but not always, since it is a convention the model follows rather than a strict rule. Write the lines in the lyrics box as plain sentences with normal punctuation, not rhyming couplets, since spoken word reads more like prose or slam poetry than song lyrics.
If the model starts singing a [Spoken] section anyway, try shortening the lines or adding a clearer instruction in the style box, such as "spoken word, no singing, talking voice throughout".
How do I describe a narration voice in the style box?
Use plain voice-timbre words rather than genre jargon: calm, deep, raspy, breathy, low, or close mic for an intimate, near-the-microphone feel. These are the kind of descriptors models respond to reliably, unlike vaguer words like "smooth" or claims about age, which are closer to a coin flip.
Pair the voice description with a mix note if you want a particular feel, for example "deep calm speaking voice, close mic, dry room sound" for something intimate, or "low raspy speaking voice, distant reverb" for something more cinematic and separated from the listener.
What tempo and backing music suits spoken word?
Spoken word usually sits over a slower, sparser beat than a sung hip-hop track, because the words need room to breathe and the beat should not compete with the pacing of speech. A boom bap style backing (a laid-back, sample-based hip-hop drum feel) at 85 to 95 BPM in a minor key, often G minor or D minor, works well as a default, but you can go slower and stripped-back for a more meditative reading.
| Feel | Tempo | Backing |
|---|---|---|
| Reflective, lo-fi | 85-90 BPM | dusty vinyl texture, soft drums, upright bass |
| Minimal, poetry reading | 78-85 BPM | finger snaps, sparse piano, no bass |
| Tense, narrative | 80-88 BPM | dark pads, deep sub bass, slow drums |
Can spoken word switch between talking and singing in the same track?
Yes, and it is common to talk the verses and sing the chorus, or to open with a spoken intro before the song proper begins. Mark each section clearly with [Spoken] for talked parts and [Verse] or [Chorus] for sung parts, and describe both voice styles in the style box, for example "calm speaking voice verses, smooth sung hook", so the model has separate instructions for each.
Keep the backing music consistent across both, usually the same tempo and key, so the switch between talking and singing feels like one track rather than two stitched together.
Lyrics skeleton
Original, for structure
[Spoken] I used to think the city slept at night but it just changes shift and keeps on moving every streetlight is a small confession every corner keeps a story it's not telling [Chorus] We stay awake, we stay awake waiting on a sign that never breaks
Common mistakes
- Writing rhymed, sung-style lyrics under a [Spoken] tag: write the lines as plain prose or loose free verse instead, since rhyme patterns tend to pull the model back towards singing.
- Leaving the beat too busy or loud, which buries the spoken voice: describe the backing as sparse or low and keep drums soft so the words stay the focus.
- Forgetting to describe the voice at all: without a voice description many models default to a generic sung tone, so always add words like calm, deep, raspy or breathy speaking voice.
- Mixing too many moods in one prompt, such as calm and aggressive together: pick one dominant mood and let the voice and backing both support it.
- Expecting the whole track to stay spoken with no singing creeping in: if this happens, add an explicit note such as "no singing, talking voice throughout" in the style box.