# Glossary of music production and AI music terms | Lyro

> Plain-English definitions of 40 terms used in music production and AI music, from BPM and stems to LUFS, seeds and watermarks.

Source: https://lyromusic.com/glossary

# Glossary of music and AI music terms

40 terms you will meet when making music with AI, from BPM and stems to LUFS and watermarking. Each entry starts with a one-sentence definition.

## 808

**An 808 is a long, deep bass note modelled on the kick drum of the Roland TR-808 drum machine. In trap and hip-hop it plays the bassline.**

The original TR-808 kick was a low, sine-like boom with an adjustable decay. Producers stretched that decay, tuned the note to the key of the song and started playing basslines with it. Today “808” usually means that tuned sub bass, often with slides (glides between notes) and a little distortion so it stays audible on phone speakers. When you prompt an AI model, ask for it by function: “booming sliding 808 bass, 140 BPM half-time feel”. An 808 carries pitch, so it has to match the key of the track. That matters when you bring your own beat or vocal to Lyro.

See also: [Key](#key), [Log drum](#log-drum), [Prompt](#prompt)

## A cappella

**An a cappella is a vocal performance with no instruments. In production it means the isolated vocal track of a song, ready to sit over a new beat.**

A clean a cappella is the best starting point when you want a beat written around your voice. Record it dry (no reverb or echo), with no music bleeding in from speakers, and keep a steady tempo, ideally to a click in your headphones. Lyro detects the tempo and key of the vocal in your browser, then composes instruments around it. If you only have a finished song, [stem separation](https://lyromusic.com/glossary#stem-separation) can pull the vocal out, but expect artefacts: traces of reverb and cymbals often stay in a separated vocal. You also need the rights to any recording you upload.

See also: [Stem separation](#stem-separation), [Topline](#topline), [BPM (beats per minute)](#bpm)

## AI disclosure

**AI disclosure means telling listeners and platforms that a track was made with generative AI, either with a visible label or with machine-readable data inside the file.**

Rules and platform policies increasingly ask for it. Article 50 of the EU AI Act, which sets transparency duties for AI-generated content, applies from 2 August 2026 ([EU AI Act, Article 50](https://artificialintelligenceact.eu/article/50/)). Deezer tags AI tracks, Spotify supports voluntary AI disclosure in credits, and TikTok requires labels on realistic AI-generated content. Lyro writes an ID3 tag (a metadata field inside the audio file) into every download of generated audio, declaring it AI-generated. Lyria 3 Pro output also carries a [SynthID watermark](https://lyromusic.com/glossary#watermark). Policies change, so check your distributor’s current rules before you release. See [Can you monetise AI music?](https://lyromusic.com/learn/can-you-monetize-ai-music)

See also: [Watermark (SynthID)](#watermark), [Blind listening test](#blind-test)

## Arrangement

**The arrangement is the plan of a song over time: which sections come in what order, and which instruments play in each one.**

A typical pop arrangement runs intro, verse, chorus, verse, chorus, bridge, final chorus. Club tracks swap those for intro, build, drop, breakdown and outro, with long intros and outros so DJs can mix them. AI models follow arrangement cues when you write them down. In lyrics, use section tags such as [Verse], [Chorus] and [Bridge]. In the prompt, describe the energy curve: “sparse intro, full drums from the first chorus, stripped breakdown before the last hook”. A good arrangement adds or removes something every 8 or 16 bars, so the ear keeps getting something new.

See also: [Verse](#verse), [Chorus](#chorus), [Bar (measure)](#bar)

## Audio-to-audio (cover)

**Audio-to-audio generation uses an existing recording as the input, not only text. The model keeps part of that audio, such as the melody or timing, and generates the rest.**

Text-only generation starts from nothing, so it cannot match a vocal or a beat you already have. Audio-conditioned models can. Common tasks are a cover (the same song in a new style), completing an arrangement around one part, and adding a new part that follows the input. On Lyro these tasks run on ACE-Step 1.5: upload a vocal and the model composes instruments around it, or upload a beat and it improvises a vocal line over it. The result follows the timing and harmony of your file, so a clean upload with a steady tempo works best. You need the rights to whatever you upload.

See also: [Text-to-music](#text-to-music), [Model routing](#model-routing), [A cappella](#a-cappella)

## Bar (measure)

**A bar, or measure, is a small group of beats that repeats through a song. In 4/4 time, the most common time signature, one bar has four beats.**

The time signature says how many beats are in a bar and which note value counts as one beat: 4/4 means four quarter-note beats. Music is built in blocks of bars. Phrases usually last 4 or 8 bars, and sections such as a verse or chorus last 8 or 16. At 120 BPM one 4/4 bar lasts exactly two seconds, so a 16-bar section is 32 seconds. Editing on bar lines keeps cuts and loops in time, which is why Lyro’s mixer snaps clips to the beat grid. The 16-step row on a drum machine is one bar split into sixteenth notes.

See also: [BPM (beats per minute)](#bpm), [Beat grid](#beat-grid), [Four-on-the-floor](#four-on-the-floor)

## Beat grid

**A beat grid is a map of where each beat and bar falls in a recording, worked out from its tempo and the position of the first downbeat.**

With a correct grid, clips snap to musical positions, loops repeat cleanly, and tempo-synced effects such as delay and pump land on the beat. The downbeat is beat one of a bar. If the tempo is right but the downbeat is wrong, everything sits a fraction early or late, and two parts played together sound like a stumble. Lyro detects BPM in your browser when you upload audio and lines parts up on the grid. You can correct the tempo if you know the real value. Music recorded without a click drifts in tempo, so a fixed grid will only fit it approximately.

See also: [BPM (beats per minute)](#bpm), [Bar (measure)](#bar), [Pump (tempo-synced ducking)](#pump)

## Blind listening test

**A blind listening test compares tracks without telling listeners which model, tool or person made each one, so brand names and expectations cannot influence the scores.**

Lyro’s model routing follows a blind test. We have run one so far, on 21 September 2026, and it was small and internal: one rater, seven Turkish-language briefs, three models, names hidden, each track scored from 1 to 5. Lyria 3 Pro scored 4.67 for pronunciation and 4.57 overall. ACE-Step 1.5 scored 2.00 and 3.29 but won the vocal-chop brief. MiniMax Music 2.6 scored 2.17 and 2.86 and was retired. The method is in [How we test AI music models](https://lyromusic.com/learn/how-we-test-ai-music-models). For context, a survey of 9,000 people found that 97% could not tell fully AI-generated music from human-made music ([Deezer and Ipsos, 12 Nov 2025](https://newsroom-deezer.com/2025/11/deezer-ipsos-survey-ai-music/)).

See also: [Model routing](#model-routing), [AI disclosure](#ai-disclosure)

## BPM (beats per minute)

**BPM, beats per minute, is the tempo of a track: how many beats pass in one minute. At 120 BPM there are two beats every second.**

Tempo decides how a track feels and what it can be mixed with. Rough ranges: lo-fi and R&B 70 to 95, reggaeton and afrobeats 90 to 110, amapiano about 110 to 115, house 118 to 128, techno 125 to 140, trap about 140 with a half-time feel, drum and bass 170 to 175. AI models follow a tempo more reliably when the prompt gives a number (“122 BPM”) than a word (“mid-tempo”). If you bring your own vocal or beat, the generated part has to match its BPM, so Lyro detects it in your browser first. You can check any file with the free [BPM and key finder](https://lyromusic.com/tools/bpm-key-finder).

See also: [Key](#key), [Beat grid](#beat-grid), [Bar (measure)](#bar)

## Chord progression

**A chord progression is the sequence of chords a song moves through, usually a loop of two to four chords that repeats under the melody.**

A chord is three or more notes played together, and the order of the chords sets the emotional colour of a track. Progressions are often written as Roman numerals counted from the key note, so the same pattern can be played in any key: I, V, vi, IV is a familiar pop loop, and i, VI, III, VII is its minor-key cousin. Deep house leans on seventh and ninth chords for a jazzy feel. In a prompt, describe harmony in words a model understands: “emotional minor-key progression”, “lush electric piano chords with jazzy seventh voicings”. Any part you add later has to fit the same chords and key.

See also: [Key](#key), [Topline](#topline), [Hook](#hook)

## Chorus

**The chorus is the section of a song that repeats with the same words and melody each time. It usually carries the title and the main hook.**

Most listeners remember the chorus first, so it gets the biggest arrangement: more layers, higher notes, stacked harmonies. It normally lasts 8 or 16 bars and returns two or three times. When you write lyrics for an AI model, mark it with a [Chorus] tag and keep the words identical on each repeat, because models tend to sing a repeated block more consistently. Keep the lines short and easy to sing. In club genres the same job is often done by a drop or a vocal hook instead of a full sung chorus. The studio effect called chorus, which thickens a sound, is unrelated.

See also: [Hook](#hook), [Verse](#verse), [Arrangement](#arrangement)

## Compression

**Compression reduces the gap between the loudest and quietest moments of a sound by turning the level down whenever it passes a set threshold.**

The main controls are threshold (the level where it starts working), ratio (how hard it turns down), attack and release (how fast it reacts and lets go) and make-up gain. A vocal usually needs it most: a ratio around 3:1 or 4:1 with a few decibels of gain reduction keeps every word audible over a beat. Too much makes a track flat and tiring. This is dynamics compression, which has nothing to do with file compression such as MP3. Every track in Lyro’s mixer has a compressor. Generated audio is usually compressed already, so go lightly there and use more on raw recordings.

See also: [Limiter](#limiter), [Sidechain](#sidechain), [Mixing](#mixing)

## Delay

**Delay is an effect that repeats a sound after a set time, producing one or more echoes that fade away.**

Two settings matter most. Time sets the gap between repeats, and in music it is usually synced to the tempo as a note value: a quarter note, an eighth, or a dotted eighth for a bouncing feel. Feedback sets how many repeats you hear. At 120 BPM a quarter-note delay is 500 milliseconds. Delay adds width and movement without the wash of reverb, which keeps a vocal clear. A common trick is the delay throw: send only the last word of a line to the delay so it echoes into the gap. Lyro’s mixer includes delay alongside reverb on every track.

See also: [Reverb](#reverb), [BPM (beats per minute)](#bpm), [Mixing](#mixing)

## EQ (equalisation)

**EQ, short for equalisation, turns chosen frequency ranges of a sound up or down, from deep bass to bright treble, to change its tone or make room for other parts.**

Human hearing runs from about 20 Hz to 20 kHz. Rough landmarks: sub bass below 60 Hz, bass and kick weight from 60 to 250 Hz, muddiness around 200 to 500 Hz, vocal presence from 2 to 5 kHz, air above 10 kHz. The most useful single move is a high-pass filter, which removes low rumble: 80 to 100 Hz on a vocal is a safe start. Cutting usually sounds more natural than boosting. When a vocal and a beat fight, dip the beat slightly around 2 to 4 kHz before you turn the vocal up. Every track in Lyro’s mixer has its own EQ.

See also: [Mixing](#mixing), [Compression](#compression), [Stems](#stems)

## Four-on-the-floor

**Four-on-the-floor is a drum pattern in 4/4 time where the kick drum hits on every beat: one, two, three, four. It drives house, techno and disco.**

The steady kick gives dancers a pulse that never breaks, and everything else plays against it: claps or snares on beats two and four, open hi-hats on the off-beats in between. On a 16-step sequencer the kick sits on steps 1, 5, 9 and 13. The pattern pairs naturally with sidechain pumping, where pads and bass dip on each kick. In a prompt, write the phrase itself (“four-on-the-floor kick”), since models recognise it. Afro house keeps the four kicks and layers syncopated percussion on top. Amapiano, trap and reggaeton are built on other kick patterns, so leave the phrase out there.

See also: [Bar (measure)](#bar), [Sidechain](#sidechain), [Pump (tempo-synced ducking)](#pump)

## Hook

**A hook is the short, catchy element of a song that sticks in the memory: a sung phrase, a riff, a chopped vocal or even a drum sound.**

A chorus can contain the hook, but they are different things: the chorus is a section, the hook is the idea people hum afterwards. Short videos use only a few seconds of a track, so the hook often decides whether a song gets replayed. Practical rules: keep it to a few words or notes, repeat it, and let it arrive early. With AI models, put the hook line at the top of the chorus and repeat it exactly. Lyro’s lyric writer drafts hooks, verses and vocal-chop phrases in 12 languages, so you can try several hook lines before you spend credits on a generation.

See also: [Chorus](#chorus), [Topline](#topline), [Vocal chop](#vocal-chop)

## Inpainting (repaint)

**Inpainting, also called repaint, regenerates one chosen time range of a track while keeping everything before and after it, so a weak section can be fixed without starting again.**

The term comes from image editing, where a model fills a masked area so that it blends with its surroundings. In music the mask is a time range, for example 0:45 to 1:05. The model hears the audio on both sides and generates a replacement that joins them, optionally with new lyrics or a new style prompt for that part. It suits small repairs: a mispronounced word, a clumsy transition, a chorus that needs more lift. It suits large rewrites less, because the new part has to stay compatible with what is kept. Results vary between attempts, so keep the original and compare the two by ear.

See also: [Audio-to-audio (cover)](#audio-to-audio), [Seed](#seed), [Text-to-music](#text-to-music)

## Key

**A song’s key is its home note plus the scale built on it, such as A minor or C major. It sets which notes and chords fit together.**

A scale is a set of notes in a fixed pattern of steps. Major keys tend to sound bright and minor keys darker, which is why most house, trap and drill sits in minor. Key matters most when you combine audio: a vocal in A minor over a beat in B flat minor sounds out of tune however good each part is. Lyro detects the key of an uploaded vocal or beat in your browser and generates the missing part in the same key. Detection can confuse a key with its relative (C major and A minor share the same notes), so correct it if you know better.

See also: [BPM (beats per minute)](#bpm), [Chord progression](#chord-progression), [Topline](#topline)

## Limiter

**A limiter is a very fast, high-ratio compressor that stops audio from going above a set ceiling. It is the last processor in a mastering chain.**

Turning a finished mix up would push its loudest peaks past 0 dBFS, the digital maximum, and clip them. A limiter catches those peaks, so the average level can rise while the ceiling holds. The two settings are the ceiling, commonly about -1 dB true peak, and the amount of gain pushed into it. A few decibels of limiting is normal. Much more and the drums lose punch, the mix starts to distort and it becomes tiring to hear. Lyro’s mastering raises your track to a chosen LUFS target with a limiter, and the A/B switch lets you compare the mastered and unmastered sound.

See also: [LUFS](#lufs), [True peak](#true-peak), [Mastering](#mastering)

## Log drum

**In amapiano and afro house, the log drum is a pitched, percussive synth bass with a woody knock. It plays the bassline and much of the rhythm at once.**

The name comes from the wooden slit drum, but the sound in these genres is synthesised: a short, hollow attack followed by a deep tuned tail, often sliding between notes. In amapiano, at roughly 110 to 115 BPM, it is the lead voice of the track, playing sparse syncopated phrases with plenty of space. Afro house, at around 120 to 125 BPM, uses a rolling log drum bassline under a four-on-the-floor kick. Models recognise the term, so ask for it directly: “wide log drum bass slides, shuffling shakers, airy piano chords”. Because it is pitched, it has to be in the key of the song.

See also: [808](#808), [Four-on-the-floor](#four-on-the-floor), [Key](#key)

## LUFS

**LUFS (loudness units relative to full scale) measures how loud audio sounds over time, not how high its peaks reach. Streaming services use it to even out playback volume.**

The measurement is defined in the ITU-R BS.1770 standard and weights frequencies roughly the way hearing does. The figure that matters for a release is integrated LUFS, the average over the whole track. Values are negative, and closer to zero means louder: -8 LUFS is much louder than -14. Many streaming services turn loud tracks down to a reference of around -14 LUFS, so mastering far above that gains nothing there and costs punch. Club and DJ masters are usually louder. One loudness unit equals one decibel. Lyro’s mastering lets you pick a loudness target in LUFS, applies a limiter and gives you an A/B switch.

See also: [Mastering](#mastering), [Limiter](#limiter), [True peak](#true-peak)

## Mastering

**Mastering is the final step before release: the finished stereo mix is brought to its target loudness, given a safe peak ceiling and exported in the delivery format.**

Mixing balances the parts inside a song. Mastering treats the whole song as one file, so it cannot fix a vocal that is too quiet: go back to the mix for that. A basic chain is gentle EQ, light compression, then a limiter that lifts the track to a loudness target measured in LUFS while holding true peaks under a ceiling such as -1 dB. Lyro’s mastering runs in your browser, so it uses no credits. You pick a target for streaming or for the club, check the difference with the A/B switch and export WAV on the Creator plan. More on the [mixing and mastering](https://lyromusic.com/features/mixing-mastering) page.

See also: [LUFS](#lufs), [Limiter](#limiter), [Mixing](#mixing)

## Mixing

**Mixing is the process of combining separate tracks, such as vocal, drums and bass, into one balanced stereo file by setting levels, panning, tone, dynamics and effects.**

Work in a fixed order and it gets easier. Set levels first, because balance is most of a mix. Then pan parts left and right for width, keeping kick, bass and lead vocal in the centre. Use EQ to stop parts masking each other, compression to steady uneven ones, and reverb and delay to place them in a space. Leave a few decibels of headroom, meaning the loudest peak stays below 0 dBFS, so mastering has room to work. Lyro’s multitrack mixer runs in the browser and uses no credits. It is part of the Creator plan. Phones have less memory than computers, so keep large projects for a laptop or desktop.

See also: [EQ (equalisation)](#eq), [Compression](#compression), [Mastering](#mastering)

## Model routing

**Model routing means sending each generation request to the AI model best suited to it, based on task, language and test results, instead of using one model for everything.**

Music models have different strengths. One may pronounce sung lyrics clearly, while another holds an exact tempo or accepts your audio as input. Lyro is multi-model, and its routing follows the one blind listening test it has run: one rater, seven Turkish-language briefs, three models, in September 2026. That test did not cover each genre or each language. At the moment Google’s Lyria 3 Pro is the default for sung lyrics, and ACE-Step 1.5, an open model released under the Apache-2.0 licence, handles vocal chops, tempo-locked beats and the audio-conditioned tasks. Both are reached through fal.ai. Each track in your library shows which model made it. Routing changes when test results change: MiniMax Music 2.6 was retired after that test. See [how we test](https://lyromusic.com/learn/how-we-test-ai-music-models).

See also: [Blind listening test](#blind-test), [Text-to-music](#text-to-music), [Audio-to-audio (cover)](#audio-to-audio)

## Prompt

**A prompt is the text description you give an AI music model. It tells the model the genre, tempo, instruments, mood and type of vocal you want to hear.**

Models respond best to concrete production language in a sensible order: genre, BPM, drums, bass, chords and lead instruments, vocal type and language, then mood and mix character. “Afro house, 122 BPM, rolling log drum bassline, congas and shakers, warm marimba motif, soft female vocal, spacious club mix” will beat “a nice summer song”. Around 40 to 60 words is enough, and very long prompts tend to be partly ignored. Describe the sound instead of naming a real artist or song. Lyrics go in their own field, not in the prompt. The free [AI music prompt generator](https://lyromusic.com/ai-music-prompt-generator) builds a prompt like this from a few choices.

See also: [Text-to-music](#text-to-music), [Seed](#seed), [BPM (beats per minute)](#bpm)

## Pump (tempo-synced ducking)

**Pump is a rhythmic dip in volume, timed to the beat, that makes pads, bass or a whole mix seem to breathe. It is a signature of house music.**

The classic method is sidechain compression: a compressor on the pads listens to the kick and turns them down on every hit. A tempo-synced ducker reaches the same result without needing the kick signal. It draws a volume curve on the beat grid, usually once per quarter note: down fast on the beat, back up before the next one. Lyro’s mixer has this as the pump control. It needs the correct BPM and downbeat to land properly, so check the detected tempo first. Use more on pads and bass, and little or none on the lead vocal, where heavy pumping makes words hard to follow.

See also: [Sidechain](#sidechain), [Four-on-the-floor](#four-on-the-floor), [Beat grid](#beat-grid)

## Reverb

**Reverb is an effect that simulates the reflections of a physical space, from a small room to a cathedral, so a dry sound seems to sit in a real place.**

The key settings are decay time (how long the tail lasts), pre-delay (a short gap before the reverb starts, which keeps a vocal clear) and the balance between dry and wet signal. Short decays under a second suit fast, busy tracks. Two seconds or more suits ballads and ambient music. Too much reverb pushes a sound to the back and blurs the mix, and low frequencies suffer first, so keep kick and bass mostly dry. Record vocals dry and add reverb afterwards, because reverb that is already in a recording cannot be removed cleanly. Lyro’s mixer includes reverb and delay for every track.

See also: [Delay](#delay), [Mixing](#mixing), [A cappella](#a-cappella)

## Sample rate and bit depth

**Sample rate is how many times per second digital audio measures the sound wave, in hertz. Bit depth is how precisely each of those measurements is stored.**

A sample rate can capture frequencies up to half its value, so 44.1 kHz covers the range of human hearing, which ends at about 20 kHz. It is the standard for music releases, while 48 kHz is the norm for video. Bit depth sets the dynamic range: 16-bit gives about 96 dB, enough for a finished master, and 24-bit gives about 144 dB, which leaves more margin while recording and mixing. Higher numbers make bigger files, not automatically better sound. Converting upwards adds nothing: an MP3 saved as a 24-bit WAV still contains only what the MP3 kept. Check your distributor’s specification before you export.

See also: [WAV vs MP3 (and FLAC)](#wav-vs-mp3), [Mastering](#mastering)

## Seed

**A seed is the number that starts a generative model’s random process. Reusing a seed with the same prompt, settings and model version gives the same or a near-identical result.**

AI music generation starts from random noise, which is why the same prompt gives a different song each time. Fixing the seed removes that randomness, so you can change one thing, for instance a single word in the prompt, and hear what it alone does. Leaving the seed random is how you get variations. Two caveats: not every model or service exposes a seed, and a seed only reproduces a result on the same model version, because an update changes what the number leads to. For most people the practical approach is simpler: generate a few versions, keep the best and refine the prompt from there.

See also: [Prompt](#prompt), [Text-to-music](#text-to-music), [Inpainting (repaint)](#inpainting)

## Sidechain

**Sidechain means controlling an effect on one track with the signal from another. The classic use is a compressor on the bass that ducks whenever the kick hits.**

Kick and bass occupy the same low frequencies, and when both sound at once the low end turns muddy and uses up headroom. Sidechain compression makes room: the kick plays, the bass dips for a fraction of a second, then returns. Set deep with a slow release, the dip becomes the audible pumping of house music. Set shallow and fast, it is inaudible and simply cleans the low end. The same idea is used in podcasts and adverts to lower music under a voice. Lyro’s mixer offers a tempo-synced version called pump, which follows the beat grid and does not need a kick track as its trigger.

See also: [Pump (tempo-synced ducking)](#pump), [Compression](#compression), [Four-on-the-floor](#four-on-the-floor)

## Stem separation

**Stem separation uses a neural network to split a finished, mixed recording back into parts such as vocals, drums, bass and everything else.**

The model has learned what each source sounds like and estimates them from the mix. It is an estimate, not a recovery of the original recordings, so expect artefacts: faint cymbals or reverb in the vocal, watery high frequencies, bass notes leaking into the “other” stem. Clean, sparse mixes separate best. Dense, distorted or heavily reverberant ones separate worst. Lyro’s [stem splitter](https://lyromusic.com/features/stem-splitter) returns four stems: vocals, drums, bass and other. It is part of the Creator plan and costs one credit plus two per started minute, so 7 credits for a three-minute track. Only upload recordings you have the right to use: separating a track does not give you any rights in it.

See also: [Stems](#stems), [A cappella](#a-cappella), [Vocal chop](#vocal-chop)

## Stems

**Stems are a song delivered as a few separate audio files, typically vocals, drums, bass and other instruments, that play back together as the full mix.**

Stems sit between a finished stereo file and the full multitrack session, which has one file for every recorded part. They let you rebalance a song, mute the vocal for an instrumental, make a remix, or send parts to a mix engineer. All stems must start at the same point and have the same length, so they line up when imported into a DAW (digital audio workstation, the software used to record and mix). On Lyro you can split a track into four stems. When Lyro sings your lyrics over your beat, the new vocal arrives as its own isolated stem, so your beat stays untouched.

See also: [Stem separation](#stem-separation), [Mixing](#mixing), [WAV vs MP3 (and FLAC)](#wav-vs-mp3)

## Swing

**Swing delays every second note in a pair of eighth or sixteenth notes, turning a rigid, even rhythm into a loping, shuffled one.**

Drum machines express swing as a percentage. At 50% the notes are evenly spaced, which is called straight. At about 66% the second note lands on a triplet, the full shuffle of jazz and blues. Most electronic music lives in between: 54 to 60% on the sixteenth-note hi-hats gives deep house and UK garage their bounce, and boom bap hip-hop leans on swung drums too. Swing is a large part of why a programmed beat feels human instead of mechanical. In a prompt, use the words models know: “shuffled hi-hats with swing”, “laid-back swung drums”, or “straight, tight, quantised” when you want the opposite.

See also: [Beat grid](#beat-grid), [BPM (beats per minute)](#bpm), [Four-on-the-floor](#four-on-the-floor)

## Text-to-music

**Text-to-music is a type of generative AI that produces a finished audio recording, instruments and optionally sung vocals, from a written description and lyrics.**

You write a prompt describing genre, tempo, instruments and mood, optionally add lyrics, and the model returns audio. The output is a single mixed file, not separate tracks, which is why stem separation matters afterwards. Quality differs between models and between languages: a model can produce a convincing beat yet mispronounce sung words. For that reason Lyro uses more than one model and picks by the job, with Google’s Lyria 3 Pro as the current default for sung lyrics. A Lyria 3 Pro song costs 4 credits, whatever its length. The quickest way to try it is to [describe a song](https://lyromusic.com/start), or read about the [AI song generator](https://lyromusic.com/features/ai-song-generator).

See also: [Prompt](#prompt), [Model routing](#model-routing), [Audio-to-audio (cover)](#audio-to-audio)

## Topline

**A topline is the vocal melody and lyrics written over an existing instrumental. Writing one is called toplining, and the person who does it is a topliner.**

In modern pop and dance music the beat often comes first and the topline is written to it, sometimes by a different person. A strong topline fits the key and chords of the beat, leaves gaps for the instrumental to answer, and puts its most memorable phrase, the hook, where the arrangement opens up. On Lyro, if you upload a beat you have two options. Give it lyrics and it sings them at the beat’s tempo and key, returning an isolated vocal stem. Or give no lyrics and it improvises a topline over the beat. The improvised version follows the music but does not sing words you have written.

See also: [Hook](#hook), [A cappella](#a-cappella), [Key](#key)

## True peak

**True peak is the highest level an audio signal reaches once it is converted back to analogue, including peaks that fall between digital samples. It is measured in dBTP.**

A normal peak meter reads only the stored sample values. The real waveform curves between those samples and can rise higher, so a file that reads 0 dBFS may clip inside a player’s converter. True-peak meters estimate those inter-sample peaks by oversampling, as set out in the ITU-R BS.1770 standard. Lossy encoding to MP3 or AAC shifts peaks as well, often upwards. A ceiling of about -1 dBTP is a common safety margin for streaming, and some engineers leave -2 dBTP on very loud masters. In Lyro’s mastering, the limiter holds the ceiling while the track is raised to its LUFS target.

See also: [Limiter](#limiter), [LUFS](#lufs), [Mastering](#mastering)

## Verse

**A verse is the section of a song that tells the story. Its melody usually repeats each time, while the words change from one verse to the next.**

Verses set up the chorus. They are normally lower in energy, with fewer instruments and a narrower melody, so the chorus feels like a lift when it arrives. A common length is 8 or 16 bars. When writing lyrics for an AI model, mark each one with a [Verse] tag, keep line lengths similar so the phrasing stays even, and write about four to eight lines per verse. Lines that are too long get rushed or cut. Rap verses are the exception: they are denser, often a full 16 bars, and carry most of the song. Lyro’s lyric writer drafts verses in 12 languages.

See also: [Chorus](#chorus), [Hook](#hook), [Arrangement](#arrangement)

## Vocal chop

**A vocal chop is a short slice of a vocal recording, often a single syllable, that is re-pitched and re-sequenced so the voice plays like an instrument.**

Chops carry rhythm and melody without carrying meaning, so the words do not need to be understood. That makes them a staple hook in house, afro house and UK garage, usually drenched in reverb and delay. It also makes them a different job for an AI model than a sung verse. In Lyro’s small internal blind test, ACE-Step 1.5 won the vocal-chop brief even though it scored well below Lyria 3 Pro on pronunciation, so Lyro routes chop requests to it. ACE-Step is priced by length: one credit per 30 seconds with the detailed pass, or one credit per minute with the fast pass. To make chops by hand, split a song into stems, slice the vocal stem and re-pitch the pieces with the vocal tools.

See also: [Hook](#hook), [Model routing](#model-routing), [Stem separation](#stem-separation)

## Watermark (SynthID)

**An audio watermark is an inaudible signal embedded in the sound itself that software can detect later. SynthID is Google’s watermark for marking content generated by its AI models.**

A watermark differs from a metadata tag. A tag, such as an ID3 field, sits beside the audio inside the file and disappears if someone rewrites the tags. A watermark is part of the audio signal, so it travels with the sound. Tracks made with Google’s Lyria 3 Pro on Lyro carry SynthID. It is added by the model, not by Lyro. Separately, Lyro writes an AI-generated declaration into the ID3 tag of every download of generated audio, whichever model made it. Both exist for the same reason: so that platforms and listeners can find out how a track was made. See [AI disclosure](https://lyromusic.com/glossary#ai-disclosure).

See also: [AI disclosure](#ai-disclosure), [Model routing](#model-routing)

## WAV vs MP3 (and FLAC)

**WAV stores audio uncompressed at full quality. MP3 shrinks the file by permanently discarding detail. FLAC shrinks it too, by about half, without losing anything.**

Use a lossless format, WAV or FLAC, for anything you will process again: mixing, mastering, stem work, or delivery to a distributor. Use MP3 for sharing, previews and messaging, where a much smaller file matters more than the last bit of detail. Every conversion to a lossy format loses a little more, so keep a lossless master and make MP3 copies from it. Converting an MP3 to WAV does not restore what was discarded. On Lyro, MP3 and FLAC downloads are included on every plan, WAV export comes with the Creator plan, and downloads are never counted or capped. Generated files carry an ID3 tag declaring them AI-generated.

See also: [Sample rate and bit depth](#sample-rate), [Mastering](#mastering), [Stems](#stems)
