# How we test AI music models: our blind listening method | Lyro

> How Lyro chose its models: one blind listening test. The method, the seven Turkish briefs, the scores for three models, and the limits of a small test.

Source: https://lyromusic.com/learn/how-we-test-ai-music-models

# How we test AI music models: our blind listening method

Updated 21 September 2026, 6 min read, by the Lyro team

We generate the same briefs on every model, hide the model names, shuffle the order and score each clip from 1 to 5 for pronunciation, production and overall. In our first and so far only test, on 21 Sep 2026, we ran seven Turkish-language briefs on three models. Google Lyria 3 Pro scored 4.67 for pronunciation and 4.57 overall. ACE-Step 1.5 scored 2.00 and 3.29. MiniMax Music 2.6 scored 2.17 and 2.86. It was a small test with one rater.

## Why the test is blind

We hide the model names because knowing them changes what you hear. A well-known name or a higher price pushes a score before the first note. In a [blind test](https://lyromusic.com/glossary#blind-test) the only input is the sound.

Labels matter most when ears cannot settle a question. In a Deezer and Ipsos survey of 9,000 people, 97% could not tell fully AI-generated music from human-made music ([Deezer and Ipsos, 12 Nov 2025](https://newsroom-deezer.com/2025/11/deezer-ipsos-survey-ai-music/)). So our scoring page shows the clips as A, B and C, in an order that is shuffled per brief. Model averages are revealed only after every clip has a score.

## The seven briefs

The seven briefs run from easy to hard. Six have Turkish lyrics and one has no vocal. A brief is the full request a model receives: a genre description, a voice and the lyrics. Every model got the same brief and produced one clip of about 90 seconds.

| Brief | What it measures |
|---|---|
| 1. Afro house vocal chops | The easiest case: short, repeated phrases whose words do not have to be understood |
| 2. Deep house soul hook | Short lines that must be sung clearly |
| 3. Afro house verse and hook | Whether words sit on the rhythm and the verse leads into the hook |
| 4. Melodic house with a male vocal | Turkish on a male voice |
| 5. Long Turkish verse | The hard case: eight lines packed with Turkish-specific sounds |
| 6. Afro house instrumental | A control with no vocal, to judge production alone |
| 7. Turkish pop | A control outside house music |

The lyrics include the sounds we expected models to struggle with: the soft g (ğ) in *yağmur* and *düğün*, the dotless i (ı) in *ışık* and *kırık*, ö and ü in *gözlerin* and *gülüm*, and ş and ç in *şehir* and *çiçek*.

## The three scores

Each clip gets three scores from 1 to 5. Each score answers one plain question.

- **Pronunciation.** Can a Turkish speaker understand the words, and are they the right words?

- **Production.** Do the mix and the sound fit the genre, and could you play it in a club?

- **Overall.** Would you pay for this result?

Pronunciation is averaged over the six briefs that have a vocal, because the instrumental has no words to score. Production and overall are averaged over all seven.

## Results of the 21 Sep 2026 test

Lyria 3 Pro led on all three scores. ACE-Step 1.5 was weak on Turkish pronunciation but won the vocal-chop brief. MiniMax Music 2.6 came last for production and overall.

| Model | Pronunciation | Production | Overall | What we did with it |
|---|---|---|---|---|
| Google Lyria 3 Pro | 4.67 | 4.43 | 4.57 | Default for sung lyrics |
| ACE-Step 1.5 | 2.00 | 3.00 | 3.29 | Used for vocal chops and tempo-locked beats |
| MiniMax Music 2.6 | 2.17 | 2.71 | 2.86 | Retired from Lyro |

On the vocal-chop brief ACE-Step scored 5 overall against 4 for Lyria 3 Pro, even though its pronunciation on that brief scored 1. When words are chopped into rhythm, clarity matters less than groove. Lyria 3 Pro had the highest overall score on the other six briefs, level with MiniMax on one of them.

## The limits of this test

This was a small internal test, and it should be read as one. We publish it because it drives our routing, not because it settles which model is best.

- **One rater.** One Turkish speaker on our team scored every clip.

- **One language.** All lyrics were Turkish. The scores say nothing about the other languages Lyro supports.

- **Small sample.** There were 21 clips: one generation per model per brief, from a single [seed](https://lyromusic.com/glossary#seed), the number that fixes a model’s random starting point. A second run could move any score.

- **Narrow genres.** Six of the seven briefs were house music.

- **Three models.** We tested models we can offer through fal.ai. Suno and Udio were not part of the test, so these scores say nothing about them.

- **One date.** The results describe the model versions available on 21 Sep 2026.

## How the results become routing rules

[Routing](https://lyromusic.com/glossary#model-routing) means choosing which model receives your request. Lyro’s rules follow one question: how well do the words need to be understood, and in which language?

- **Sung lyrics go to Lyria 3 Pro.** For Turkish lyrics, ACE-Step is switched off for sung vocals unless you confirm that you want it anyway. Lyro detects Turkish even when it is typed without Turkish letters.

- **Vocal chops go to ACE-Step 1.5.** It won that brief. It is priced by length (one credit per 30 seconds with the detailed pass, or one credit per minute with the fast pass), so a short hook costs less than the 4 credits of a Lyria 3 Pro song, and a full-length track with the detailed pass costs more.

- **Tempo-locked beats go to ACE-Step 1.5.** It accepts an exact tempo and key, which matters when a beat has to sit under your own vocal.

- **MiniMax Music 2.6 was retired.** It was last for production and overall, and did not win a brief outright. That describes this version on these briefs, not the model in general.

These rules come from one test in one language. We have not run a test for each genre or each language, so outside sung Turkish they are a working rule, not a measurement. Lyro recommends a model, and you can choose another.

## How the test continues in the app

Every generated track in the app has a thumbs up and a thumbs down, and those votes continue the test at a scale one rater cannot reach. We group them by model, genre and language. If real use disagrees with the blind test, we test again.

Votes are not blind, so we treat them as a signal to re-test, not as a final score. A new model goes through the same blind test before it receives requests. You can hear the current routing in the samples on the [homepage](https://lyromusic.com/).

## Questions

### Which AI music model is best for Turkish vocals?

In our small internal blind test on 21 Sep 2026, Google Lyria 3 Pro scored 4.67 out of 5 for Turkish pronunciation, ACE-Step 1.5 scored 2.00 and MiniMax Music 2.6 scored 2.17. It was a small test with one rater. More in [our Turkish guide](https://lyromusic.com/learn/best-ai-music-generator-for-turkish-songs).

### Why did a lower-scoring model win one brief?

The brief was vocal chops, where words are cut into short rhythmic phrases and do not have to be understood. ACE-Step 1.5 scored 5 overall there against 4 for Lyria 3 Pro, so Lyro routes vocal chops to it.

### Did you test Suno or Udio?

No. The test covered three models that Lyro can offer through fal.ai: Google Lyria 3 Pro, ACE-Step 1.5 and MiniMax Music 2.6. The scores say nothing about Suno or Udio. Our [comparison pages](https://lyromusic.com/compare/lyro-vs-suno) stick to dated, sourced facts.

### How many people scored the clips?

One. A Turkish speaker on our team scored all 21 clips without seeing the model names. That is a real limit, which is why votes in the app are used to check the results.

## Sources

- [Deezer and Ipsos, 12 Nov 2025](https://newsroom-deezer.com/2025/11/deezer-ipsos-survey-ai-music/)

## Hear it for yourself.

[Make a song](https://lyromusic.com/start)[Generate a prompt](https://lyromusic.com/ai-music-prompt-generator)
