Benchmark TTSFR

Test d'écoute — agrégation

3 auditeur(s) — 40 votes A/B, 60 notes MOS, 40 votes émotion.

A/B — win-rate par modèle

modèlewin-rate(victoires / duels)défauts entendus
firered_tts379%15.0 / 19accent:2
voxcpm271%10.0 / 14
cosyvoice3_05b50%6.0 / 12voix_diff:6 accent:2 tronque:1 repetition:1 artefact:1
chatterbox_v333%4.0 / 12voix_diff:5 accent:2 artefact:2 repetition:1
moss_tts_local_v1525%2.5 / 10voix_diff:1 artefact:1
xtts_v219%2.5 / 13voix_diff:7 artefact:4 accent:1

A/B — matrice des duels (gauche bat haut)

chatterbox_v3cosyvoice3_05bfirered_tts3moss_tts_local_v15voxcpm2xtts_v2
chatterbox_v3·0/10/31/20/11/1
cosyvoice3_05b1/1·1/31/10/21/1
firered_tts33/32/3·3/32/34/5
moss_tts_local_v151/20/10/3·1/3
voxcpm21/12/21/32/3·3/3
xtts_v20/10/11/50/3·

A/B — win-rate par voix de référence (victoires / duels)

modèle?aurore2_narrationjohnnypapy_narrationtonton_marc_narration
chatterbox_v330% (1.5/5)0% (0.0/1)50% (0.5/1)50% (2.0/4)0% (0.0/1)
cosyvoice3_05b20% (1.0/5)67% (2.0/3)75% (1.5/2)75% (1.5/2)
firered_tts373% (8.0/11)100% (3.0/3)100% (2.0/2)67% (2.0/3)
moss_tts_local_v1525% (1.0/4)0% (0.0/2)0% (0.0/1)100% (1.0/1)25% (0.5/2)
voxcpm278% (7.0/9)100% (1.0/1)0% (0.0/1)67% (2.0/3)
xtts_v225% (1.5/6)0% (0.0/2)50% (0.5/1)17% (0.5/3)0% (0.0/1)

MOS — moyenne ± IC 95 % par axe

modèleNaturelIntelligibiliteSimilariteExpressivite
chatterbox_v33.20 ±0.573.30 ±0.592.00 ±0.582.90 ±0.54
cosyvoice3_05b2.89 ±0.693.22 ±0.442.89 ±0.692.22 ±0.54
firered_tts34.12 ±0.584.75 ±0.324.38 ±0.363.38 ±0.73
moss_tts_local_v153.75 ±0.604.00 ±0.594.25 ±0.432.83 ±0.63
voxcpm23.43 ±0.403.57 ±0.403.43 ±0.582.86 ±1.00
xtts_v23.36 ±0.793.64 ±0.671.86 ±0.542.71 ±0.52

MOS — note globale moyenne (4 axes) par voix de référence

modèle?aurore2_narrationjohnnypapa_narrationpapy_narrationtonton_marc_narration
chatterbox_v32.83 (n=3)3.12 (n=2)2.25 (n=1)3.50 (n=1)2.75 (n=2)2.50 (n=1)
cosyvoice3_05b2.42 (n=3)3.25 (n=2)3.50 (n=1)2.75 (n=2)2.50 (n=1)
firered_tts34.42 (n=3)4.25 (n=2)4.12 (n=2)3.25 (n=1)
moss_tts_local_v153.70 (n=5)3.25 (n=2)3.50 (n=1)4.75 (n=1)3.75 (n=3)
voxcpm23.69 (n=4)3.00 (n=1)2.75 (n=2)
xtts_v22.75 (n=2)3.25 (n=3)2.50 (n=5)3.25 (n=2)3.12 (n=2)

Corrélation MOS (humain) ↔ métrique automatique (par clip)

axe humainmétrique autoPearson rn
intelligibilite1-WER0.1260
similariteSIM0.2260

Émotion — quel modèle rend le mieux l'émotion ?

A/B en aveugle : deux modèles disent la même phrase, tous deux clonés depuis la même voix de réf émotionnelle. Win-rate = victoires + ½ nuls.

modèlewin-rate(victoires / duels)
firered_tts388%7.0 / 8
xtts_v262%2.5 / 4
chatterbox_v350%3.0 / 6
voxcpm244%4.0 / 9
cosyvoice3_05b42%2.5 / 6
moss_tts_local_v1514%1.0 / 7

Émotion — win-rate par (modèle, émotion)

modèlecolerejoiepeurtristesse
chatterbox_v367% (2.0/3)0% (0.0/1)0% (0.0/1)100% (1.0/1)
cosyvoice3_05b25% (0.5/2)0% (0.0/1)0% (0.0/1)100% (2.0/2)
firered_tts380% (4.0/5)100% (1.0/1)100% (2.0/2)
moss_tts_local_v150% (0.0/4)100% (1.0/1)0% (0.0/2)
voxcpm250% (2.0/4)100% (1.0/1)50% (1.0/2)0% (0.0/2)
xtts_v275% (1.5/2)100% (1.0/1)0% (0.0/1)

Émotion — protocole archivé (réf. émotionnelle vs neutre)

Sessions antérieures : « lequel sonne le plus <émotion> », clip cloné depuis la réf émotionnelle vs depuis la réf neutre.

modèlenchoix réf. émo.égalitéchoix réf. neutretaux de transfert
moss_tts_local_v154400100%
firered_tts32200100%
voxcpm24400100%
cosyvoice3_05b1100100%
xtts_v24400100%
chatterbox_v3541080%