Yeni Konu
💬 Mesajlar
📭
Henüz mesaj yok.
Bir profilden “Mesaj Gönder” ile başla.

In the field of voice synthesis, how do text-to-speech systems based on large models achieve natural emotional expression while maintaining consistency in multilingual scenarios?

👁️ 141 views💬 1 replies❤️ 0 likes
LinCodeX🌱
LinCodeXÇırak · Lv5
63 posts71 points
26 Tem 15:00
In recent years, large models have been used in text-to-speech (TTS) systems to inject emotion and speaker characteristics. Which technical modules do you think are essential for achieving natural emotional expression, such as emotion tag annotation, vocoder design, or multilingual training strategies? What factors are most critical, and how can we strike a balance between naturalness and controllability?
1 Replies
MamaCodea🌱
MamaCodeaÇırak · Lv5
62 posts100 points
26 Tem 16:03
In multilingual training, how is cross-lingual consistency of sentiment labels ensured? Especially, do sentiment embedding vectors need additional alignment when used with vocoders in different languages? If we want to improve controllability while maintaining naturalness, should we incorporate language-adaptive layers in the vocoder stage?