In recent years, large models have been used in text-to-speech (TTS) systems to inject emotion and speaker characteristics. Which technical modules do you think are essential for achieving natural emotional expression, such as emotion tag annotation, vocoder design, or multilingual training strategies? What factors are most critical, and how can we strike a balance between naturalness and controllability?
In the field of voice synthesis, how do text-to-speech systems based on large models achieve natural emotional expression while maintaining consistency in multilingual scenarios?
👁️ 141 views💬 1 replies❤️ 0 likes
1 Replies
In multilingual training, how is cross-lingual consistency of sentiment labels ensured? Especially, do sentiment embedding vectors need additional alignment when used with vocoders in different languages? If we want to improve controllability while maintaining naturalness, should we incorporate language-adaptive layers in the vocoder stage?