Papers
arxiv:2501.17790

BreezyVoice: Adapting TTS for Taiwanese Mandarin with Enhanced Polyphone Disambiguation -- Challenges and Insights

Published on Jan 29, 2025
Authors:
,
,
,
,
,
,
,
,

Abstract

BreezyVoice, a specialized Text-to-Speech system for Taiwanese Mandarin, employs advanced models for phonetic control and polyphone disambiguation, demonstrating superior performance and generalizability.

We present BreezyVoice, a Text-to-Speech (TTS) system specifically adapted for Taiwanese Mandarin, highlighting phonetic control abilities to address the unique challenges of polyphone disambiguation in the language. Building upon CosyVoice, we incorporate a S^{3} tokenizer, a large language model (LLM), an optimal-transport conditional flow matching model (OT-CFM), and a grapheme to phoneme prediction model, to generate realistic speech that closely mimics human utterances. Our evaluation demonstrates BreezyVoice's superior performance in both general and code-switching contexts, highlighting its robustness and effectiveness in generating high-fidelity speech. Additionally, we address the challenges of generalizability in modeling long-tail speakers and polyphone disambiguation. Our approach significantly enhances performance and offers valuable insights into the workings of neural codec TTS systems.

Community

•
This comment has been hidden
•
This comment has been hidden (marked as Off-Topic)

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2501.17790
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 1

Datasets citing this paper 1

Spaces citing this paper 24

Browse 24 spaces citing this paper

Collections including this paper 1