ZICQ
中 Log in / Sign up
Newsroom LLMs & Foundation Models #Hugging Face #Speech Synthesis #Cross-Lingual #TTS #Accent Consistency

Hugging Face Proposes AAG: Enhancing Speaker Similarity and Accent Consistency in Cross-Lingual Voice Cloning

Avatar of Mr.Xu

By Mr.Xu

Published:

中文阅读 (Chinese) English Version

Summary:Hugging Face's research team introduces Accent Analogy Guidance (AAG), a novel technique addressing the issue of reference accent leakage in cross-lingual zero-shot text-to-speech synthesis. AAG subtracts an estimated accent direction from the model's predictions, preserving the accent while canceling the voice identity, thereby enhancing speaker similarity. Evaluations across multiple open TTS models demonstrate that AAG significantly improves speaker similarity while maintaining accent consist


Background and Challenge

In cross-lingual zero-shot text-to-speech (TTS) synthesis, the accent of the reference speech often leaks into the target speech, causing inconsistencies in accent and reducing the naturalness of the synthesized speech. This limitation hinders the effectiveness of voice cloning technologies in multilingual applications.

Technical Breakthrough: Accent Analogy Guidance (AAG)

Hugging Face's research team introduces Accent Analogy Guidance (AAG), a novel method designed to address the issue of accent leakage. The core idea of AAG involves the following steps:

  1. Estimate Accent Direction: Estimate the accent direction from the model's predictions for a synthetic voice rendered in both languages.
  2. Subtract Accent Direction: Subtract the estimated accent direction to preserve the accent while canceling the voice identity.
  3. Reweight Classifier-Free Guidance: Reweight classifier-free guidance between the reference and text, and use its variants to maintain the identity-accent trade-off curve.

Experiments and Results

The research team evaluated AAG on four open TTS models:

  • OmniVoice: On three test sets, the ΔSIM metric increased by 0.11 to 0.27, with the accent score on a 1-5 scale improving from 3.51 to 4.28, while the speaker similarity remained at 0.29.
  • MaskGCT and CosyVoice 2: Both models also lay above their respective trade-off curves, demonstrating AAG's general applicability.
  • F5-TTS: AAG outperformed any reweighting setting in terms of accent consistency.

Additionally, the study used an LLM-free language-ID measure and a twelve-listener panel to validate AAG's effectiveness, with results consistently showing AAG's superiority in enhancing accent consistency and speaker similarity.

Industry Impact and Future Directions

AAG's introduction offers a new technical path for the cross-lingual TTS field, particularly in multilingual applications. Potential applications of AAG include:

  • Multilingual Voice Assistants: Improving accent consistency across different language environments.
  • Film Dubbing: Achieving more natural cross-lingual dubbing effects.
  • Virtual Reality and Gaming: Enhancing the voice expressiveness of virtual characters.

Developer Recommendations

For developers, AAG can be integrated as a modular component into existing TTS systems to enhance the quality of cross-lingual speech synthesis. Additionally, developers can combine AAG with other technologies, such as emotion-aware speech synthesis, to further boost the expressiveness of synthesized speech.


Source: Hugging Face Daily Papers (2026-09-24)

— END —

Tags: #Hugging Face #Speech Synthesis #Cross-Lingual #TTS #Accent Consistency

Community Comments

Loading live comments and annotations…