muhamadul commited on
Commit
69cf7a5
·
verified ·
1 Parent(s): 104ca21

Upload fine-tuned XTTS-v2 for Wolof (3 epochs, A100)

Browse files
Files changed (5) hide show
  1. README.md +202 -0
  2. config.json +159 -0
  3. model.pth +3 -0
  4. speakers_xtts.pth +3 -0
  5. vocab.json +0 -0
README.md ADDED
@@ -0,0 +1,202 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ language:
3
+ - wo
4
+ - fr
5
+ - en
6
+ license: other
7
+ license_name: coqui-public-model-license
8
+ license_link: https://coqui.ai/cpml
9
+ base_model: coqui/XTTS-v2
10
+ tags:
11
+ - text-to-speech
12
+ - tts
13
+ - wolof
14
+ - xtts
15
+ - coqui
16
+ - fine-tuned
17
+ - african-languages
18
+ - senegal
19
+ datasets:
20
+ - WaxalNLP/wolof_speech
21
+ - google/fleurs
22
+ - keithito/lj_speech
23
+ library_name: coqui-tts
24
+ pipeline_tag: text-to-speech
25
+ ---
26
+
27
+ # XTTS-v2 Fine-tuned for Wolof 🇸🇳
28
+
29
+ A fine-tuned version of [Coqui XTTS-v2](https://huggingface.co/coqui/XTTS-v2) with improved Wolof language support. This model was fine-tuned to bring high-quality text-to-speech synthesis to the **Wolof language** — one of the most widely spoken languages in Senegal and West Africa.
30
+
31
+ ## Model Description
32
+
33
+ XTTS-v2 is a multilingual text-to-speech model that supports voice cloning. This fine-tuned version enhances performance on Wolof while retaining capabilities in French and English.
34
+
35
+ | Property | Value |
36
+ |----------|-------|
37
+ | **Base Model** | [coqui/XTTS-v2](https://huggingface.co/coqui/XTTS-v2) |
38
+ | **Architecture** | GPT-2 based encoder + HiFi-GAN decoder |
39
+ | **Parameters** | ~467M |
40
+ | **Audio Sample Rate** | 24,000 Hz |
41
+ | **Model Size** | 1.7 GB |
42
+
43
+ ## Training Details
44
+
45
+ ### Datasets
46
+
47
+ The model was fine-tuned on a curated multilingual dataset of **10,861 audio samples**:
48
+
49
+ | Dataset | Language | Samples | Description |
50
+ |---------|----------|---------|-------------|
51
+ | [WaxalNLP/wolof_speech](https://huggingface.co/datasets/WaxalNLP/wolof_speech) | Wolof | 1,042 | Wolof speech corpus |
52
+ | [google/fleurs](https://huggingface.co/datasets/google/fleurs) (wo) | Wolof | 2,819 | FLEURS Wolof split |
53
+ | [google/fleurs](https://huggingface.co/datasets/google/fleurs) (fr) | French | 2,000 | FLEURS French (subset) |
54
+ | [google/fleurs](https://huggingface.co/datasets/google/fleurs) (en) | English | 2,000 | FLEURS English (subset) |
55
+ | [keithito/lj_speech](https://huggingface.co/datasets/keithito/lj_speech) | English | 3,000 | LJSpeech (subset) |
56
+
57
+ French and English data was included to prevent catastrophic forgetting of multilingual capabilities during fine-tuning.
58
+
59
+ ### Fine-tuning Technique
60
+
61
+ - **Method**: GPT encoder fine-tuning (Stage 1) using `GPTTrainer` from Coqui TTS 0.22.0
62
+ - **All 898 parameters** were trained (full fine-tuning, no freezing)
63
+ - **Optimizer**: AdamW with learning rate 5e-6
64
+ - **Batch size**: 4 (effective batch size 16 with gradient accumulation of 4)
65
+ - **Epochs**: 3 (7,740 total steps)
66
+ - **Precision**: FP32 (mixed precision disabled — FP16 caused NaN losses on A100)
67
+ - **Hardware**: NVIDIA A100-SXM4-40GB
68
+ - **Training loss**: 0.924 → 0.767
69
+ - **Best eval loss**: 3.049
70
+
71
+ ### Key Training Decisions
72
+
73
+ 1. **FP32 over FP16**: Mixed precision training produced NaN losses with TTS 0.22.0 on A100 GPUs. FP32 training was stable and produced valid gradients throughout.
74
+ 2. **Multilingual data mix**: Including French, English, and LJSpeech alongside Wolof prevented the model from losing its multilingual voice cloning ability.
75
+ 3. **Low learning rate (5e-6)**: A conservative learning rate preserved the pre-trained model's strengths while allowing adaptation to Wolof phonology.
76
+
77
+ ## Evaluation Results
78
+
79
+ Comparison against the original XTTS-v2 base model on 9 test sentences (5 Wolof, 2 French, 2 English):
80
+
81
+ ### Overall
82
+
83
+ | Metric | Fine-tuned | Original | Δ |
84
+ |--------|-----------|----------|---|
85
+ | Speaker Similarity (↑) | **0.8273** | 0.8175 | +0.0099 |
86
+ | Wins (speaker match) | **7/9** | 2/9 | — |
87
+
88
+ ### Per-Language Speaker Similarity
89
+
90
+ | Language | Fine-tuned | Original | Δ | FT Win Rate |
91
+ |----------|-----------|----------|---|-------------|
92
+ | **Wolof** 🟢 | **0.8396** | 0.8193 | **+0.0203** | 5/5 (100%) |
93
+ | **French** 🟢 | **0.8266** | 0.8187 | +0.0080 | 1/2 |
94
+ | **English** 🔴 | 0.7974 | **0.8116** | -0.0142 | 1/2 |
95
+
96
+ **Key finding**: The fine-tuned model shows a **+2% improvement in speaker similarity for Wolof** while maintaining competitive performance in French and English.
97
+
98
+ ## Usage
99
+
100
+ ### With Coqui TTS
101
+
102
+ ```python
103
+ from TTS.api import TTS
104
+ import torch
105
+
106
+ # Load model
107
+ tts = TTS("tts_models/multilingual/multi-dataset/xtts_v2", gpu=False)
108
+
109
+ # Override with fine-tuned weights
110
+ model_path = "path/to/model.pth"
111
+ checkpoint = torch.load(model_path, map_location="cpu", weights_only=False)
112
+ tts.synthesizer.tts_model.load_state_dict(checkpoint["model"], strict=True)
113
+
114
+ # Generate Wolof speech
115
+ tts.tts_to_file(
116
+ text="Jàmm nga fanaan. Nanga def?",
117
+ speaker_wav="reference_audio.wav",
118
+ language="fr", # Wolof maps to French in XTTS-v2
119
+ file_path="output.wav"
120
+ )
121
+ ```
122
+
123
+ ### Direct with XttsModel
124
+
125
+ ```python
126
+ from TTS.tts.configs.xtts_config import XttsConfig
127
+ from TTS.tts.models.xtts import Xtts
128
+ import torch
129
+
130
+ config = XttsConfig()
131
+ config.load_json("config.json")
132
+ model = Xtts.init_from_config(config)
133
+ model.load_checkpoint(config, checkpoint_path="model.pth", vocab_path="vocab.json")
134
+ model.eval()
135
+
136
+ outputs = model.synthesize(
137
+ text="Jàmm nga fanaan.",
138
+ config=config,
139
+ speaker_wav="reference.wav",
140
+ language="fr",
141
+ )
142
+ ```
143
+
144
+ ## Files
145
+
146
+ | File | Size | Description |
147
+ |------|------|-------------|
148
+ | `model.pth` | 1.7 GB | Fine-tuned model weights (wrapped with `{"model": state_dict}`) |
149
+ | `config.json` | 4.3 KB | Model configuration |
150
+ | `vocab.json` | 353 KB | Tokenizer vocabulary |
151
+ | `speakers_xtts.pth` | 7.4 MB | Speaker embeddings |
152
+
153
+ ## Limitations
154
+
155
+ - **Wolof is mapped to French** (`language="fr"`) since XTTS-v2 has no native Wolof language token. This works well because Wolof and French share phonological similarities in the Senegalese context.
156
+ - English speaker similarity shows a slight regression (-1.4%) compared to the base model.
157
+ - The model was fine-tuned for 3 epochs only — longer training or larger Wolof datasets could yield further improvements.
158
+ - Voice quality depends on the reference audio provided for cloning.
159
+
160
+ ## Citation
161
+
162
+ If you use this model, please cite the original XTTS-v2 and the datasets:
163
+
164
+ ```bibtex
165
+ @misc{xtts-v2-wolof,
166
+ title={XTTS-v2 Fine-tuned for Wolof},
167
+ author={Muhamad Ul},
168
+ year={2026},
169
+ url={https://huggingface.co/muhamadul/xtts-v2-wolof}
170
+ }
171
+
172
+ @misc{casanova2024xtts,
173
+ title={XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model},
174
+ author={Casanova, Edresson and others},
175
+ year={2024},
176
+ publisher={Coqui AI}
177
+ }
178
+
179
+ @misc{waxalnlp-wolof,
180
+ title={Wolof Speech Dataset},
181
+ author={WaxalNLP},
182
+ url={https://huggingface.co/datasets/WaxalNLP/wolof_speech}
183
+ }
184
+
185
+ @inproceedings{conneau2023fleurs,
186
+ title={FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech},
187
+ author={Conneau, Alexis and others},
188
+ booktitle={IEEE SLT},
189
+ year={2023}
190
+ }
191
+ ```
192
+
193
+ ## Acknowledgments
194
+
195
+ - **[Coqui AI](https://coqui.ai/)** for the XTTS-v2 base model and TTS framework
196
+ - **[WaxalNLP](https://huggingface.co/WaxalNLP)** for the Wolof speech dataset
197
+ - **[Google FLEURS](https://huggingface.co/datasets/google/fleurs)** for multilingual speech data
198
+ - **[Daande](https://daande.aliune.com)** — The TTS studio powered by this model, built for Wolof speakers
199
+
200
+ ---
201
+
202
+ *Daande (ދައންދެ) means "voice" in Wolof. This project aims to bring modern AI voice technology to West African languages.*
config.json ADDED
@@ -0,0 +1,159 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "output_path": "output",
3
+ "logger_uri": null,
4
+ "run_name": "run",
5
+ "project_name": null,
6
+ "run_description": "\ud83d\udc38Coqui trainer run.",
7
+ "print_step": 25,
8
+ "plot_step": 100,
9
+ "model_param_stats": false,
10
+ "wandb_entity": null,
11
+ "dashboard_logger": "tensorboard",
12
+ "save_on_interrupt": true,
13
+ "log_model_step": null,
14
+ "save_step": 10000,
15
+ "save_n_checkpoints": 5,
16
+ "save_checkpoints": true,
17
+ "save_all_best": false,
18
+ "save_best_after": 10000,
19
+ "target_loss": null,
20
+ "print_eval": false,
21
+ "test_delay_epochs": 0,
22
+ "run_eval": true,
23
+ "run_eval_steps": null,
24
+ "distributed_backend": "nccl",
25
+ "distributed_url": "tcp://localhost:54321",
26
+ "mixed_precision": false,
27
+ "precision": "fp16",
28
+ "epochs": 1000,
29
+ "batch_size": 32,
30
+ "eval_batch_size": 16,
31
+ "grad_clip": 0.0,
32
+ "scheduler_after_epoch": true,
33
+ "lr": 0.001,
34
+ "optimizer": "radam",
35
+ "optimizer_params": null,
36
+ "lr_scheduler": null,
37
+ "lr_scheduler_params": {},
38
+ "use_grad_scaler": false,
39
+ "allow_tf32": false,
40
+ "cudnn_enable": true,
41
+ "cudnn_deterministic": false,
42
+ "cudnn_benchmark": false,
43
+ "training_seed": 54321,
44
+ "model": "xtts",
45
+ "num_loader_workers": 0,
46
+ "num_eval_loader_workers": 0,
47
+ "use_noise_augment": false,
48
+ "audio": {
49
+ "sample_rate": 22050,
50
+ "output_sample_rate": 24000
51
+ },
52
+ "use_phonemes": false,
53
+ "phonemizer": null,
54
+ "phoneme_language": null,
55
+ "compute_input_seq_cache": false,
56
+ "text_cleaner": null,
57
+ "enable_eos_bos_chars": false,
58
+ "test_sentences_file": "",
59
+ "phoneme_cache_path": null,
60
+ "characters": null,
61
+ "add_blank": false,
62
+ "batch_group_size": 0,
63
+ "loss_masking": null,
64
+ "min_audio_len": 1,
65
+ "max_audio_len": Infinity,
66
+ "min_text_len": 1,
67
+ "max_text_len": Infinity,
68
+ "compute_f0": false,
69
+ "compute_energy": false,
70
+ "compute_linear_spec": false,
71
+ "precompute_num_workers": 0,
72
+ "start_by_longest": false,
73
+ "shuffle": false,
74
+ "drop_last": false,
75
+ "datasets": [
76
+ {
77
+ "formatter": "",
78
+ "dataset_name": "",
79
+ "path": "",
80
+ "meta_file_train": "",
81
+ "ignored_speakers": null,
82
+ "language": "",
83
+ "phonemizer": "",
84
+ "meta_file_val": "",
85
+ "meta_file_attn_mask": ""
86
+ }
87
+ ],
88
+ "test_sentences": [],
89
+ "eval_split_max_size": null,
90
+ "eval_split_size": 0.01,
91
+ "use_speaker_weighted_sampler": false,
92
+ "speaker_weighted_sampler_alpha": 1.0,
93
+ "use_language_weighted_sampler": false,
94
+ "language_weighted_sampler_alpha": 1.0,
95
+ "use_length_weighted_sampler": false,
96
+ "length_weighted_sampler_alpha": 1.0,
97
+ "model_args": {
98
+ "gpt_batch_size": 1,
99
+ "enable_redaction": false,
100
+ "kv_cache": true,
101
+ "gpt_checkpoint": null,
102
+ "clvp_checkpoint": null,
103
+ "decoder_checkpoint": null,
104
+ "num_chars": 255,
105
+ "tokenizer_file": "",
106
+ "gpt_max_audio_tokens": 605,
107
+ "gpt_max_text_tokens": 402,
108
+ "gpt_max_prompt_tokens": 70,
109
+ "gpt_layers": 30,
110
+ "gpt_n_model_channels": 1024,
111
+ "gpt_n_heads": 16,
112
+ "gpt_number_text_tokens": 6681,
113
+ "gpt_start_text_token": null,
114
+ "gpt_stop_text_token": null,
115
+ "gpt_num_audio_tokens": 1026,
116
+ "gpt_start_audio_token": 1024,
117
+ "gpt_stop_audio_token": 1025,
118
+ "gpt_code_stride_len": 1024,
119
+ "gpt_use_masking_gt_prompt_approach": true,
120
+ "gpt_use_perceiver_resampler": true,
121
+ "input_sample_rate": 22050,
122
+ "output_sample_rate": 24000,
123
+ "output_hop_length": 256,
124
+ "decoder_input_dim": 1024,
125
+ "d_vector_dim": 512,
126
+ "cond_d_vector_in_each_upsampling_layer": true,
127
+ "duration_const": 102400
128
+ },
129
+ "model_dir": null,
130
+ "languages": [
131
+ "en",
132
+ "es",
133
+ "fr",
134
+ "de",
135
+ "it",
136
+ "pt",
137
+ "pl",
138
+ "tr",
139
+ "ru",
140
+ "nl",
141
+ "cs",
142
+ "ar",
143
+ "zh-cn",
144
+ "hu",
145
+ "ko",
146
+ "ja",
147
+ "hi"
148
+ ],
149
+ "temperature": 0.75,
150
+ "length_penalty": 1.0,
151
+ "repetition_penalty": 5.0,
152
+ "top_k": 50,
153
+ "top_p": 0.85,
154
+ "num_gpt_outputs": 1,
155
+ "gpt_cond_len": 30,
156
+ "gpt_cond_chunk_len": 4,
157
+ "max_ref_len": 30,
158
+ "sound_norm_refs": false
159
+ }
model.pth ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:43f387a3bd6d2dadcb61c641fe6c7aa528e667d670ed01e9953c589e2012c72d
3
+ size 1867854631
speakers_xtts.pth ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:f0f6137c19a4eab0cbbe4c99b5babacf68b1746e50da90807708c10e645b943b
3
+ size 7754818
vocab.json ADDED
The diff for this file is too large to render. See raw diff