Instructions to use BNarayanaReddy/modhak-tts-mr-lora with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use BNarayanaReddy/modhak-tts-mr-lora with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("/root/models/indic-speak") model = PeftModel.from_pretrained(base_model, "BNarayanaReddy/modhak-tts-mr-lora") - Notebooks
- Google Colab
- Kaggle
Modhak-TTS β Marathi domain-adaptation LoRA for Indic-Speak
Built with Indic-Speak from Bodhan AI / AI4Bharat.
A LoRA adapter that domain-adapts bodhan-ai/indic-speak (a closed-voice autoregressive codec-LM TTS model: Llama-3.2-3B backbone + SNAC 24 kHz + fine-tuned Vocos) toward Marathi, using only the model's existing named voices β no new voices, no cloning, no style.
Honest result: this is a negative / research artifact
This adapter did not improve Marathi synthesis quality by ASR-WER β it regressed it, worst on the studio anchor voice. It is published as a reproducible domain-adaptation experiment with an honest negative result, not as a quality upgrade. Do not deploy it expecting better Marathi TTS.
The regression is exactly the drift the project's design predicted in advance (training the studio voices Anagha/Chinmay on general-quality SPRINGLab Marathi drifts them toward that noisier timbre, trading anchor fidelity for domain coverage). The eval measured the pre-registered risk rather than hiding it β that, plus the decision log, is the point of the artifact.
Base vs fine-tuned (ASR-WER via bodhan-ai/indic-transcribe-core)
| bucket | base | fine-tuned | Ξ |
|---|---|---|---|
| marathi_adversarial | 0.353 | 0.353 | +0.000 |
| marathi_general_unseen | 0.194 | 0.343 | +0.149 |
| marathi_stem_seen | 0.202 | 0.262 | +0.060 |
| retention_rasa (anchor voice) | 0.167 | 0.630 | +0.463 |
| retention_hindi | 0.281 | 0.281 | +0.000 (held) |
| retention_english | 0.452 | 0.524 | +0.071 |
| marathi_longform (3-gram repetition) | 0.000 | 0.001 | no runaway |
Caveat on magnitudes: this eval was run fast under a deadline (--limit 6 per bucket,
--max-new-tokens 1000 vs the configured 2520, WER-only). A low token cap truncates longer fine-tuned
outputs and inflates their WER, so treat magnitudes β especially the large anchor delta β as
directional. The direction (no improvement; anchor drift; Hindi retention held) is robust.
Audio samples (A/B)
samples/base_model/ vs samples/finetuned/ hold the same held-out texts synthesized by each
model (same voice, seed, settings) so the regression is audible. The clearest case is the Rasa anchor
(retention_rasa_marathi__Chinmay.wav): for one sentence the base model emits 65 SNAC frames, the
fine-tuned model 171 β it rambles ~2.6Γ longer, which is exactly the intelligibility loss the WER
jump (0.167β0.630) reports. samples/*/samples.json lists each file's text/speaker/frame count.
Training
| Method | LoRA (r=32, Ξ±=64, dropout=0.05) on q/v_proj (all 28 layers) + up/gate_proj (layers 8β24) |
| Trainable | 21,430,272 params (0.645%); base, embeddings/lm_head (tied), SNAC, Vocos all frozen |
| Data | 2,500 unique clips, mixture 56% SPRINGLab-MR / 24% Rasa-MR (anchor) / 15% Hindi / 5% English replay |
| Objective | per-frame-position-weighted cross-entropy on audio + stop token |
| Steps | 3,000 (~4.8 epochs/pool) Β· per-token loss 4.25 β 3.15 Β· ~103 min on one A100-40GB |
How to use
You need the gated base model (bodhan-ai/indic-speak, accept its terms) plus its SNAC 24 kHz codec
and fine-tuned Vocos decoder β this repo contains only the LoRA adapter deltas, not the base weights.
from peft import PeftModel
from transformers import AutoModelForCausalLM
base = AutoModelForCausalLM.from_pretrained("bodhan-ai/indic-speak", dtype="bfloat16")
model = PeftModel.from_pretrained(base, "BNarayanaReddy/modhak-tts-mr-lora")
# then drive generation + SNAC/Vocos decode via the Indic-Speak inference contract
# (see the Modhak-TTS pipeline: scripts/sample.py, scripts/eval.py)
Only the closed 44-voice library is supported (Marathi natives: Anagha, Chinmay). No cloning, no new voices, no impersonation.
License & attribution
This adapter is a derivative of bodhan-ai/indic-speak and inherits the Bodhan AI / Indic Open
Model License; it is distributed under that same license. The attribution
"Built with Indic-Speak from Bodhan AI / AI4Bharat" is required in any product, demo, or research
output built on it. Source datasets (SPRINGLab IndicTTS MR/HI/EN, ai4bharat/Rasa) retain their own
licenses; no source audio is redistributed here.
Reproducibility
Full pipeline, decision log (D1βD18, incl. the pre-registered drift risk), and eval methodology live in
the Modhak-TTS codebase (DECISIONS.md, docs/RESULTS.md, docs/EVAL.md).
- Downloads last month
- 16