Modhak-TTS β€” Marathi domain-adaptation LoRA for Indic-Speak

Built with Indic-Speak from Bodhan AI / AI4Bharat.

A LoRA adapter that domain-adapts bodhan-ai/indic-speak (a closed-voice autoregressive codec-LM TTS model: Llama-3.2-3B backbone + SNAC 24 kHz + fine-tuned Vocos) toward Marathi, using only the model's existing named voices β€” no new voices, no cloning, no style.

Honest result: this is a negative / research artifact

This adapter did not improve Marathi synthesis quality by ASR-WER β€” it regressed it, worst on the studio anchor voice. It is published as a reproducible domain-adaptation experiment with an honest negative result, not as a quality upgrade. Do not deploy it expecting better Marathi TTS.

The regression is exactly the drift the project's design predicted in advance (training the studio voices Anagha/Chinmay on general-quality SPRINGLab Marathi drifts them toward that noisier timbre, trading anchor fidelity for domain coverage). The eval measured the pre-registered risk rather than hiding it β€” that, plus the decision log, is the point of the artifact.

Base vs fine-tuned (ASR-WER via bodhan-ai/indic-transcribe-core)

bucket base fine-tuned Ξ”
marathi_adversarial 0.353 0.353 +0.000
marathi_general_unseen 0.194 0.343 +0.149
marathi_stem_seen 0.202 0.262 +0.060
retention_rasa (anchor voice) 0.167 0.630 +0.463
retention_hindi 0.281 0.281 +0.000 (held)
retention_english 0.452 0.524 +0.071
marathi_longform (3-gram repetition) 0.000 0.001 no runaway

Caveat on magnitudes: this eval was run fast under a deadline (--limit 6 per bucket, --max-new-tokens 1000 vs the configured 2520, WER-only). A low token cap truncates longer fine-tuned outputs and inflates their WER, so treat magnitudes β€” especially the large anchor delta β€” as directional. The direction (no improvement; anchor drift; Hindi retention held) is robust.

Audio samples (A/B)

samples/base_model/ vs samples/finetuned/ hold the same held-out texts synthesized by each model (same voice, seed, settings) so the regression is audible. The clearest case is the Rasa anchor (retention_rasa_marathi__Chinmay.wav): for one sentence the base model emits 65 SNAC frames, the fine-tuned model 171 β€” it rambles ~2.6Γ— longer, which is exactly the intelligibility loss the WER jump (0.167β†’0.630) reports. samples/*/samples.json lists each file's text/speaker/frame count.

Training

Method LoRA (r=32, Ξ±=64, dropout=0.05) on q/v_proj (all 28 layers) + up/gate_proj (layers 8–24)
Trainable 21,430,272 params (0.645%); base, embeddings/lm_head (tied), SNAC, Vocos all frozen
Data 2,500 unique clips, mixture 56% SPRINGLab-MR / 24% Rasa-MR (anchor) / 15% Hindi / 5% English replay
Objective per-frame-position-weighted cross-entropy on audio + stop token
Steps 3,000 (~4.8 epochs/pool) Β· per-token loss 4.25 β†’ 3.15 Β· ~103 min on one A100-40GB

How to use

You need the gated base model (bodhan-ai/indic-speak, accept its terms) plus its SNAC 24 kHz codec and fine-tuned Vocos decoder β€” this repo contains only the LoRA adapter deltas, not the base weights.

from peft import PeftModel
from transformers import AutoModelForCausalLM
base = AutoModelForCausalLM.from_pretrained("bodhan-ai/indic-speak", dtype="bfloat16")
model = PeftModel.from_pretrained(base, "BNarayanaReddy/modhak-tts-mr-lora")
# then drive generation + SNAC/Vocos decode via the Indic-Speak inference contract
# (see the Modhak-TTS pipeline: scripts/sample.py, scripts/eval.py)

Only the closed 44-voice library is supported (Marathi natives: Anagha, Chinmay). No cloning, no new voices, no impersonation.

License & attribution

This adapter is a derivative of bodhan-ai/indic-speak and inherits the Bodhan AI / Indic Open Model License; it is distributed under that same license. The attribution "Built with Indic-Speak from Bodhan AI / AI4Bharat" is required in any product, demo, or research output built on it. Source datasets (SPRINGLab IndicTTS MR/HI/EN, ai4bharat/Rasa) retain their own licenses; no source audio is redistributed here.

Reproducibility

Full pipeline, decision log (D1–D18, incl. the pre-registered drift risk), and eval methodology live in the Modhak-TTS codebase (DECISIONS.md, docs/RESULTS.md, docs/EVAL.md).

Downloads last month
16
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for BNarayanaReddy/modhak-tts-mr-lora

Adapter
(2)
this model