Index-Translate-9B

Online demo · GitHub · Technical report · Hugging Face collection · ModelScope collection

Index-Translate-9B is the 9B-parameter multilingual text model of the Index-Translate family, built on Qwen3.5. It supports text translation across 150 languages and follows translation instructions involving terminology, formatting, style, structure, context, and output length.

Capabilities

  • General multilingual translation: translate sentences, articles, subtitles, and other text across the supported language inventory.
  • Translation with instructions: preserve specified terms, JSON or other structures, code and placeholders; adapt style and resolve meaning from context.
  • Social and cultural translation: interpret community aliases, playful spelling, memes, and nonliteral expressions with attention to their intended meaning.
  • Scores 0.8789 FLORES COMET-22, 75.35 WMT26 Judge, and 0.8209 instTrans IFscore on the report's main evaluation.
  • Achieves the highest low-resource instTrans IFscore (0.7725) and lowest off-target rate (3.47%) among the report's comparison systems; low-resource quality is 0.5222.
  • Leads the report's compared models on MuST-Cinema COMET-22 (0.8979) and scores 0.7387 MEME.

Language coverage is documented in the technical report's language inventory. The speech and long-document models have their own packaged interfaces and language coverage.

Training

The technical report describes a shared multilingual mid-training recipe, followed by task-specific post-training:

  1. Multilingual mid-training: general text replay, monolingual text, ordinary parallel translations, and pivot-organized multilingual groups. The constant stage uses general / parallel / monolingual data at 1:1:1; decay uses general / core-pivot / full-pivot data at 1:4:2. The shared recipe totals 167.77B tokens.
  2. Specialist SFT and RL: general translation, instruction following, and meme translation specialists receive task-specific supervision. General translation RL combines XCOMET-XXL with language-validity and adequacy judgments; instruction RL uses Rubric-as-Reward with hard checks and graded constraints. RIVAL provides adaptive judge supervision.
  3. Expert integration and targeted distillation: parameter interpolation combines complementary specialists, then multi-teacher on-policy distillation (MOPD) addresses task types that remain weak after merging.

The evaluated 2B and 9B models combine the general, instruction, and meme specialists with weights 0.8 / 0.1 / 0.1, then receive targeted MOPD. Reported training-stage ablations use separate configurations; they are not the final model scores below.

Evaluation

Results below reproduce the latest technical report. Comparisons apply to the report's evaluation data and settings, and metric scales differ. Higher is better except for off-target rates. These are individual benchmark scores, separate from the demo's normalized radar categories.

The tables compare this checkpoint with the other Index-Translate sizes, Hy-MT2-7B, TranslateGemma-12B, and the Qwen3.5-9B base at a similar scale, and larger translation models and API baselines.

Main translation benchmarks

WMT26 Judge uses a 0–100 scale; the other columns use 0–1. instTrans quality and IFscore measure translation quality and instruction adherence separately. IFMTBench reports reference-based XCOMET-XXL and instruction adherence. Vertical is the equal-weight mean of the five domain scores below. North-Small-Translate is the 218B-A25B model; TranslateGemma is translategemma-12b-it.

Model FLORES COMET-22 WMT24++ COMET-22 WMT26 Judge instTrans Quality instTrans IFscore IFMTBench XCOMET-XXL IFMTBench IFscore Vertical mean MEME
Index-Translate-35B-A3B (preview) 0.8794 0.8586 76.76 0.6901 0.8336 0.7926 0.8991 0.8438 0.7405
Index-Translate-9B 0.8789 0.8601 75.35 0.6771 0.8209 0.7957 0.8760 0.8451 0.7387
Index-Translate-2B 0.8655 0.8489 60.26 0.5391 0.7569 0.7712 0.7584 0.8377 0.6443
Hy-MT2-1.8B 0.8522 0.8401 49.35 0.3181 0.4932 0.7493 0.7161 0.8314 0.3643
Hy-MT2-7B 0.8747 0.8593 60.51 0.5143 0.6079 0.8049 0.8741 0.8335 0.5139
Hy-MT2-30B-A3B 0.8787 0.8624 66.81 0.5725 0.6415 0.8177 0.9029 0.8459 0.5812
TranslateGemma-12B 0.8732 0.8524 71.19 0.4515 0.3068 0.8023 0.2892 0.8347 0.4281
North-Small-Translate (218B-A25B) 0.8784 0.8578 68.37 0.5697 0.5294 0.7657 0.8635 0.8357 0.6836
Qwen3.5-2B (base) 0.6983 0.6933 32.11 0.0999 0.2431 0.6197 0.3836 0.7557 0.2062
Qwen3.5-9B (base) 0.8316 0.8073 60.31 0.2467 0.0609 0.7341 0.5980 0.8199 0.5728
Qwen3.5-35B-A3B 0.8570 0.8290 71.33 0.3690 0.5204 0.7589 0.7822 0.8267 0.6447
DeepSeek-V4.1-Flash 0.8762 0.8510 83.55 0.6068 0.6374 0.7817 0.9090 0.8432 0.7424
GPT-5.6-Sol 0.8650 0.8469 89.10 0.6902 0.7624 0.7946 0.9367 0.8311 0.7194
Gemini 3.5 Flash Lite 0.8750 0.8497 79.52 0.6068 0.6374 0.7764 0.8854 0.8131 0.7034

Low-resource translation and instruction following

FLORES_minor_pair contains 104,000 inputs across 1,040 directions among 62 languages. instTrans_minor contains 2,793 instruction-translation tasks from Chinese or English into low-resource languages. Off-target is the percentage of outputs in a language other than the requested target; lower is better. The other scores use a 0–1 scale.

Model FLORES_minor_pair COMET-22 ↑ FLORES_minor_pair XCOMET-XXL ↑ FLORES_minor_pair off-target ↓ instTrans_minor Quality ↑ instTrans_minor IFscore ↑ instTrans_minor off-target ↓
Index-Translate-35B-A3B (preview) 0.8168 0.7164 2.4% 0.5151 0.7715 4.05%
Index-Translate-9B 0.7992 0.6805 4.0% 0.5222 0.7725 3.47%
Index-Translate-2B 0.7377 0.4817 4.2% 0.3050 0.6586 3.97%
Hy-MT2-1.8B 0.3163 0.2014 54.2% 0.0412 0.1313 65.27%
Hy-MT2-7B 0.4626 0.3334 35.7% 0.1121 0.2405 45.40%
Hy-MT2-30B-A3B 0.6746 0.5359 14.5% 0.2246 0.4449 15.47%
TranslateGemma-12B 0.8021 0.6217 1.7% 0.2412 0.2682 5.37%
North-Small-Translate (218B-A25B) 0.6855 0.4919 13.2% 0.2436 0.2902 26.89%
Qwen3.5-2B (base) 0.3716 0.1852 39.1% 0.0069 0.0549 73.54%
Qwen3.5-9B (base) 0.6774 0.4502 11.2% 0.1318 0.3532 18.51%
Qwen3.5-35B-A3B 0.7854 0.6257 3.5% 0.2874 0.4458 11.60%
DeepSeek-V4.1-Flash 0.8333 0.7297 1.3% 0.4793 0.5854 5.73%
GPT-5.6-Sol 0.7669 0.6918 10.4% 0.5757 0.6866 7.73%
Gemini 3.5 Flash Lite 0.8122 0.6995 3.3% 0.3927 0.5584 12.18%

Domain translation

COMET-22 is used for OPUS-Books, WMT Biomedical, MTNT, and MuST-Cinema; GuoFeng-Webnovel uses document-COMET. Higher is better. The complete comparison is in the report.

Model OPUS-Books Biomedical MTNT MuST-Cinema GuoFeng
Index-Translate-35B-A3B (preview) 0.8013 0.8681 0.8244 0.8962 0.8291
Index-Translate-9B 0.8010 0.8674 0.8246 0.8979 0.8345
Index-Translate-2B 0.7845 0.8635 0.8132 0.8930 0.8345

General capabilities

These answer-accuracy results (%) use matched evaluation subsets. Translation training can change general capabilities; these scores describe the translation checkpoints rather than the original Qwen3.5 models.

Model C-Eval GPQA-Diamond INCLUDE MMMLU
Index-Translate-35B-A3B (preview) 77.6 49.1 68.8 71.9
Index-Translate-9B 69.6 36.3 63.1 66.6
Index-Translate-2B 42.6 27.4 39.1 42.7

Evaluation protocol

  • The main FLORES evaluation uses 126,000 inputs, 420 directions among 21 languages; WMT24++ uses 6,000 inputs from English into 20 languages.
  • instTrans has 3,000 Chinese-to-20-language tasks across ten constraint types. GPT-5.6-Sol judges quality and instruction adherence separately. MEME contains 3,638 Chinese-to-English cases, covering 703 terms and 857 disambiguated senses.
  • Text translation evaluation uses greedy decoding and a 16,384-token context window, except WMT26, which uses zero-shot Chinese instructions, a 3,584-token context, and a 1,536-token output limit. WMT26 reference-free GPT-5.6-Sol scores are averaged over 20 target languages.
  • IFMTBench starts with 7,344 examples. Preprocessing removes 280 single-constraint examples with reference or language-label issues, leaving 7,064 inputs; 7,063 receive valid judgments. See the preprocessing note.

The 1,024-token output budget in the inference client below is a usage default, not the output budget used to produce every reported evaluation result. Detailed protocols and valid-judgment counts are in the technical report. The report describes instTrans and MEME; benchmark release plans are tracked in the GitHub TODO.

Inference

Default settings

The following settings describe the released translate.py client. Its decoding defaults apply across Index-Translate sizes; the commands below explicitly select this card's model.

Setting Client default and override
Decoding Greedy, temperature=0; override with --temperature
Output budget max_tokens=1024; override with --max-tokens
Thinking Disabled through chat_template_kwargs={"enable_thinking": False}; no thinking toggle in this CLI
Top-p / top-k / min-p Not explicitly set by this client; inference-server defaults apply
Presence / frequency / repetition penalties Not explicitly set by this client; inference-server defaults apply
Random seed Not explicitly set by this client; inference-server behavior applies
Stop strings / stop-token IDs / ignore-EOS Not explicitly set by this client; inference-server and model defaults apply
Model ID IndexTeam/Index-Translate-9B; INDEX_MODEL changes the default, and --model overrides it
Source language auto; override with --source / -s; auto omits the source-language name from the prompt
Target language en; override with --target / -t
Endpoint http://127.0.0.1:8000/v1; OPENAI_BASE_URL changes the default, and --base-url overrides it
API key EMPTY for local serving; OPENAI_API_KEY changes the default, and --api-key overrides it
Request format One user message; no system message is added by the client
Response streaming Not enabled by this client; it prints the complete response
Transport timeout and retries No explicit values set by the script; installed OpenAI / HTTPX client behavior applies
Proxy handling Environment proxies disabled for 127.0.0.1, localhost, and ::1; remote endpoints use the OpenAI client's usual behavior
Input / output Positional text or stdin; input is stripped and empty input is rejected; output is stripped and any emitted thinking block is removed

The client maps supported short language codes to Chinese language names in the prompt. Provide the full language name for a code outside its mapping. Its training-aligned prompt is:

请将以下{source-language}文本翻译为{target-language},直接输出翻译结果,不要进行任何解释。

{source-text}

With --source auto, {source-language} is omitted. Prompt length, source text, and generated output all count toward the served context. Parameters left unspecified above are not assigned new fixed values by this model card.

The shipped configuration sets max_position_embeddings=262144. The examples below use --max-model-len 32768, matching the GitHub 2B/9B Translate serving preset. Input plus output must fit within the served context, and available GPU memory constrains practical lengths. Increase the client output budget for longer translations.

The GitHub inference guide gives an approximate BF16 GPU-memory starting point of 24 GB for 9B, with additional memory needed for KV cache and longer contexts.

Serve with vLLM

Use a CUDA environment and a vLLM build with Qwen3.5 support. Start the server in one terminal:

pip install -U vllm
vllm serve IndexTeam/Index-Translate-9B \
  --served-model-name IndexTeam/Index-Translate-9B \
  --host 127.0.0.1 --port 8000 --max-model-len 32768

For a short-text smoke test, --max-model-len 4096 reduces the requested context. Multi-GPU deployments can add appropriate tensor-parallel arguments. See the serving guide for runtime details and other OpenAI-compatible servers.

OpenAI-compatible Python client

Install the client dependencies in another terminal:

pip install openai httpx

This example uses the training-aligned Chinese translation prompt and disables thinking explicitly:

import httpx
from openai import OpenAI

client = OpenAI(
    base_url="http://127.0.0.1:8000/v1",
    api_key="EMPTY",
    http_client=httpx.Client(trust_env=False),
)
response = client.chat.completions.create(
    model="IndexTeam/Index-Translate-9B",
    messages=[{
        "role": "user",
        "content": "请将以下文本翻译为英语,直接输出翻译结果,不要进行任何解释。\n\n你好,世界!",
    }],
    temperature=0,
    max_tokens=1024,
    extra_body={"chat_template_kwargs": {"enable_thinking": False}},
)
text = response.choices[0].message.content or ""
# Match the repository client's fallback if a thinking block is emitted.
if "</think>" in text:
    text = text.split("</think>", 1)[1]
print(text.strip().removeprefix("<think>").strip())

Repository client

After starting the server, use the public client with this model ID:

git clone https://github.com/bilibili/Index-Translate.git
cd Index-Translate
pip install -r inference/llm/requirements.txt
python inference/llm/translate.py \
  "你好,世界。今天天气不错。" --target en \
  --model IndexTeam/Index-Translate-9B

Use --source to specify the input language and --target for the output language. For languages without a mapped code in the client, provide the full language name. The client also accepts --base-url, --api-key, --max-tokens, and --temperature.

Examples and practical limits

The following is a recorded Index-Translate-9B case from the official demo. The task asks for Korean translation while preserving the JSON structure, stars, and Chinese hashtag.

Source:

{"title": "⭐2月13日例行维护公告⭐", "content": "#热血航线大和登场#"}

Recorded translation:

{"title": "⭐2월 13일 정기 점검 공지⭐", "content": "#热血航线大和登场#"}

See the captured inference cases and prompt examples for more directions and instruction formats.

Scores vary across languages and tasks. Combined constraints and low-resource directions remain harder than core-language translation, and successful individual examples do not guarantee that every instruction will be followed. For explicit target syllable counts, use Index-Homura; for native full-document translation, use Index-NativeLong with its dedicated templates.

Model family

Family Models Task
Index-Translate 2B / 9B / 35B-A3B (preview) Text translation and translation instructions
Index-Echo S2TT 2B / 9B Speech to translated subtitles
Index-Echo S2ST 2B / 9B Speech to speech with source-voice conditioning
Index-Homura 2B / 9B Translation toward a target syllable count
Index-NativeLong 2B / 9B Full-document translation

Naming: Index-NativeLong retains the published model IDs IndexTeam/Index-Nailong-2B and IndexTeam/Index-Nailong-9B. The family repository includes prompt guides, browser-extension code, and a video-dubbing pipeline. This model's ModelScope page mirrors the model card.

Citation

@techreport{indextranslate2026,
  author={Tianjiao Li and Mengran Yu and Chenyu Shi and Lusheng Zhang and
          Qisi Chen and Yanshan Zhou and Ji Qi and Jingying Liu and
          Yuang Feng and Ziang Cui and Tianxing Yan},
  title={Index-Translate: A Multilingual Translation Model Family --- Text, Speech, Controlled Dubbing, and Long-Document Translation},
  institution={Index LLM Team},
  year={2026},
  month={September},
  url={https://github.com/bilibili/Index-Translate/blob/main/docs/Index_Translate_Series_Technical_Report.pdf}
}

License

Apache-2.0. Questions and feedback are welcome through GitHub Issues.

Downloads last month
22
Safetensors
Model size
10B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for IndexTeam/Index-Translate-9B

Quantizations
3 models

Space using IndexTeam/Index-Translate-9B 1

Collection including IndexTeam/Index-Translate-9B