Back to Model List

Index-Translate – Bilibili's Open-Source Multilingual Translation Model Family

AI Tech Editorial
RSS Feed
Index-Translate – Bilibili's Open-Source Multilingual Translation Model Family official screenshot
(Image source: official screenshot)

Executive Summary:

Index-Translate is an open-source multilingual translation model family developed by Bilibili's Index LLM team. It is built upon the Qwen3.5 foundation and opens its weights under the Apache-2.0 licen...

1. What is Index-Translate

Index-Translate is an open-source multilingual translation model family developed by Bilibili's Index LLM team. It is built upon the Qwen3.5 foundation and opens its weights under the Apache-2.0 license. This family supports text translation across 150 languages, offering three parameter scales: 2B, 9B, and 35B-A3B. It supports instruction constraints such as terminology unification, format preservation, and field-specific translation. In addition to text translation, the family includes three specialized capabilities: Index-Echo for speech translation, Index-Homura for syllable-controlled translation, and Index-NativeLong for long document translation, forming a comprehensive translation technology matrix that covers multiple scenarios including text, speech, audio-visual dubbing, and long documents.

index-translate-b official website screenshot
(Image source: official screenshot)

Technical positioning and domain: It belongs to the machine translation direction within the field of natural language processing. However, unlike traditional single-text translation models, Index-Translate integrates multiple capabilities—such as instruction-following, text-to-speech synthesis, syllable control, and long-context modeling—into a unified family system, positioning itself as a comprehensive translation infrastructure for content localization and internationalization.

Development background: Developed by Bilibili's Index LLM team, it leverages Bilibili's deep experience in video content, community culture, and multilingual content distribution. The team chose to fine-tune the Qwen3.5 model rather than train from scratch, reducing development costs while inheriting the base model's strong multilingual and reasoning capabilities. The open-source initiative aims to promote the democratization of translation technology and accelerate iteration through community feedback.

Core value: It addresses practical pain points in content localization scenarios, such as limited support for low-resource languages, adherence to instruction constraints, consistency in long documents, and matching dubbing duration. Its open-source license allows commercial use, enabling small and medium-sized teams to access near-industrial-grade translation capabilities, significantly lowering the technical barriers and costs of content localization.

Technical features: It employs a unified architecture to support both multilingual translation and instruction constraints. The speech translation pipeline incorporates speaker embeddings to achieve voice cloning. Syllable-controlled translation injects the target syllable count as an explicit condition. Long document translation uses native full-text input instead of traditional chunking strategies. Together, these features form a differentiated technical moat.

2. Key Features

  • Multilingual Text Translation: Supports translation of text and structured content across 150 languages, with the ability to execute instructions such as terminology unification, format preservation, and field-specific translation simultaneously. The model integrates translation capabilities with constraint-following within a single architecture, achieved through fine-tuning on multilingual corpora and instruction data. It is suitable for various scenarios, including general text, technical documents, and community content.

  • Index-Echo Speech Translation: Offers two modes: speech-to-speech (S2ST dubbing) and speech-to-text (S2TT subtitles). In S2ST mode, after translation, the model connects with CosyVoice speech synthesis, extracting speaker embeddings from the source audio and injecting them into the synthesis module to achieve voice cloning. Audio is processed in 60-second segments, with the first five segments' historical context retained to maintain coherence across the discourse. Ideal for movie and TV show dubbing, as well as course localization.

  • Index-Homura Syllable-Controlled Translation: Accepts target syllable count as an explicit input condition, automatically adjusting wording and sentence structure to meet the syllable budget. This feature injects syllable count into the prompt and uses sampling decoding to enable the model to simultaneously select wording and plan length during translation. On the SandGlass benchmark, 81.92% of outputs have a syllable deviation of no more than 10%, directly addressing the duration-matching needs of video dubbing.

  • Index-NativeLong Long Document Translation: Uses a native full-text input method, reading the entire document at once before generating the translation in a continuous manner. It leverages an ultra-long context window to maintain consistency in character names, terminology, and conceptual hierarchy throughout the generation process. Unlike traditional chunk-based translation methods that rely on adjacent context and automatic vocabulary tables, this approach does not require external vocabulary tables and can handle long texts such as novels and historical records.

  • Community Slang and Jargon Translation: Specifically optimized for community expressions such as homophones, abbreviations, and humorous language, with the MEME benchmark covering these scenarios. It can accurately translate Chinese community slang like "结芬" and "狒瘾" into corresponding semantic expressions in the target language, while preserving hashtags and emoji structures. Suitable for translating content such as bullet comments, dynamic posts, and forum threads for international distribution.

  • Structured Content Translation: Supports translating specified columns or fields in structured data formats such as CSV and JSON, while preserving non-translatable elements like headers and numbering. The translated output can be directly integrated into downstream business systems, reducing the cost of automation upgrades in localization workflows for e-commerce product information and announcements.

3. How to Use

  1. Environment Requirements: For local deployment, prepare a vLLM environment that supports Qwen3.5. It is recommended to use a Linux operating system with an NVIDIA GPU (at least 8GB VRAM for the 2B model, and more than 24GB VRAM for the 9B model). Python 3.9+ and CUDA environment must be installed in advance.

  2. Clone and Install: Execute git clone https://github.com/bilibili/Index-Translate.git to clone the repository and enter the project directory. Then run pip install -U vllm to install the latest version of vLLM, followed by pip install -r inference/llm/requirements.txt to install the project dependencies.

  3. Start the Service: Run vllm serve IndexTeam/Index-Translate-2B --host 127.0.0.1 --port 8000 --max-model-len 4096 to launch the model service locally. The default maximum context length is 4096 tokens, which can be adjusted based on hardware conditions. For the 9B and 35B-A3B models, the VRAM configuration must be increased accordingly.

  4. Perform Translation: Open a new terminal and run python inference/llm/translate.py "你好,世界" --target en --model IndexTeam/Index-Translate-2B to obtain the translated text. The client defaults to requesting http://127.0.0.1:8000/v1, and can be connected to a custom service using --base-url, --api-key, or environment variables OPENAI_BASE_URL and OPENAI_API_KEY.

  5. Specialized Capability Invocation: For syllable-controlled translation, use syllable_translate.py and pass the target syllable count via --syllables. For long document translation, use doc_translate.py and specify a fixed template direction such as zh-en or zh-ja with --direction. For speech translation, follow the S2TT subtitle guidelines or the S2ST dubbing guidelines in the repository to use s2tt.py or dub.py for processing audio and video.

  6. No-Deployment Experience: You can directly access the online Demo (https://index-translate.bilibili.com/) to experience the four types of translation: text, speech, syllable, and long documents. After deploying the model locally, you can also use Chrome/Edge/Firefox browser extensions to translate web pages, enabling seamless reading.

4. Pros and Cons Analysis

Pros
Broad Language Support: The text model supports translation among 150 languages, with particularly strong performance on less commonly used languages. The 35B-A3B model achieves an off-target rate of just 2.4% on the FLORES_minor_pair benchmark, significantly outperforming mainstream commercial APIs in terms of minor language coverage.
Strong Instruction-Following Ability: During translation, the model simultaneously enforces 10 types of constraints, including terminology consistency, format preservation, and field specification. The 9B model achieves an instruction-following score of 0.7725 on the instTrans_minor benchmark, the highest among all compared models, meeting complex production requirements.
Open Source and Open Access: Licensed under the Apache-2.0 protocol, the model allows for weight downloads, local deployment, and commercial use. Combined with an online demo and browser plugin, it reduces the experience barrier, offering greater transparency and controllability than closed-source APIs.
Comprehensive Specialized Capabilities: Features such as syllable control, long document consistency, and voice cloning directly address real-world pain points in scenarios like video dubbing and online content localization, creating a clear competitive advantage.

5. Comparative Analysis with Similar Tools

Comparison Dimension Index-Translate (Bilibili) Doubao-Seed-Translation (ByteDance) Google Translate API
Product Nature Open-source model family, weights downloadable for local deployment Closed-source commercial API, available only via ByteDance's Volcano Engine Closed-source commercial API
Language Coverage 150 languages 28 major languages (Chinese, English, Japanese, Korean, German, French, Spanish, Russian, etc.) 130+ languages
Instruction-following Supports 10 types of constraints including terminology, formatting, and content retention No dedicated constraint-following capability, relies on general prompt engineering Limited terminology table functionality, no complex constraints
Context Length 2B supports 262K tokens, 9B supports 229K tokens 4K context, 3K output length Dynamic window, segmented processing for long texts
Long Document Translation Index-NativeLong supports coherent full-text translation at the 64K level Basically unavailable (4K window cannot handle long texts) Segmented translation, consistency relies on terminology tables
Special Features Syllable control, voice cloning, community slang translation Chinese-English translation approaches DeepSeek-R1 Document and website translation integration
Deployment Method Local deployment (vLLM) + Online Demo Cloud API invocation Cloud API invocation
Open Source License Apache-2.0 Closed-source Closed-source

Selection Recommendations: For teams requiring support for less common languages, long document translation, or local audio-visual dubbing, Index-Translate's open-source nature and specialized capabilities offer clear advantages, particularly suitable for scenarios such as video content localization, online literature translation, and game localization. If the goal is to achieve top-tier Chinese-English translation quality and there are no data privacy concerns, Doubao-Seed-Translation leverages ByteDance's large model capabilities and performs well in general translation tasks. However, its 4K context limit makes it unsuitable for handling long documents.

For enterprise-level applications requiring stable API services, a mature ecosystem, and compliance guarantees, Google Translate and DeepL remain reliable choices. DeepL traditionally holds an edge in professional translation for European languages. However, neither supports local deployment, and both lack specialized capabilities for audio-visual scenarios such as syllable control and voice cloning. Overall, Index-Translate demonstrates strong comprehensive competitiveness among open-source solutions, though it still requires time to build up in terms of ecosystem maturity and out-of-the-box usability.

6. Editor's Summary

Index-Translate, as an open-source masterpiece from Bilibili's Index LLM team, demonstrates a unique and profound understanding of translation technology from a content platform perspective. Its technological innovation is reflected in three aspects: first, it deeply integrates instruction-following capabilities into the translation model, making constraints such as terminology consistency and format preservation no longer reliant on external post-processing, but rather internalized model capabilities. The instruction-following score of 0.7725 achieved by the 9B model on the instTrans_minor benchmark is a clear testament to this. Second, the syllable-controllable mechanism of Index-Homura provides a new perspective for combining translation with audiovisual production. The data showing that 81.92% of outputs have a syllable deviation of no more than 10% indicates that length control has already achieved practical value. Third, Index-NativeLong's native full-text translation strategy breaks away from the conventional framework of segmented translation, leveraging long context windows to address the long-standing issue of consistency in long documents.

In terms of practical value, this family of models covers the complete matrix of translation needs, ranging from text to voice, short sentences to long documents, and general translation to community slang. Moreover, the Apache-2.0 license permits commercial use, making it directly valuable for scenarios such as content localization, game localization, and e-commerce internationalization. The target users include: technical teams requiring support for less commonly used languages, video content creators and MCN organizations, professionals in online literature and publishing industries, as well as developers looking to break free from reliance on commercial APIs.

In terms of potential for development, Index-Translate's open strategy allows it to rapidly expand language coverage and scenario adaptability through community contributions. If it continues to invest in areas such as end-to-end speech translation architecture, specialized optimization for more language pairs, and the improvement of its toolchain ecosystem, it has the potential to become a significant force in the open-source translation field. Of course, as a new project, it still has room for growth in areas such as expanding evaluation systems, improving documentation, and strengthening community governance, making it worth keeping an eye on.

7. Application Scenarios

  • Video Localization Dubbing: Use Index-Echo to translate video content into target language dubbing, preserving the original speaker's voice characteristics. Combined with Index-Homura's syllable control, the dialogue can be synchronized with the length of the video clips. This is suitable for scenarios such as film and television dubbing for international markets, localization of online courses, and multilingual distribution of Vlogs, solving the problems of traditional dubbing that require re-recording and difficulty in matching lip movements and duration.

  • Multilingual Subtitle Generation: Automatically generate Chinese, Japanese, English, Spanish, and other multilingual subtitles for speeches, live streams, and short videos using the S2TT speech-to-text capability. Subtitle translation can be executed with format preservation instructions, ensuring that the timeline and style information remain intact, meeting the needs of cross-language content distribution and accessible viewing.

  • Community Content Localization: When translating comments, posts, and dynamic content, retain the original meaning of internet slang (e.g., "狒瘾" → "FFXIV itch"), topic tags, and emoji structures, ensuring that gaming and anime culture content remains contextually accurate in overseas markets. Special optimization on the MEME benchmark dataset ensures that such content is translated smoothly, avoiding literal and awkward translations.

  • Long Document and Web Novel Translation: Use Index-NativeLong to translate entire books, historical documents, and other long texts, maintaining consistency in character names, terminology, and conceptual hierarchy without the need for manual maintenance of a vocabulary list. This differs from block-by-block translation, which often leads to name drift, and is ideal for web novel localization, academic material translation, and publishing localization.

  • E-commerce and Product Documentation Localization: Translate specified columns in product CSV files according to instructions, preserving headers and numbering, or perform structured translation of announcements in JSON format, allowing the translated text to be directly integrated into subsequent business systems. This is suitable for process-driven tasks such as multilingual product listings on cross-border e-commerce platforms and batch translation of product manuals.

8. FAQ

Q: Can the Index-Translate model be used commercially?
A: Yes. The entire model family is open-sourced under the Apache-2.0 license, allowing free use, modification, and commercial deployment without requiring any licensing fees. However, the original copyright notice must be retained, and it is recommended to cite the model's source when using it.

Q: How should I choose between the 2B, 9B, and 35B-A3B models?
A: The 2B model is suitable for quick deployment in resource-constrained environments and simple translation tasks; the 9B model offers a good balance between instruction-following and translation quality, making it the recommended choice for most scenarios; the 35B-A3B model uses a MoE architecture, with only 3B activated parameters but the strongest overall capabilities, ideal for high-quality production environments. It is recommended to start with the 9B model if hardware conditions allow.

Q: Which languages does Index-Translate support? How effective is it for less commonly spoken languages?
A: The text model supports translation between 150 languages, with strong performance on less commonly spoken languages. The 35B-A3B model has an off-target rate of just 2.4% on the FLORES_minor_pair benchmark, indicating high accuracy in identifying and translating minor languages. However, the quality of specific language pairs may vary, so it is recommended to test with the online demo first for your target language pair.

Q: How can I ensure consistency in names and terminology during long document translation?
A: Use the native full-text translation mode of Index-NativeLong, inputting the entire document at once. The model leverages its ultra-long context window to remember character names and terminology usage from earlier parts of the document during generation, without requiring an external vocabulary table. The 2B model supports a context length of 262K tokens, while the 9B model supports 229K tokens.

Q: What is the latency and voice quality like for speech translation?
A: Index-Echo uses a cascaded pipeline, processing audio in 60-second segments. Latency depends on the length of the audio and GPU performance. Voice cloning is achieved by extracting the speaker embedding from the source audio and injecting it into the CosyVoice synthesis module. The fidelity of the voice restoration is affected by the quality of the source audio. For dubbing scenarios, it is recommended to use source audio that is clear and free of background noise for the best results.

9. Project Links

Related AI Model Articles

© All Rights Reserved. Some content on this site is partially generated by AI with human review.