ZICQ
中 Log in / Sign up
Newsroom LLMs & Foundation Models #Machine Translation #Multilingual Models #Indian Languages #arXiv #Parallel Corpus

arXiv Releases COILD: A New Benchmark and Parallel Corpus for Indian Language Machine Translation

Avatar of Mr.Xu

By Mr.Xu

Published:

中文阅读 (Chinese) English Version

Summary:arXiv has released COILD, a new resource specifically designed for Indian language machine translation (MT), addressing the scarcity of high-quality, Indic-centric data. COILD comprises 1.16 million human-translated and verified sentence pairs across 20 Indian language pairs, spanning the Indo-Aryan, Dravidian, Tibeto-Burman, and Austro-Asiatic language families, and covering eight domains with direct real-world applicability. Additionally, COILD introduces a domain-centric benchmark of 2,000 ex


Key Breakthroughs

  • High-Quality Parallel Corpus: COILD includes 1.16 million human-translated and verified sentence pairs across 20 Indian language pairs, providing a rich and high-quality data resource for Indian language machine translation.
  • Multi-Domain Coverage: The data spans eight domains with real-world applicability, including law, healthcare, and education, ensuring broad applicability.
  • Domain-Centric Benchmark: A domain-centric benchmark of 2,000 expert-verified sentences is introduced for consistent multilingual and cross-lingual evaluation, providing a standardized metric for model performance.

Technical Highlights

  • Multilingual Support: Covers multiple language pairs across Indo-Aryan, Dravidian, Tibeto-Burman, and Austro-Asiatic language families, supporting a wider range of multilingual machine translation research.
  • Human Verification: All sentence pairs are human-translated and verified, ensuring data accuracy and quality.
  • Domain Diversity: Data spans multiple domains, catering to diverse application scenarios.

Experimental Results

Experiments on two multilingual neural machine translation models, IndicTrans2-Distilled and NLLB-200, demonstrate that COILD significantly enhances model performance across language pairs, domains, automatic evaluation metrics, and human evaluation, underscoring the critical role of high-quality Indian language data in advancing machine translation.

Industry Impact and Developer Recommendations

  • Advancing Indian Language Machine Translation: COILD provides a valuable data resource for Indian language machine translation research, helping improve model performance and application effectiveness.
  • Promoting Multilingual Model Development: This resource not only benefits Indian language machine translation but also offers new insights and directions for multilingual model development.
  • Developer Recommendations: Machine translation researchers and developers are encouraged to leverage COILD to enhance existing models and explore new application areas.

Future Outlook

The release of COILD marks a significant milestone in the field of Indian language machine translation. As more high-quality data is introduced and models are continuously optimized, the application prospects for Indian language machine translation will be even more promising.


Source: ArXiv NLP/LLM (cs.CL) (2026-09-25)

— END —

Tags: #Machine Translation #Multilingual Models #Indian Languages #arXiv #Parallel Corpus

Community Comments

Loading live comments and annotations…