arXiv Releases COILD: A New Benchmark and Parallel Corpus for Indian Language Machine Translation
By Mr.Xu
Published:
Summary:arXiv has released COILD, a new resource specifically designed for Indian language machine translation (MT), addressing the scarcity of high-quality, Indic-centric data. COILD comprises 1.16 million human-translated and verified sentence pairs across 20 Indian language pairs, spanning the Indo-Aryan, Dravidian, Tibeto-Burman, and Austro-Asiatic language families, and covering eight domains with direct real-world applicability. Additionally, COILD introduces a domain-centric benchmark of 2,000 ex
Key Breakthroughs
- High-Quality Parallel Corpus: COILD includes 1.16 million human-translated and verified sentence pairs across 20 Indian language pairs, providing a rich and high-quality data resource for Indian language machine translation.
- Multi-Domain Coverage: The data spans eight domains with real-world applicability, including law, healthcare, and education, ensuring broad applicability.
- Domain-Centric Benchmark: A domain-centric benchmark of 2,000 expert-verified sentences is introduced for consistent multilingual and cross-lingual evaluation, providing a standardized metric for model performance.
Technical Highlights
- Multilingual Support: Covers multiple language pairs across Indo-Aryan, Dravidian, Tibeto-Burman, and Austro-Asiatic language families, supporting a wider range of multilingual machine translation research.
- Human Verification: All sentence pairs are human-translated and verified, ensuring data accuracy and quality.
- Domain Diversity: Data spans multiple domains, catering to diverse application scenarios.
Experimental Results
Experiments on two multilingual neural machine translation models, IndicTrans2-Distilled and NLLB-200, demonstrate that COILD significantly enhances model performance across language pairs, domains, automatic evaluation metrics, and human evaluation, underscoring the critical role of high-quality Indian language data in advancing machine translation.
Industry Impact and Developer Recommendations
- Advancing Indian Language Machine Translation: COILD provides a valuable data resource for Indian language machine translation research, helping improve model performance and application effectiveness.
- Promoting Multilingual Model Development: This resource not only benefits Indian language machine translation but also offers new insights and directions for multilingual model development.
- Developer Recommendations: Machine translation researchers and developers are encouraged to leverage COILD to enhance existing models and explore new application areas.
Future Outlook
The release of COILD marks a significant milestone in the field of Indian language machine translation. As more high-quality data is introduced and models are continuously optimized, the application prospects for Indian language machine translation will be even more promising.
— END —Source: ArXiv NLP/LLM (cs.CL) (2026-09-25)
Tags: #Machine Translation #Multilingual Models #Indian Languages #arXiv #Parallel Corpus
Community Comments