PERCEPT: First Large-Scale Persian-English Code-Mixed Corpus Released to Empower Multilingual NLP Research
By Mr.Xu
Published: · 5 views
Summary:PERCEPT is the first publicly available large-scale Persian-English code-mixed corpus annotated with Universal Dependencies (UD) part-of-speech (POS) tags. The dataset consists of 6,800 social media posts from platforms like X, Instagram, and Digikala. It employs an LLM-assisted annotation framework to ensure high reliability of the annotations. PERCEPT fills a critical gap in Persian-English code-mixed language resources, providing essential support for multilingual NLP model development and li
PERCEPT: A Large-Scale Persian-English Code-Mixed Corpus Released
Key Breakthroughs
- First Large-Scale Corpus: PERCEPT is the first publicly available large-scale Persian-English code-mixed corpus annotated with Universal Dependencies (UD) part-of-speech (POS) tags, filling a critical gap in multilingual NLP resources.
- Multi-Platform Data: The dataset includes 6,800 posts from social media platforms like X, Instagram, and Digikala, capturing diverse interaction scenarios.
- LLM-Assisted Annotation: The corpus employs an LLM-assisted annotation framework to ensure high accuracy and consistency in the annotations.
Technical Highlights
- Data Scale and Quality: With 6,800 high-quality annotated entries, PERCEPT offers a rich resource for multilingual NLP research.
- Annotation Reliability: Human evaluation confirms the high agreement between the automatic annotations and gold-standard annotations, demonstrating the reliability of the dataset.
- Multi-Dimensional Analysis: The release of PERCEPT enables the first comprehensive linguistic analysis of Persian-English code-mixing across multiple platforms, revealing variations in language use patterns.
Industry Impact
- Multilingual NLP Model Development: PERCEPT provides crucial data for developing more accurate Persian-English code-mixed NLP models.
- Linguistic Research: It offers new perspectives and tools for linguistic analysis of code-mixing phenomena.
- Cross-Cultural Communication: The corpus promotes better understanding and communication between Persian and English speakers.
Developer Recommendations
- Data Utilization: NLP researchers and developers are encouraged to leverage PERCEPT for model training and evaluation to enhance the performance of code-mixed language processing.
- Extended Research: Further linguistic and sociolinguistic studies are recommended to explore more dimensions of code-mixing phenomena using PERCEPT.
Original Source
https://github.com/kalhorghazal/PERCEPT
— END —Tags: #Multilingual NLP #Persian #Code-Mixing #LLMs & Foundation Models #Corpus
Community Comments