ZICQ
中 Log in / Sign up
Newsroom LLMs & Foundation Models #Multilingual NLP #Persian #Code-Mixing #LLMs & Foundation Models #Corpus

PERCEPT: First Large-Scale Persian-English Code-Mixed Corpus Released to Empower Multilingual NLP Research

Avatar of Mr.Xu

By Mr.Xu

Published: · 5 views

中文阅读 (Chinese) English Version

Summary:PERCEPT is the first publicly available large-scale Persian-English code-mixed corpus annotated with Universal Dependencies (UD) part-of-speech (POS) tags. The dataset consists of 6,800 social media posts from platforms like X, Instagram, and Digikala. It employs an LLM-assisted annotation framework to ensure high reliability of the annotations. PERCEPT fills a critical gap in Persian-English code-mixed language resources, providing essential support for multilingual NLP model development and li


PERCEPT: A Large-Scale Persian-English Code-Mixed Corpus Released

Key Breakthroughs

  • First Large-Scale Corpus: PERCEPT is the first publicly available large-scale Persian-English code-mixed corpus annotated with Universal Dependencies (UD) part-of-speech (POS) tags, filling a critical gap in multilingual NLP resources.
  • Multi-Platform Data: The dataset includes 6,800 posts from social media platforms like X, Instagram, and Digikala, capturing diverse interaction scenarios.
  • LLM-Assisted Annotation: The corpus employs an LLM-assisted annotation framework to ensure high accuracy and consistency in the annotations.

Technical Highlights

  • Data Scale and Quality: With 6,800 high-quality annotated entries, PERCEPT offers a rich resource for multilingual NLP research.
  • Annotation Reliability: Human evaluation confirms the high agreement between the automatic annotations and gold-standard annotations, demonstrating the reliability of the dataset.
  • Multi-Dimensional Analysis: The release of PERCEPT enables the first comprehensive linguistic analysis of Persian-English code-mixing across multiple platforms, revealing variations in language use patterns.

Industry Impact

  • Multilingual NLP Model Development: PERCEPT provides crucial data for developing more accurate Persian-English code-mixed NLP models.
  • Linguistic Research: It offers new perspectives and tools for linguistic analysis of code-mixing phenomena.
  • Cross-Cultural Communication: The corpus promotes better understanding and communication between Persian and English speakers.

Developer Recommendations

  • Data Utilization: NLP researchers and developers are encouraged to leverage PERCEPT for model training and evaluation to enhance the performance of code-mixed language processing.
  • Extended Research: Further linguistic and sociolinguistic studies are recommended to explore more dimensions of code-mixing phenomena using PERCEPT.

Original Source

https://github.com/kalhorghazal/PERCEPT

— END —

Tags: #Multilingual NLP #Persian #Code-Mixing #LLMs & Foundation Models #Corpus

Community Comments

Loading live comments and annotations…