ZICQ
中 Log in / Sign up
ZICQ Info Open Source AI #ArXiv #Multi-Dialect Arabic #Detoxification #NLP #Dataset

ArXiv Releases AraDetox: A Multi-Dialect Arabic Detoxification Dataset

Avatar of Mr.Xu

By Mr.Xu

Published: · 4 views

中文阅读 (Chinese) English Version

Summary:ArXiv has released AraDetox, a multi-dialect Arabic detoxification dataset comprising 10,500 harmful social media posts and 84,000 detoxified rewrites across Modern Standard Arabic, Gulf, Levantine, and Egyptian Arabic. The dataset addresses the gap in Arabic text detoxification research, which has received less attention compared to harmful language detection. Human evaluation and automatic analyses show that detoxification primarily involves meaning-preserving rewriting with substantial lexica


Key Breakthroughs

ArXiv has released AraDetox, a multi-dialect Arabic detoxification dataset, addressing the gap in detoxification research within the field of Arabic harmful language detection. Key features of AraDetox include:

  • Multi-Dialect Coverage: The dataset spans Modern Standard Arabic, Gulf, Levantine, and Egyptian Arabic, providing extensive dialect support.
  • Large-Scale Data: It contains 10,500 harmful social media posts and 84,000 detoxified rewrites, offering ample data for model training.
  • Detoxification Mechanism: Research indicates that detoxification primarily involves meaning-preserving rewriting, with substantial lexical and structural changes while maintaining high semantic similarity.
  • Evaluation Methods: The dataset employs both human evaluation and automatic analyses, including lexical change, semantic preservation, sentiment, and dialectal style analysis, ensuring the quality and accuracy of the rewrites.

Technical Highlights

  1. Multi-Modal Detoxification: AraDetox focuses not only on the toxicity of the text but also on dialect style alignment, ensuring that the rewritten text is consistent with the original in both semantics and style.
  2. Automated and Human Verification: The dataset uses GPT-5 and Gemini 2.5 Flash to generate detoxified texts and verifies their quality through human evaluation, ensuring high-quality data.
  3. Open Access: The dataset is publicly available on GitHub, supporting further research and application development.

Industry Impact

The release of AraDetox brings new opportunities to the Arabic NLP field, particularly in the following areas:

  • Safe Text Generation: It provides critical data support for developing safer Arabic text generation models.
  • Multi-Dialect NLP Research: It promotes the understanding and application of different Arabic dialects.
  • Social Applications: It helps build safer social media platforms and content moderation systems.

Recommendations for Developers

  • Data Application: Researchers are encouraged to leverage AraDetox for training and evaluating Arabic detoxification models.
  • Model Improvement: Improve the semantic preservation and dialect style alignment capabilities of existing models using AraDetox data.
  • Cross-Domain Application: Explore the potential of AraDetox in multi-modal tasks, such as text-image detoxification.

Source: ArXiv cs.AI (2026-08-24)

— END —

Tags: #ArXiv #Multi-Dialect Arabic #Detoxification #NLP #Dataset

Community Comments

Loading live comments and annotations…