SHAP-RTL: Addressing SHAP and LIME Visualization Issues in Right-to-Left Languages
By Mr.Xu
Published:
Summary:This research addresses the visualization issues of post-hoc explanation methods like SHAP and LIME when applied to right-to-left (RTL) languages such as Arabic, Urdu, Persian, and Hebrew. The proposed SHAP-RTL rendering layer corrects the reading direction and script shaping while preserving the original attribution values, feature ordering, and model outputs. Evaluations on hate and offensive language datasets in Urdu, Arabic, Hebrew, and Persian show that SHAP-RTL significantly improves rende
Background and Problem
Post-hoc explanation methods like SHAP and LIME are widely used to interpret text classifiers, but their visualizations are primarily designed for left-to-right (LTR) languages. When applied to right-to-left (RTL) languages such as Arabic, Urdu, Persian, and Hebrew, these methods suffer from issues like out-of-sequence tokens, broken connected letterforms, and plot layouts that do not follow the natural reading direction. These problems not only affect the readability of the explanations but also may lead to misinterpretations or misuse.
Main Contributions
This study proposes SHAP-RTL, a rendering layer specifically designed to address the visualization issues of SHAP and LIME in RTL languages. Its key features include:
- Reading Direction Correction: Adjusts the display order of tokens to align with the natural reading direction of RTL languages.
- Script Shaping Adjustment: Ensures the integrity of connected letters, avoiding breaks.
- Font Selection: Chooses appropriate fonts based on the language to enhance readability.
- Attribution Value Preservation: Maintains the original attribution values and feature ordering while correcting the visualization issues.
Experiments and Results
The research tested SHAP-RTL on hate and offensive language datasets in Urdu, Arabic, Hebrew, and Persian using TF-IDF and logistic regression classifiers. The evaluation, conducted through OCR round-trip testing over 200 feature words per language, showed that SHAP-RTL significantly improved rendering correctness, reducing character error rates from 0.820-0.979 to near zero.
Industry Impact
The introduction of SHAP-RTL fills a critical gap in the visualization of post-hoc explanation methods for RTL languages, making these methods more practical in a wider range of language environments. This is particularly important for the development and deployment of multilingual AI systems, especially in regions like the Middle East and South Asia where RTL languages are widely used.
Developer Recommendations
- Integrate SHAP-RTL: AI developers are advised to integrate SHAP-RTL into their existing explanation toolchains to enhance the visualization of RTL languages.
- Multilingual Support: Consider using SHAP-RTL when developing multilingual AI systems to ensure the effectiveness of explanation methods across different language environments.
- Continuous Optimization: Continuously optimize SHAP-RTL to adapt to more complex application scenarios as more languages and datasets are introduced.
— END —Source: ArXiv Machine Learning (cs.LG) (2026-09-26)
Tags: #SHAP #LIME #Visualization #RTL Languages #AI Explanation
Community Comments