ZICQ
中 Log in / Sign up
Newsroom LLMs & Foundation Models #Large Language Models #Emotion Analysis #Requirements Engineering #Data Augmentation #Imbalance Mitigation

ArXiv Research: Leveraging Large Language Models for Fine-Grained Emotion Classification in Mobile App Reviews

Avatar of Mr.Xu

By Mr.Xu

Published:

中文阅读 (Chinese) English Version

Summary:ArXiv has released a study on leveraging Large Language Models (LLMs) for fine-grained emotion classification in mobile app reviews. The research, based on Plutchik's taxonomy, investigates the performance of LLMs in multi-label emotion classification under severe class imbalance. The results demonstrate that LLMs can significantly improve the accuracy of emotion classification, particularly for rare emotions. The study also releases the experimental pipeline, synthetic corpora, and fine-tuned c


Background and Objectives

Fine-grained emotion classification of mobile app reviews is a crucial component of requirements engineering, helping developers better understand user needs and feedback. However, due to its complexity, fine-grained emotion classification remains a challenging task. This study investigates the performance of Large Language Models (LLMs) in fine-grained, multi-label emotion classification, particularly under severe class imbalance, based on Plutchik's taxonomy.

Methodology

The study employed the following methods:

  • Model Comparison: Compared encoder-decoder models and decoder-only models under different prompting strategies.
  • Imbalance Mitigation Strategies: Applied a variety of imbalance mitigation strategies, including loss reweighting, resampling, and generative data augmentation.
  • Experimental Design: Selected the synthetic review generator and prompting strategy via intrinsic augmentation-utility ranking.

Key Findings

  1. Model Performance: The decoder-only few-shot prompting strategy performed best at baseline (macro-F1 0.642), while encoder-tuned models underperformed in both multi-label and binary ensemble formulations (0.387 and 0.450 respectively).
  2. Improvement Strategy: Pairing the best multi-label encoder with generative data augmentation and positive-weighted loss significantly closed the gap (+0.204) with up to three orders of magnitude lower inference latency than the decoders.
  3. Rare Emotion Handling: The largest gains were observed in the rarest emotions, with improvements from undetected to gains of up to +0.501 F1.

Conclusion and Impact

LLMs make fine-grained, multi-label emotion classification feasible for requirements engineering pipelines, albeit with modest macro-F1. The best formulation and mitigation strategy are backbone- and formulation-dependent. The research team has released the experimental pipeline, synthetic corpora, and fine-tuned checkpoints, providing valuable resources for future research and applications.

Developer Recommendations

  • Model Selection: For fine-grained emotion classification tasks, prioritize decoder models and combine them with generative data augmentation strategies.
  • Resource Optimization: For resource-constrained applications, consider using encoder models with positive-weighted loss strategies to achieve high performance at lower latency.
  • Data Augmentation: Utilize synthetic data generators to effectively mitigate class imbalance and improve the model's ability to recognize rare emotions.

Source: ArXiv NLP/LLM (cs.CL) (2026-10-06)

— END —

Tags: #Large Language Models #Emotion Analysis #Requirements Engineering #Data Augmentation #Imbalance Mitigation

Community Comments

Loading live comments and annotations…