Majority-Class Collapse in Low-Resource E-Commerce Emotion Mining: Diagnostic Evaluation of LSTM on Imbalanced and Linguistically Noisy Indonesian Reviews
Product reviews on e-commerce platforms such as Shopee provide valuable emotional information that can support customer behavior analysis and business decision-making. However, automatically identifying fine-grained emotions remains challenging because user-generated reviews often contain informal language, abbreviations, spelling variations, and severe class imbalance. Although Long Short-Term Memory (LSTM) networks have demonstrated strong capabilities in sequential text modeling, their robustness under highly imbalanced and low-resource real-world conditions remains insufficiently investigated. This study evaluates an LSTM-Word2Vec architecture for multi-class emotion classification (positive, neutral, and negative) using 306 Indomie product reviews collected from Shopee through web crawling. The proposed methodology includes comprehensive text preprocessing, custom Word2Vec semantic embedding generation, LSTM-based classification, and performance evaluation using Accuracy, Precision, Recall, and F1-Score. To further quantify prediction deviations, numerical error metrics including Mean Squared Error (MSE), Root Mean Squared Error (RMSE), and Mean Absolute Error (MAE) are also employed. Experimental results reveal a pronounced majority-class collapse, where the model achieved only 31% accuracy by predicting nearly all testing samples as the negative class. Detailed diagnostic analysis indicates that this failure is primarily caused by the combined effects of severe class imbalance, the limited dataset size, and noisy labels generated through heuristic proxy annotation. These findings demonstrate that conventional LSTM architectures are highly vulnerable to noisy, colloquial marketplace language when adequate data quality is unavailable. Consequently, this research establishes a realistic baseline for e-commerce emotion mining and emphasizes that model complexity alone cannot overcome fundamental data limitations.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex