Related Experiment Videos
A Multi-layer Linguistic Profiling Framework for Emotion Recognition in Text: A Data-Driven Natural Language
Roaa S Bogdadi1, Marivel M De Guzman2, Rahaf Al Hasheem3
1Research, King Abdullah Medical Complex, Jeddah, SAU.
Abstract:
Background Emotion recognition in text underpins a wide range of human-centered artificial intelligence (AI) applications, yet existing frameworks either prioritize accuracy at the expense of interpretability or analyze emotional language through a single lens. This study developed and evaluated the Multi-layer Emotion Linguistic Profiling Framework (MELPF), a reproducible protocol integrating lexical, structural, and model-level interpretability analyses to characterize how emotional meaning is encoded in short text. Methods The Kaggle Emotions Dataset (16,000 training, 2,000 validations, and 2,000 test instances; six emotion categories) was used. The corpus exhibits pronounced class imbalance: joy (33.5%) and sadness (29.2%) dominate, while surprise is underrepresented (3.6%); in addition, all statistics were derived strictly from training instances to prevent data leakage. MELPF operates across five stages: preprocessing; lexical profiling (unigram/bigram distributions, term frequency-inverse document frequency {TF-IDF}, pointwise mutual information (PMI), and co-occurrence analysis); structural profiling (part-of-speech {POS} distribution and template identification); supervised classification using logistic regression, a linear support vector machine (SVM), random forest, a bidirectional long short-term memory (BiLSTM) network, and Bidirectional Encoder Representations from Transformers (BERT) (base-uncased variant); and cross-layer synthesis via feature-importance mapping and attention visualization. Results BERT achieved the highest macro-F1 (the macro-averaged F1 score, i.e., the harmonic mean of precision and recall; 0.92, accuracy 0.93), followed by BiLSTM (0.89), linear SVM (0.88), logistic regression (0.86), and random forest (0.81). Ablation confirmed that PMI-derived and POS-based features improved macro-F1 from 0.88 to 0.89 over the TF-IDF baseline. Lexical profiling identified category-discriminative unigrams and high-PMI bigrams, while structural profiling revealed a dominant first-person copular frame. Feature-importance scores and attention visualization confirmed alignment between salient tokens and model decisions. Conclusion MELPF offers a transparent, reproducible methodology for examining emotional meaning across multiple linguistic layers, bridging theoretical linguistics and applied NLP.
Related Concept Videos
Labeling Emotion
Non-Verbal Cues