Back

Leveraging Generative Artificial Intelligence for Enhanced Data Augmentation in Emotion Intensity Classification: A Comprehensive Framework for Cross-Dataset Transfer Learning

Wieczorek, J.; Jiang, X.; Palade, V.; Trela, J.

2026-03-03 health informatics
10.64898/2026.02.23.26346928 medRxiv
Show abstract

Data scarcity and stylistic heterogeneity pose major challenges for emotion intensity classification. This paper presents a cross-dataset augmentation framework that leverages prompt-conditioned generative models alongside deterministic and heuristic transformations to synthesize target-style examples for improved transfer learning. We introduce a unified taxonomy of augmentation strategies--Heuristic Lexical Perturbation (HLA), Prompt-Conditioned Generative Augmentation (CGA), Sequential Hybrid Pipeline (SHA), Rule-Guided Style Adaptation (DSGA), and Enhanced Hybrid Augmentation (EHA)--and detail an interpretability-oriented prompt engineering approach that conditions LLMs on authentic target exemplars and stylistic features extracted from the target dataset. Augmented datasets were evaluated using multi-dimensional quality metrics (transformation quality, stylistic consistency, BLEU/CHRF, Self-BLEU, uniqueness) and downstream classification via a two-phase BERT-LSTM training with rigorous statistical testing. During source dataset pretraining and subsequent target dataset fine-tuning, CGA achieved the highest single-method gains in F1 and accuracy (F1 = 0.8816; accuracy = 0.8819, 95% CI recalculated). HLA and SHA exhibited improved cross-domain stability, suggesting stronger domain-generalizable features. We observe systematic trade-offs between fluency, lexical diversity, and emotion fidelity: high surface similarity often correlates with classifier performance but does not fully capture affective authenticity. We discuss methodological pitfalls, propose best practices for emotion-aware augmentation, and provide reproducible artifacts (prompts, example transformations, evaluation scripts) to facilitate further research in affective NLP.

Matching journals

The top 9 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.