Enhanced Hierarchical Attention Fusion Network for Multimodal Conversational Emotion Recognition using IEMOCAP and MELD Datasets
Keywords:
Multimodal Emotion Recognition, Conversational Emotion Recognition, Hierarchical Attention, Attention Fusion Network, IEMOCAP, MELD, Deep Learning, CNN, Bi-LSTM, Transformer, Multimodal Fusion, Facial Expression Recognition, Speech Emotion Recognition, Text Emotion Analysis.Abstract
Emotion recognition is still facing lot challenges to handle complex and contextual emotions of human in real time as well as in the field of healthcare too. Lot of unimodal and advanced machine learning techniques were used to handle those data. From the unimodal to multimodal approach which comprises of handling different contexts like facial expressions, speech and text related information using advanced techniques namely convolutional neural networks (CNN), bidirectional long short-term memory (Bi-LSTM), and Transformer architectures. Still the accuracy is challenging one with the existing datasets like CMU-MOSI and CMU-MOSEI and achieved nearly 80% accuracy with the help of fusion techniques with some limitations in balancing factor with different modality, high complex during computation and less adoptability with real time data. To propose a work with enhanced fusion techniques for multimodal emotion recognition with the IEMOCAP and MELD dataset since it is having large number of data than the CMU based dataset. Effective facial based features are handled by the CNN; speech variations are captured and analyzed efficiently with the help of Bi-LSTM and for text analysis transformers are used with the help of specialized encoders. The classification of emotional cues and multimodal contexts achieves better result than the existing one nearly 8-10% with the help of enhanced fusion techniques.