Multi-Model Emotion Recognition from voice, face, and text sources using optimized Equivariant Quantum Dense Nested Neural Networks

Main Article Content

K. Murugesan, Dr. S. Sivakumar, Dr. J Venkatesh, S. Balaguru and N. Revathi

Abstract

The term Emotion Recognition(ER) is used for the Recognition and comprehension of human emotions have become essential aspects of technological progress in today's fast-paced, linked world where human–computer interaction is an essential part of everyday life. It is necessary to identify ER from audio,video and text data, because of that many techniques are implemented. However,the existing methods have lack of accuracy, precision and high error rate. To overcome the aforementioned problem, Equivariant Quantum Dense Nested Neural Networks(EQDNNN) with Fire Hawk Optimizer (EQDNN-FHO) is proposed for accurately identifying ER from audio, video and text data. In this input image is taken from two datasets such as IEMOCAP dataset and RAVDESS dataset. To improve the quality of the speech signals, text, and video signals, undesired artifacts are eliminated using methods including tokenization for text, Spectral Noise Gate (SNG) for audio and Color Wiener Filtering (CWF) for video. After that, the features such as audio, video and text are extracted using Short-Time Fourier Transform, Video Feature Extraction and Bag-of-Words is used for subsequent extraction. After that classification are done using Equivariant Quantum Dense Nested Neural Network (EQDNN)and optimization are done using Fire Hawk Optimizer (FHO) for detecting the different types of actions done by humans. The efficiency of the proposed EQDNN -FHO is analyzed using a dataset and attains 99.7% accuracy, 99.27% recall and attains better results compared with the existing methods. This method performed similarly in identifying positive and negative emotions when compared to human annotation. Moreover, it has demonstrated efficacy in precisely identifying the sentiment of ambiguous emotion parts that were deemed unsuitable in prior research. Advances in audio emotion recognition may have consequences for voice-activated virtual assistants, medical fields, and other industrial uses of automated communication.

Article Details

Section
Articles