RSIS Repository Open-access research from RSIS International journals

International Journal of Research and Scientific Innovation (IJRSI)

Emotion Recognition Using Machine Learning and Computer Vision: A Hybrid CNN–ViT Multimodal Framework

byDeepa Chandrashekhar Rathod; Bhumika B K; Arpitha G A; Brunda U Jajur; Usha K

Published May 19, 2026  •  Vol. 13, Issue 4, pp. 2922–2933Open Access
DOI: 10.51244/IJRSI.2026.1304000249

Abstract

Automatic recognition of human emotions from facial expressions and multimodal signals constitutes a foundational challenge in affective computing and human–computer interaction, with broad applications spanning healthcare monitoring, autonomous vehicle safety, educational technology, and social robotics. Despite remarkable progress driven by deep learning, particularly convolutional neural networks (CNNs), recurrent neural networks (RNNs), and Vision Transformers (ViT), achieving robust emotion recognition in unconstrained, real-world environments remains an open problem. This paper presents a comprehensive synthesis of over twenty-five state-of-the-art studies on facial and multimodal emotion recognition, encompassing CNN-based systems trained on FER2013, CK+, RAF-DB, and AffectNet; transformer-based hybrid architectures; and multimodal fusion systems integrating facial, speech, and electroencephalography (EEG) cues evaluated on RAVDESS, IEMOCAP, CMU-MOSEI, eNTERFACE'05, and MAHNOB-HCI. Building upon these insights, this work proposes a novel Hybrid CNN–ViT Multimodal Emotion Recognition (HCV-MER) framework comprising: (i) a squeeze-and-excitation ResNet combined with a Vision Transformer facial backbone incorporating region-specific attention over eyes and mouth; (ii) a lightweight temporal aggregation unit for video-level inference; and (iii) a cross-modal attention fusion module integrating facial and speech streams. Experimental evaluations target FER2013, RAF-DB, CK+, and RAVDESS using TensorFlow, PyTorch, and OpenCV. Expected improvements over baseline CNN architectures range from five to ten percentage points on challenging in-the-wild benchmarks. The paper further analyzes unresolved challenges including cross-domain generalization, demographic fairness, micro-expression recognition, and privacy-preserving deployment.

Keywords: Facial emotion recognition; Convolutional neural networks

JournalInternational Journal of Research and Scientific Innovation (IJRSI)
ISSN2321-2705
Volume / IssueVolume 13, Issue 4
Pages2922–2933
Publication dateMay 19, 2026
DOI10.51244/IJRSI.2026.1304000249
PublisherRSIS International
LicenseOpen Access

How to cite this article

Deepa Chandrashekhar Rathod, Bhumika B K, Arpitha G A, Brunda U Jajur, & Usha K (2026). Emotion Recognition Using Machine Learning and Computer Vision: A Hybrid CNN–ViT Multimodal Framework. International Journal of Research and Scientific Innovation (IJRSI), 13(4), 2922-2933. https://doi.org/10.51244/IJRSI.2026.1304000249

BibTeX

@article{Deepa2026,
  title   = {Emotion Recognition Using Machine Learning and Computer Vision: A Hybrid CNN–ViT Multimodal Framework},
  author  = {Deepa Chandrashekhar Rathod and Bhumika B K and Arpitha G A and Brunda U Jajur and Usha K},
  journal = {International Journal of Research and Scientific Innovation (IJRSI)},
  volume  = {13},
  number  = {4},
  pages   = {2922--2933},
  year    = {2026},
  doi     = {10.51244/IJRSI.2026.1304000249},
  publisher = {RSIS International}
}