RSIS Repository Open-access research from RSIS International journals

International Journal of Research and Scientific Innovation (IJRSI)

Multimodal Fusion and Explainable Deep Learning for Synthetic Voice and Vishing Detection

byMahima BG; Pallavi GB

Published July 22, 2026  •  Vol. 13, Issue 7, pp. 218–229Open Access
DOI: 10.51244/IJRSI.2026.1307000015

Abstract

Synthetic speech generation and voice-cloning technologies have achieved unprecedented levels of realism, enabling numerous applications in accessibility, virtual assistants, and media production. However, these advancements also introduce significant risks, including identity fraud, impersonation attacks, misinformation, and security breaches. This paper proposes a multimodal fusion framework for synthetic voice detection that combines handcrafted acoustic features with deep spectrogram representations to improve detection robustness and generalization. The proposed architecture employs a Convolutional Neural Network–Bidirectional Long ShortTerm Memory (CNN-BiLSTM) network to capture both spectral artifacts and temporal inconsistencies characteristic of AI-generated speech. To enhance transparency and interpretability, an explainability module incorporating attention visualization and feature attribution techniques is integrated into the detection pipeline. Furthermore, the framework is deployed through a real-time inference interface, demonstrating its practical applicability in cybersecurity, digital forensics, and media authentication scenarios. The findings highlight the effectiveness of combining deep learning, multimodal feature fusion, and explainable artificial intelligence to address the growing challenge of synthetic speech detection.

Keywords: Multimodal, Fusion, Explainable

JournalInternational Journal of Research and Scientific Innovation (IJRSI)
ISSN2321-2705
Volume / IssueVolume 13, Issue 7
Pages218–229
Publication dateJuly 22, 2026
DOI10.51244/IJRSI.2026.1307000015
PublisherRSIS International
LicenseOpen Access

How to cite this article

Mahima BG, & Pallavi GB (2026). Multimodal Fusion and Explainable Deep Learning for Synthetic Voice and Vishing Detection. International Journal of Research and Scientific Innovation (IJRSI), 13(7), 218-229. https://doi.org/10.51244/IJRSI.2026.1307000015

BibTeX

@article{Mahima2026,
  title   = {Multimodal Fusion and Explainable Deep Learning for Synthetic Voice and Vishing Detection},
  author  = {Mahima BG and Pallavi GB},
  journal = {International Journal of Research and Scientific Innovation (IJRSI)},
  volume  = {13},
  number  = {7},
  pages   = {218--229},
  year    = {2026},
  doi     = {10.51244/IJRSI.2026.1307000015},
  publisher = {RSIS International}
}