Vocal-Emotion Visualizer: Voice-Based Emotion Detection & Image Generation

by Dasari Yaswanth Sri Balachandra, Jami Anjali Devi, P. Srinivasa Rao, Penumatcha Satya Sai Shivani, Veluvarthi Lasya Mounika

Published: July 17, 2026 • DOI: 10.51584/IJRIAS.2026.11060277

Abstract

Human feelings are essential in how people communicate and engage with computer systems. Spotting feelings from the manner human beings communicate has become a key area of study in synthetic intelligence. This Study introduces a machine named Vocal Emotion Visualizer, which identifies human emotions from speech, the usage of deep mastering and turns them into clean, visible paperwork. The gadget uses the Librosa library and Mel-Frequency Cepstral Coefficients (MFCC) to extract sound capabilities from recorded audio, and Wav2Vec 2.0 is used for live speech to obtain specified speech facts. These capabilities are then handled through a long short-term memory (LSTM) network to categorise feelings. The emotions discovered are changed into descriptive words to create a visible picture based on emotions, the usage of AI that generates snapshots. This machine facilitates people to apprehend emotions better in voice-based generation and makes interacting with computers greater herbal. Testing suggests that the device can correctly discover numerous feelings like happy, unhappy, angry, calm, and neutral.