Acoustic Classification for Deaf Persons Using Machine Learning: A CNN-Based Approach with MFCC Feature Extraction on the UrbanSound8K Dataset

https://doi.org/10.5281/zenodo.21134101

Authors

  • Muhammad Usama Ahmed*
  • Laila Younas

Abstract

This work focuses on enhancing the availability and safety of individuals who are deaf by utilizing machine learning techniques for the automatic classification of everyday environmental sounds. A diverse dataset of urban sounds, UrbanSound8K, was used to train a Convolutional Neural Network (CNN) to categorize ten distinct urban sound classes. The system is designed to provide a visual or vibrotactile representation when a recognized sound is detected, offering a tool for improving accessibility and communication for people unable to hear. Mel-Frequency Cepstral Coefficients (MFCCs) were extracted using the Librosa library and used as the primary feature representation for classification. The proposed CNN architecture, comprising three convolutional blocks with ReLU activation, max-pooling, and dropout regularisation, was trained using the Adam optimiser and sparse categorical cross-entropy loss. The CNN model achieved a training accuracy of 92%, with an overall test classification accuracy of 92.01%, sensitivity of 97.12%, specificity of 90.01%, precision of 92.93%, and an F1-score of 96.45%, outperforming both Artificial Neural Network (ANN) and Recurrent Neural Network (RNN) baselines evaluated under identical experimental conditions. Various visualisation techniques — including bar graphs, line graphs, heat maps, spectrograms, and comparison tables — were employed to gain deeper insight into the urban sound content. A conceptual real-time embedded deployment using a Raspberry Pi 3B+ and a wrist-worn vibration motor is proposed to translate the trained model into a practical assistive device. The results confirm the strong potential of MFCC-based CNN models for urban acoustic classification in the deaf-assistive-technology domain.

Index Terms: Acoustic classification; assistive technology; convolutional neural network; deep learning; hearing impairment; mel-frequency cepstral coefficients; urban sound recognition; UrbanSound8K; vibrotactile feedback; sound awareness.

Downloads

Published

2024-11-25

How to Cite

Muhammad Usama Ahmed*, & Laila Younas. (2024). Acoustic Classification for Deaf Persons Using Machine Learning: A CNN-Based Approach with MFCC Feature Extraction on the UrbanSound8K Dataset: https://doi.org/10.5281/zenodo.21134101. Spectrum of Engineering Sciences, 2(4), 592–616. Retrieved from https://thesesjournal.com.medicalsciencereview.com/index.php/1/article/view/3403