Speech Enhancement Using Convolutional Denoising Auto-Encoder

dc.contributor.advisorAkhand, Prof. Dr. Muhammad Aminul Haque
dc.contributor.authorShahriyar, Shaikh Akib
dc.date.accessioned2019-05-15T06:12:34Z
dc.date.available2019-05-15T06:12:34Z
dc.date.issued2019-05
dc.descriptionThis thesis is submitted to the Department of Computer Science and Engineering, Khulna University of Engineering & Technology in partial fulfillment of the requirements for the degree of Master of Science in Computer Science and Engineering, May 2019.
dc.descriptionCataloged from PDF Version of Thesis.
dc.descriptionIncludes bibliographical references (pages 31-35).
dc.description.abstractSpeech signals are complex in nature with respect to other forms of communication media such as text or image. Different forms of noises (e.g., additive noise, channel noise, babble noise) interfere with the speech signals and drastically hamper the quality of the speech. Enhancement of speech signals is a daunting task considering multiple forms of noises while denoising a speech signal. Certain analog noise eliminator models have been studied over the years for this purpose. Researchers have also delved into some machine learning techniques (e.g., artificial neural network) to enhance speech signals. In this study, a speech enhancement system is investigated using Convolutional Denoising Autoencoder (CDAE). Convolutional neural network (CNN) is a special kind of deep neural networks which is suitable for 2D structured input (e.g., image) and CDAE is a CNN based special kind of Denoising Autoencoder. CDAE takes advantages from the 2D structured inputs of the features extracted from speech signals and also considers the local temporal relationship among the features. In the proposed system, CDAE is trained considering features from noisy speech signal as input and clean speech features as desired output. The proposed CDAE based method has been tested on a benchmark dataset, called Speech Command Dataset, and attained 80% similarity between denoised speech and actual clean speech. The proposed system achieved perceptual evaluation of speech quality (PESQ) value of 2.43 which outperformed other related existing methods.
dc.identifier.otherID 1707556
dc.identifier.otherhttp://dspace.kuet.ac.bd/handle/20.500.12228/516
dc.identifier.urihttp://hdl.handle.net/20.500.12228/516
dc.language.isoen_US
dc.publisherKhulna University of Engineering & Technology (KUET), Khulna, Bangladesh
dc.sourceKUET Institutional Repository
dc.subjectSpeech Enhancement
dc.subjectSpeech Cleaning
dc.subjectConvolutional Denoising Autoencoder (CDAE)
dc.subjectMel Frequency Cepstral Coefficients (MFCC)
dc.subjectClean Audio
dc.titleSpeech Enhancement Using Convolutional Denoising Auto-Encoder
dc.typeThesis

Files

Original bundle

Now showing 1 - 1 of 1
Thumbnail Image
Name:
Full Thesis.pdf
Size:
1.81 MB
Format:
Adobe Portable Document Format

Collections