Bangla speech to text conversion using CMU sphinx
| dc.contributor.advisor | Arif, Hossain | |
| dc.contributor.author | Bristy, Israt Jerin | |
| dc.contributor.author | Shakil, Nadim Imtiaz | |
| dc.contributor.author | Musavee, Tesnim | |
| dc.contributor.author | Choton, Akibur Rahman | |
| dc.date.accessioned | 2020-01-20T04:24:59Z | |
| dc.date.available | 2020-01-20T04:24:59Z | |
| dc.date.issued | 2019-08 | |
| dc.description | Cataloged from PDF version of thesis. | |
| dc.description | Includes bibliographical references (pages 30-32). | |
| dc.description | This thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science, 2019. | |
| dc.description.abstract | Speech is the most normal type of communication and association between people while content (text) and images are the most basic types of exchange in the computer system. Therefore, enthusiasm in regards to transformation between speech and text is expanding day by day for integrating the human-computer relation. Understanding speech for a human is not a challenge but for a machine it is a big deal because a machine does not catch expression or human nature. For the conversion of speech into text, this proposed model requires the usage of the open sourced framework Sphinx 4 which is written in Java. For the proposed system, it requires certain steps which are training an acoustic model, creating a language model and building a dictionary with CMUSphinx. For training, the audio files were recorded by 8 speakers both male and female for more accuracy. Among them, 6 speakers recorded each word 3 times. To test the accuracy, we took audio recordings from 2 speakers among them one speaker is unknown to the system. After testing, we got the accuracy around 59.01%. For known speakers we got 78.57% accuracy. We gave audio files as input only to check accuracy as our main purpose was to make a system which works in real time. In our system, user can speak in real time and the system converts it into text. | |
| dc.identifier.other | ID 15301006 | |
| dc.identifier.other | ID 15301037 | |
| dc.identifier.other | ID 15101110 | |
| dc.identifier.other | ID 15301102 | |
| dc.identifier.other | https://dspace.bracu.ac.bd/server/api/core/items/2d37af9d-b96e-4788-9626-c51b7eb29e3b | |
| dc.identifier.uri | http://hdl.handle.net/10361/13632 | |
| dc.language.iso | en | |
| dc.publisher | BRAC University | |
| dc.source | BRAC University Institutional Repository | |
| dc.subject | Bangla | |
| dc.subject | Voice recognition | |
| dc.subject | CMUSphinx | |
| dc.subject | Acoustic model | |
| dc.subject | Language model | |
| dc.title | Bangla speech to text conversion using CMU sphinx | |
| dc.type | Thesis |
Files
Original bundle
1 - 1 of 1
- Name:
- 15301006,15301037,15101110,15301102_CSE.pdf
- Size:
- 579.49 KB
- Format:
- Adobe Portable Document Format
