Repository logo
Communities & Collections
All of DSpace
  • English
  • العربية
  • বাংলা
  • Català
  • Čeština
  • Deutsch
  • Ελληνικά
  • Español
  • Suomi
  • Français
  • Gàidhlig
  • हिंदी
  • Magyar
  • Italiano
  • Қазақ
  • Latviešu
  • Nederlands
  • Polski
  • Português
  • Português do Brasil
  • Srpski (lat)
  • Српски
  • Svenska
  • Türkçe
  • Yкраї́нська
  • Tiếng Việt
Log In
New user? Click here to register.Have you forgotten your password?
  1. Home
  2. Browse by Author

Browsing by Author "Abujar, Sheikh"

Filter results by typing the first few letters
Now showing 1 - 20 of 65
  • Results Per Page
  • Sort Options
  • No Thumbnail Available
    Item
    A Bengali Text Generation Approach in Context of Abstractive Text Summarization Using RNN
    (Lecture Notes in Networks and Systems, Springer, 2020-03-04) Abujar, Sheikh; Masum, Abu Kaisar Mohammad; Islam, Md. Sanzidul; Faisal, Fahad; Hossain, Syed Akhter
    Automatic text summarization is one of the mentionable research areas of natural language processing. The amount of data is increasing rapidly, and the necessity of understanding the gist of any text is just a mandatory tool, nowadays. The area of text summarization has been developing since many years. Mentionable research has been already done through extractive summarization approach; in other side, abstractive summarization approach is the way to summarize any text as like human. Machine will be able to provide a new type of summarization, where the understanding of given summary may found as like as human-generated summary. Several research developments have already been done for abstractive summarization in English language. This paper shows a necessary method—“text generation” in context of Bengali abstractive text summarization development. Text generation helps the machine to understand the pattern of human-written text and then produce the output as is human-written text. A basis recurrent neural network (RNN) has been applied for this text generation approach. The most applicable and successful RNN—long short-term memory (LSTM)—has been applied. Contextual tokens have been used for the better sequence prediction. The proposed method has been developed in the context of making it useable for further development of abstractive text summarization.
  • No Thumbnail Available
    Item
    A Continuous Word Segmentation of Bengali Noisy Speech
    (Springer, 2020-11-28) Hossain, Md. Fahad; Hasan, Md. Mehedi; Ali, Hasmot; Abujar, Sheikh
    Human voice is an important concern of efficient and modern communication in the era of Alexa, Siri, or Google Assistance. Working with voice or speech is going to be easy by preprocessing the unwanted entities when real speech data contains a lot of noise or continuous delivery of a speech. Working with Bangla language is also a concern of enriching the scope of efficient communication over Bangla language. This paper presented a method to reduce noise from speech data collected from a random noisy place, and segmentation of word from continuous Bangla voice. By filtering the threshold of noise with fast Fourier transform (FFT) of audio frequency signal for reduction of noise and compared each chunk of audio signal with minimum dBFS value to separate silent period and non-silent period and on each silent period, segment the signal for word segmentation.
  • No Thumbnail Available
    Item
    A Continuous Word Segmentation of Bengali Noisy Speech
    (Scopus, 2021) Hossain, Md. Fahad; Hasan, Md. Mehedi; Ali, Hasmot; Abujar, Sheikh
    Human voice is an important concern of efficient and modern communication in the era of Alexa, Siri, or Google Assistance. Working with voice or speech is going to be easy by preprocessing the unwanted entities when real speech data contains a lot of noise or continuous delivery of a speech. Working with Bangla language is also a concern of enriching the scope of efficient communication over Bangla language. This paper presented a method to reduce noise from speech data collected from a random noisy place, and segmentation of word from continuous Bangla voice. By filtering the threshold of noise with fast Fourier transform (FFT) of audio frequency signal for reduction of noise and compared each chunk of audio signal with minimum dBFS value to separate silent period and non-silent period and on each silent period, segment the signal for word segmentation.
  • No Thumbnail Available
    Item
    A heuristic approach of text summarization for Bengali documentation
    (IEEE Xplore, 2017-12-14) Abujar, Sheikh; Hasan, Mahmudul; Shahin, M.S.I; Hossain, Syed Akhter
    Automated Text Summarization is a technique of summarizing any document or text automatically. Summarized text is the concise form of the given text. In Natural language processing many text summarization techniques are available for English language, but only a few for Bangla language. Bangla is one of the most taught and used language all over the world. Most of the text summarization techniques are implemented in two different ways, known as abstractive or extractive approach. This paper deal with the summarization of Bangla text based on extractive method. A new efficient extractive summarization method is proposed in this work. The other summarization tools developed for Bangla language seems not much appropriate from application point of view. The proposed analysis models are applicable for Bangla text summarization. In the proposed approach, basic extractive summarization is applied with new proposed model and a set of Bangla text analysis rules derived from the heuristics. Every Bangla sentences and words from original text is analyzed properly with Bangla sentence clustering method. This work proposed a new type of sentence scoring processes for Bangla text summarization. In the evaluation of this technique, the system reflects good accuracy of results, comparing to that of the human generated summarized result and other Bangla text summarization tools. Full Text Link: http://doi.org/10.1109/ICCCNT.2017.8204166
  • No Thumbnail Available
    Item
    A Knowledge Base Data Mining Based on Parkinson's Disease
    (IEEE, 2019-11) Hassan, Md. Redone; Kadir, S.K. Obidul; Islam, Md. Aminul; Abujar, Sheikh; Zannat, Raihana; Ohidujjaman
    The approaches to detecting Parkinson's disease in the human body from voice data by using Classification techniques apply three different algorithms for finding the growth rate of this disease. Unified Parkinson's disease rating scale deals with motor fluctuations and changes over voice after a certain period and that can measure the people affected by this disease and the difference with healthy people. Hoehn & Yahr scale measures the symptoms which are being working through the improvement of Parkinson's disease in the human body. Classifier algorithms used to detect the factors and symptoms which are involved in the advancement of this disease in the human body using voice data. From the distinctions of all algorithms measures the growth rate and find out which algorithm gives the best result for several approaches to diagnosis Parkinson's disease and chances of had this disease in the human body.
  • No Thumbnail Available
    Item
    A Potent Model to Recognize Bangla Sign Language Digits Using Convolutional Neural Network
    (Elsevier B.V., 2018-11-19) (Md.), Sanzidul Islam; Mousumi, Sadia Sultana Sharmin; Rabby, AKM Shahariar Azad; Hossain, Sayed Akhter; Abujar, Sheikh
    Hearing impaired people have own language called Sign Language but it is difficult for understanding to general people. Sign language is the basic method of communication for deaf people during their everyday of life. Sign digits are also a major part of sign language. So machine translator is necessary to allow them to communicate with general people. For making their language understandable to general people, computer vision based solutions are well known nowadays. In this research work we aim at constructing a model in deep learning approach to recognize Bangla Sign Language (BdSL) digits. In this approach there used Convolutional Neural Network (CNN) to train particular signs with a respective training dataset (Eshara-Lipi) for acquiring our aim. The model trained and tested with respectively 860 training images and 215 (20%) test images of tent classes of digits. Finally, the training model gained about 95% accuracy at recognition of Bangla sign language digits. This model will contribute for moving one step forward to make BdSL machine translator.
  • No Thumbnail Available
    Item
    A Systematic Way of Collecting Data of Insomniac Patients
    (IEEE, 2020-07) Islam, Md. Muhaiminul; Masum, Abu Kaisar Mohammad; Abujar, Sheikh; Hossain, Syed Akhter
    Insomnia (a sleeping disorder) can also be defined as sleeplessness that means facing trouble in falling or staying asleep. In this 20 th century, this disorder is very common among the olds and also teenagers. This disorder can harm vastly in our physical and mental health. So this is a serious fact in medical science. But these patients are rarely been hospitalized. Doctors usually predict this disorder considering the symptoms of patients. They use a questionnaire for it. Sleeping status, mental and physical conditions of patients both include in it. From the early days, it was a challenge for researchers to collect the data of these patients and analyze it. And the collection of these kinds of data is also a very difficult task to do. As this disorder expresses the mental and physical condition of a patient, one feels very uncomfortable to share their personal information with others. So there is no standard dataset available related to Insomnia. We have decided to collect the data of insomniac people for further research in the medical field. And an intelligent method is developed for the collection of data. Our dataset is preserved in such a way that machine learning algorithms can classify them in categories. Therefore the main objective of this study is to clarify the symptoms of this disorder and make a large dataset of it so that researchers can easily analyze the data and can make the appropriate output from their researches.
  • No Thumbnail Available
    Item
    A Universal Way to Collect and Process Handwritten Data for Any Language
    (Elsevier B.V., 2018-11-19) Rabby, AKM Shahariar Azad; Haque, Sadeka; Shahinoor, Shammi Akther; Abujar, Sheikh; Hossain, Syed Akhter
    In recent years researches based on Machine learning and Deep learning have achieved much interest and one of its handwritten recognition. Handwritten recognition is very difficult due to its lack of dataset and also for collecting data from people. This research introduces a fast and comprehensive way to collect and process handwritten data to develop a way of Handwritten Recognition (HWR) algorithm for any languages. In this research handwritten characters wrote on a paper and then scanned to get the data into a JPEG format. We also focused on some of the other issues and requirements while collecting handwritten data, creating form, data collection methodology, process, using software and relevant tools. We described these issues in the context of our own effort to create a handwritten database for the Bangla language. Our designed Graphical User Interface (GUI) is also able to process 100 scanned images per minute where each scanned image contains 120 characters.
  • No Thumbnail Available
    Item
    Abstractive Method of Text Summarization with Sequence to Sequence RNNs
    (Scopus, 2019-07-08) Masum, Abu Kaisar Mohammad; Rabby, AKM Shahariar Azad; Talukder, Md. Ashraful Islam; Abujar, Sheikh; Hossain, Syed Akhter
    Text summarization is one of the famous problems in natural language processing and deep learning in recent years. Generally, text summarization contains a short note on a large text document. Our main purpose is to create a short, fluent and understandable abstractive summary of a text document. For making a good summarizer we have used amazon fine food reviews dataset, which is available on Kaggle. We have used reviews text descriptions as our input data, and generated a simple summary of that review descriptions as our output. To assist produce some extensive summary, we have used a bi-directional RNN with LSTM's in encoding layer and attention model in decoding layer. And we applied the sequence to sequence model to generate a short summary of food descriptions. There are some challenges when we working with abstractive text summarizer such as text processing, vocabulary counting, missing word counting, word embedding, the efficiency of the model or reduce value of loss and response machine fluent summary. In this paper, the main goal was increased the efficiency and reduce train loss of sequence to sequence model for making a better abstractive text summarizer. In our experiment, we've successfully reduced the training loss with a value of 0.036 and our abstractive text summarizer able to create a short summary of English to English text.
  • No Thumbnail Available
    Item
    An Approach for Bengali Automatic Question Answering System using Attention Mechanism
    (IEEE, 2020-07) Bhuiyan, Md. Rafiuzzaman; Masum, Abu Kaisar Mohammad; Abdullahil-Oaphy, Md.; Hossain, Syed Akhter; Abujar, Sheikh
    Question answering is a set of tools for obtaining detailed answers from user questions. At present, it is gaining very popularity day by day in the area of NLP research. There is a lot of work done in English. Still, become the seventh spoken language has not notable development at all. In Bengali very little work we've seen so far. Many types of problems can be solved by answering questions. Automatic question answering system is very much needed to solve various problems through Q&A. It is a very challenging task to create this type of system. In our paper we developed an automatic context based Question&Answering system using sequence to sequence architecture. An encoder layer will be used with a bi-directional LSTM and a decoder layer followed by an attention mechanism. The main challenge of this work is - data collection, finding the right vocabulary for word mapping and lots more. The main function of our model is to answer the questions. We have been able to successfully answer the question and reduce our training loss to 0.003.
  • No Thumbnail Available
    Item
    An Approach for Bengali Text Summarization using Word2Vector
    (Scopus, 2019-12-30) Abujar, Sheikh; Masum, Abu Kaisar Mohammad; Mohibullah, Md.; Ohidujjaman; Hossain, Syed Akhter
    Text Summarization is one of the mentionable research areas of Natural language processing. Several approaches have already been developed in this concern. Such as - Abstractive approach and extractive approach. Most recent recurrent neural network methods are producing much better results. Several mentionable research has already been discussed for English language summarizer, but a few have already done for the Bengali language. There are so many prerequisites for data analysis purpose-word2vector is one of them. Understanding the vector representation of any text leads the way to identify the key main points of that specific text and helps to measure the relationship of that text with other texts in similarity/dissimilarity [11]. Generated matrix using word2vector can easily applicable for identifying top-ranked sentence/words, either domain specific or in general form. In this paper, a word2vector approach has been discussed in the context of text summarization for the Bengali language.
  • Thumbnail Image
    Item
    An Insight Into the Intricacies of Lingual Paraphrasing Pragmatic Discourse on the Purpose of Synonyms
    (Daffodil International University, 22-10-08) Nahian, Jabir Al; Masum, Abu Kaisar Mohammad; Syed, Muntaser Mansur; Abujar, Sheikh
    The term "paraphrasing" refers to the process of presenting the sense of an input text in a new way while preserving fluency. Scientific research distribution is gaining traction, allowing both rookie and experienced scientists to participate in their respective fields. As a result, there is now a massive demand for paraphrase tools that may efficiently and effectively assist scientists in modifying statements in order to avoid plagiarism. natural language processing (NLP) is very much important in the realm of the process of document paraphrasing. We analyze and discuss existing studies on paraphrasing in the English language in this paper. Finally, we develop an algorithm to paraphrase any text document or paragraphs using WordNet and natural language tool kit (NLTK) and maintain "Using Synonyms" techniques to achieve our result. For 250 paragraphs, our algorithm achieved a paraphrase accuracy of 94.8%.
  • No Thumbnail Available
    Item
    Analysis of Bangladeshi People's Emotion during Covid-19 in Social Media Using Deep Learning
    (IEEE, 2020-07) Pran, Md. Sabbir Alam; Bhuiyan, Md. Rafiuzzaman; Hossain, Syed Akhter; Abujar, Sheikh
    World is passing through a very uncertain circumstance as Coronavirus becoming a great threat. Staying in home is the best solution now to be safe. People are now passing their most of the time in social platform. They're reacting in public posts, news, articles and also commenting there. And a persons comment can talk about his sentiment. Emotion exploration is a very famous topic in the field of data mining. Lots of work have been done yet. In this piece of research, Bangladeshi people's comments on several Facebook news post related to coronavirus have been analyzed to observe the sentiment of them toward this situation. Using three classes investigation have been done on their emotions. which are Analytical, Depressed, Angry. The data set was developed in Bangla language. Several deep learning algorithms have been applied and found the maximum accuracy in CNN 97.24% and in LSTM 95.33%. Result shows that most people commented analytically. The outcome draw up the public psychology of Bangladesh toward the pandemic.
  • Thumbnail Image
    Item
    BAAD: A Multipurpose Dataset for Automatic Bangla Offensive Speech Recognition
    (Elsevier, 2023-03-24) Hossain, Md. Fahad; Supto, Md. Al Abid; Chowdhury, Zannat; Chowdhury, Hana Sultan; Abujar, Sheikh
    In spite of being the fifth most spoken native language in the world, Bangla has barely received any attention in the domain of audio and speech recognition. This article represents a speech dataset of Bengali Abusive Words with some non-abusive wors which are very close to the abusive ones. In this work, a multipurpose dataset is presented to recognize automatic slang speech for Bangla language, which was prepared by collection, annotation, and refinement of data. It consists of 114 slang words and 43 non-slang words with 6100 audio clips. For the collection of slang words, 60 native speakers and for non-abusive words, 23 native speakers participated who were, speaking in various dialects from over 20 districts of Bangladesh, and 10 university students participated to evaluate this dataset including annotation and refinements. Researchers can use this dataset to develop an automatic Bengali Slang speech recognition system, and also it can be used as a new benchmark for creating speech recognition-based machine learning models. This dataset can be enrich-ed further, and some background noise in the dataset can be used to simulate a more real-world scenario if desired. Otherwise, these noises could also be removed.
  • No Thumbnail Available
    Item
    Bangla Continuous Handwriting Character and Digit Recognition Using CNN
    (Springer, 2020-03-04) Hasan, Fuad; Shuvo, Shifat Nayme; Abujar, Sheikh; Mohibullah, Md.; Hossain, Syed Akhter
    There are several works in Bangla handwritten character recognition. Here a new methodology proposed to recognize the character from continuous Bangla handwritten character. The system’s main components are preprocessing, feature extraction, and recognition. There is a strong possibility that is found in Bangla words, and characters are overlapped. This problem often happens in handwritten texts like a consecutive character appears on another character. When it comes to Bangla characters, segmentation becomes much more difficult. To build an effective OCR system of Bangla handwritten text, recognition of characters is important as much as segmentation of characters. Here the main purpose is creating a system, which takes continuous Bangla handwritten text images as an input and then segments the input texts into its constituent words and finally segments each word into individual characters. In this present study, here we used EkushNet dataset model which includes 50 basic characters, 10 character modifiers, 52 frequently used conjunct characters, and 10 digits. By using our algorithm, we are able to segment 95% words from text and 90% characters from the words. Overall, in this present OCR system here recognition and segmentation of characters from handwritten Bangla texts are effectively dealing with the probable problems.
  • No Thumbnail Available
    Item
    Bangla Handwritten Digit Recognition Using Convolutional Neural Network
    (Advances in Intelligent Systems and Computing, Springer, 2018-12-12) Rabby, AKM Shahariar Azad; Abujar, Sheikh; Haque, Sadeka; Hossain, Syed Akhter
    Handwritten digit recognition has always a big challenge due to its variation of shape, size, and writing style. Accurate handwritten recognition is becoming more thoughtful to the researchers for its educational and economic values. There had several works been already done on the Bangla Handwritten Recognition, but still there is no robust model developed yet. Therefore, this paper states development and implementation of a lightweight CNN model for classifying Bangla Handwriting Digits. The proposed model outperforms any previous implemented method with fewer epochs and faster execution time. This Model was trained and tested with ISI handwritten character database Bhattacharya and Chaudhuri (IEEE Trans Pattern Anal Mach Intell 31:444–457, 2009, [1], BanglaLekha Isolated Biswas et al. (Data Brief 12, 103–107, 2017, [2]) and CAMTERDB 3.1.1 Sarkar et al. (Int J Doc Anal Recogn (IJDAR) 15(1):71–83, 2012, [3]). As a result, it was successfully achieved validation accuracy of 99.74% on ISI handwritten character database, 98.93% on BanglaLekha Isolated, 99.42% on CAMTERDB 3.1.1 dataset and lastly 99.43% on a mixed (combination of BanglaLekha Isolated, CAMTERDB 3.1.1 and ISI handwritten character dataset) dataset. This model achieved the best performance on different datasets and found very lightweight, it can be used on a low processing device like-mobile phone.The pre-train model and code for all these datasets can be found on this link https://github.com/shahariarrabby/Bangla_Digit_Recognition_CNN.
  • No Thumbnail Available
    Item
    Bangla Speaker Accent Variation Detection by MFCC Using Recurrent Neural Network Algorithm
    (Springer, 2020-03-04) Mamun, Rezaul Karim; Abujar, Sheikh; Islam, Rakibul; Been Md. Badruzzaman, Khalid; Hasan, Mehedi
    There are a number of languages accent differential applications that detect the different accents in assorted languages. The studies which have done before most of them are based on the English language and different languages throughout the world. A few researches have been performed in Bangla regional language accent differential applications, which is not conclusive for the system to be able to manage Bangla accented speakers. In this paper, we report regional language accent detection experiments of different types of Bangladesh. We demonstrate a strategy to observe Bangladeshi different accents which exploit Mel frequency cepstral coefficient (MFCC) and recurrent neural network (RNN). Listening from the people of different places in Bangladesh creates an accent differentiation results performed by the speakers. This experimental result shows the adaptation of the people to adapt of the regional languages.
  • No Thumbnail Available
    Item
    Bengali Abstractive Text Summarization Using Sequence to Sequence RNNs
    (10th International Conference on Computing, Communication and Networking Technologies, IEEE, 2019-07-08) Talukder, Md Ashraful Islam; Abujar, Sheikh; Masum, Abu Kaisar Mohammad; Faisal, Fahad; Hossain, Syed Akhter
    Text summarization is one of the leading problem of natural language processing and deep learning in recent years. Text summarization contains a condensed short note on a large text document. Our purpose is to create an efficient and effective abstractive Bengali text summarizer what can generate an understandable and meaningful summary from a given Bengali text document. To do this we have collected various texts such as newspaper articles, Facebook posts etc. and to generate summary from those text we will be using our model. Our model works with bi-directional RNNs with LSTM in encoding layer and attention model at decoding layer. Our model works as sequence to sequence model to generate summary. There are some challenges we have faced while building this model such as text pre-processing, vocabulary counting, missing words counting, word embedding, unknown words find out and so on. In this model, our main goal was to make an abstractive summarizer and reduce the train loss of that. During our research experiment, we have successfully reduced the train loss to 0.008 and able to generate a fluent short summary note from a given text.
  • No Thumbnail Available
    Item
    Bengali Accent Classification from Speech Using Different Machine Learning and Deep Learning Techniques
    (Scopus, 2021) Badhon, S. M. Saiful Islam; Rahaman, Habibur; Rupon, Farea Rehnuma; Abujar, Sheikh
    The work starts with a question “Does human vocal folds produce different wavelength when they speak in different accent of same language?” Generally, when humans hear the language, they can easily classify the accent and region from the language. But the challenge was how we give this capability to the machine. By calculating discrete Fourier transform, Mel-spaced filter-bank and log filter-bank energies, we got Mel-frequency cepstral coefficients (MFCCs) of a voice which is the numeric representation of an analog signal. And then, we used different machine learning and deep learning algorithms to find the best possible accuracy. By detecting the region of speaker from voice, we can help security agencies and e-commerce marketing. Working with human natural language is a part of Natural Language Processing (NLP) which is branch of artificial intelligence. For feature extraction, we used MFCCs, and for classification, we used linear regression, decision tree, gradient boosting, random forest and neural network. And we got max 86% accuracy on 9303 data. The data was collected from eight different regions (Dhaka, Khulna, Barisal, Rajshahi, Sylhet, Chittagong, Mymensingh and Noakhali) of Bangladesh. We follow a simple workflow for getting the ultimate result.
  • No Thumbnail Available
    Item
    Bengali Accent Classification from Speech Using Different Machine Learning and Deep Learning Techniques
    (Scopus, 2021) Badhon, S. M. Saiful Islam; Rahaman, Habibur; Rupon, Farea Rehnuma; Abujar, Sheikh
    The work starts with a question “Does human vocal folds produce different wavelength when they speak in different accent of same language?” Generally, when humans hear the language, they can easily classify the accent and region from the language. But the challenge was how we give this capability to the machine. By calculating discrete Fourier transform, Mel-spaced filter-bank and log filter-bank energies, we got Mel-frequency cepstral coefficients (MFCCs) of a voice which is the numeric representation of an analog signal. And then, we used different machine learning and deep learning algorithms to find the best possible accuracy. By detecting the region of speaker from voice, we can help security agencies and e-commerce marketing. Working with human natural language is a part of Natural Language Processing (NLP) which is branch of artificial intelligence. For feature extraction, we used MFCCs, and for classification, we used linear regression, decision tree, gradient boosting, random forest and neural network. And we got max 86% accuracy on 9303 data. The data was collected from eight different regions (Dhaka, Khulna, Barisal, Rajshahi, Sylhet, Chittagong, Mymensingh and Noakhali) of Bangladesh. We follow a simple workflow for getting the ultimate result.
  • «
  • 1 (current)
  • 2
  • 3
  • 4
  • »

© Open Research Bangladesh

  • Privacy policy
  • End User Agreement
  • Send Feedback