Repository logo
Communities & Collections
All of DSpace
  • English
  • العربية
  • বাংলা
  • Català
  • Čeština
  • Deutsch
  • Ελληνικά
  • Español
  • Suomi
  • Français
  • Gàidhlig
  • हिंदी
  • Magyar
  • Italiano
  • Қазақ
  • Latviešu
  • Nederlands
  • Polski
  • Português
  • Português do Brasil
  • Srpski (lat)
  • Српски
  • Svenska
  • Türkçe
  • Yкраї́нська
  • Tiếng Việt
Log In
New user? Click here to register.Have you forgotten your password?
  1. Home
  2. Browse by Author

Browsing by Author "Parves, Abdul Bari"

Filter results by typing the first few letters
Now showing 1 - 5 of 5
  • Results Per Page
  • Sort Options
  • Thumbnail Image
    Item
    Bangla Language Mode (Sadhu/Cholito) Classification
    (Daffodil International University, 2019-05-01) Parves, Abdul Bari; Rakib, Emranul Haque
    This project addresses the problem of distinguishing between two form of Bangla language, namely Sadhubhasha and Cholitobhasha. The classifier would be beneficial for finding the right word choice for Bangla literature. The main vision of this project is to different the modern era’s early Bangla form of Sadhubhasha to the current form of Cholitobhasha. As far as we know there has been no single work done addressing this particular issue. From another perspective, only a few works have been done on “Bangla Language”. So, it has been difficult to conduct advance linguistic works on Bangla language like extracting information or summarizing. We had to face difficulties when collecting Bangla data due to the limited availability, but finally we have collected around total 100000 words dataset for this project. Among which 80% of the data is used for training and rest 20% is test data. Machine learning algorithms Random forest, Naïve Bayes, Support Vector Machine, Knearest neighbor and Decision tree are applied to classify the language and the Term Frequency-Inverse Document Frequency and Bag of Words are used for the numerical representation. With these classifiers 91% to 99.5% accuracy is observed. The promising outcome of this project is, "sadhu and cholito Language classifier" can be used as the first step on that ladder from where others will be influenced to do further research on Bangla language.
  • Thumbnail Image
    Item
    Bangla language mode (Sadhu/Cholito) Classification
    (Daffodil International University, 2019-05-05) Parves, Abdul Bari; Rakib, Emranul Haque
    This project addresses the problem of distinguishing between two form of Bangla language, namely Sadhubhasha and Cholitobhasha. The classifier would be beneficial for finding the right word choice for Bangla literature. The main vision of this project is to different the modern era’s early Bangla form of Sadhubhasha to the current form of Cholitobhasha. As far as we know there has been no single work done addressing this particular issue. From another perspective, only a few works have been done on “Bangla Language”. So, it has been difficult to conduct advance linguistic works on Bangla language like extracting information or summarizing. We had to face difficulties when collecting Bangla data due to the limited availability, but finally we have collected around total 100000 words dataset for this project. Among which 80% of the data is used for training and rest 20% is test data. Machine learning algorithms Random forest, Naïve Bayes, Support Vector Machine, Knearest neighbor and Decision tree are applied to classify the language and the Term Frequency-Inverse Document Frequency and Bag of Words are used for the numerical representation. With these classifiers 91% to 99.5% accuracy is observed. The promising outcome of this project is, "sadhu and cholito Language classifier" can be used as the first step on that ladder from where others will be influenced to do further research on Bangla language.
  • No Thumbnail Available
    Item
    Incorporating Supervised Learning Algorithms with NLP Techniques to Classify Bengali Language Forms
    (Scopus, 2020-01-10) Parves, Abdul Bari; Imran, Abdullah Al; Rahman, Md. Riazur
    Every language has its own root, form, and grammar, and so does Bengali. Bengali language has two core forms: "Sadhu-bhasha" and "Cholito-bhasha" which have been widely used from regular communication to literary publications. At present, Sadhu-bhasha can be only found in old books and literary publications, whereas Cholito-bhasha is mostly used everywhere. However, so many Bengali linguists are still researching on these two forms to preserve its root, understand and develop Bengali, and also extract knowledge from the historical publications which were mainly written in Sadhu-bhasha. Unfortunately, till now they do not have any digital tool that can assist their research by automatically identifying these core forms of Bengali from the large archive of Bengali literature. This study aims to build such an automatic intelligent system that can accurately identify these two language forms by harnessing the power of Natural Language Processing (NLP). In this study, we have applied advanced NLP techniques and six Supervised learning algorithms to classify "Sadhu-bhasha" and "Cholito-bhasha" from text corpora. Results of this study show that all the six models yielded very promising results, however, the Multinomial Naive Bayes outperformed all the models with 99.5% accuracy, 99.0% precision, 100% recall, 0.995 AUC score and, 0.995 F1 score. Additionally, this study also performs qualitative analysis using t-SNE algorithm to visualize the difference between Sadhu-bhasha and Cholito-bhasha.
  • No Thumbnail Available
    Item
    Incorporating Supervised Learning Algorithms with NlP Techniques to Classify Bengali Language Forms
    (ACM International Conference Proceeding Series, 2020-01-10) Parves, Abdul Bari; Imran, Abdullah Al; Rahman, Md. Riazur
    Every language has its own root, form, and grammar, and so does Bengali. Bengali language has two core forms: "Sadhu-bhasha" and "Cholito-bhasha" which have been widely used from regular communication to literary publications. At present, Sadhu-bhasha can be only found in old books and literary publications, whereas Cholito-bhasha is mostly used everywhere. However, so many Bengali linguists are still researching on these two forms to preserve its root, understand and develop Bengali, and also extract knowledge from the historical publications which were mainly written in Sadhu-bhasha. Unfortunately, till now they do not have any digital tool that can assist their research by automatically identifying these core forms of Bengali from the large archive of Bengali literature. This study aims to build such an automatic intelligent system that can accurately identify these two language forms by harnessing the power of Natural Language Processing (NLP). In this study, we have applied advanced NLP techniques and six Supervised learning algorithms to classify "Sadhu-bhasha" and "Cholito-bhasha" from text corpora. Results of this study show that all the six models yielded very promising results, however, the Multinomial Naive Bayes outperformed all the models with 99.5% accuracy, 99.0% precision, 100% recall, 0.995 AUC score and, 0.995 F1 score. Additionally, this study also performs qualitative analysis using t-SNE algorithm to visualize the difference between Sadhu-bhasha and Cholito-bhasha.
  • No Thumbnail Available
    Item
    Incorporating Supervised Learning Algorithms with NLP Techniques to Classify Bengali Language Forms
    (ACM International Conference Proceeding Series, 2020-01-10) Parves, Abdul Bari; Imran, Abdullah Al; Rahman, Md. Riazur
    Every language has its own root, form, and grammar, and so does Bengali. Bengali language has two core forms: "Sadhu-bhasha" and "Cholito-bhasha" which have been widely used from regular communication to literary publications. At present, Sadhu-bhasha can be only found in old books and literary publications, whereas Cholito-bhasha is mostly used everywhere. However, so many Bengali linguists are still researching on these two forms to preserve its root, understand and develop Bengali, and also extract knowledge from the historical publications which were mainly written in Sadhu-bhasha. Unfortunately, till now they do not have any digital tool that can assist their research by automatically identifying these core forms of Bengali from the large archive of Bengali literature. This study aims to build such an automatic intelligent system that can accurately identify these two language forms by harnessing the power of Natural Language Processing (NLP). In this study, we have applied advanced NLP techniques and six Supervised learning algorithms to classify "Sadhu-bhasha" and "Cholito-bhasha" from text corpora. Results of this study show that all the six models yielded very promising results, however, the Multinomial Naive Bayes outperformed all the models with 99.5% accuracy, 99.0% precision, 100% recall, 0.995 AUC score and, 0.995 F1 score. Additionally, this study also performs qualitative analysis using t-SNE algorithm to visualize the difference between Sadhu-bhasha and Cholito-bhasha.

© Open Research Bangladesh

  • Privacy policy
  • End User Agreement
  • Send Feedback