Beyond Words: Unraveling Text Complexity with Novel Dataset and a Classifier Application

dc.contributor.authorIslam, Mohammad Shariful
dc.contributor.authorRony, Mohammad Abu Tareq
dc.contributor.authorSaha, Pritom
dc.contributor.authorAhammad, Mejbah
dc.contributor.authorAlam, Shah Md Nazmul
dc.contributor.authorRahman, Md Saifur
dc.date.accessioned2024-05-04T06:21:20Z
dc.date.available2024-05-04T06:21:20Z
dc.date.issued2023-02-27
dc.description.abstractText classification is a fundamental aspect of Natural Language Processing (NLP). This research presents a novel human-annotated English sentence dataset categorized into four classes (simple, complex, compound, complex-compound) containing 22331 sentences and a sophisticated sentence classifier tool offering the capability to analyze and classify sentences within English text with particular relevance to literature writing. This study explores its performance using three distinct feature representation methods: Bag-of-Words (BoW), Term Frequency-Inverse Document Frequency (TF-IDF), and Word Embedding Features. The study involves the evaluation of four machine learning and two deep learning classifier models. BoW combined with Support Vector Classifier (SVC) and Logistic Regression (LR) demonstrated impressive accuracy rates, excelling in distinguishing sentence complexity. Word Embedding Features, specifically LSTM and RNN, offer a more profound semantic representation. LSTM stands out with the highest accuracy of 98.03% and balanced precision and recall, yielding an average F1-score of 97%. RNN, slightly less accurate at 97.75%, nevertheless exhibits competence in grasping sentence structure dependencies. It offers valuable insights for practical applications and contributes to the broader understanding of sentence structures and semantics.
dc.identifier.otherhttp://dspace.daffodilvarsity.edu.bd:8080/handle/123456789/12216
dc.identifier.urihttp://dspace.daffodilvarsity.edu.bd:8080/handle/123456789/12216
dc.language.isoen_US
dc.publisherIEEE
dc.sourceDIU Institutional Repository
dc.subjectClassification
dc.subjectDatasets
dc.subjectNatural language
dc.titleBeyond Words: Unraveling Text Complexity with Novel Dataset and a Classifier Application
dc.typeArticle

Files

Original bundle

Now showing 1 - 1 of 1
No Thumbnail Available
Name:
Beyond Words Unraveling Text Complexity with Novel Dataset and A Classifier Application.docx
Size:
14.13 KB
Format:
Adobe Portable Document Format

Collections