2022

Browse

Search Results

Now showing 1 - 10 of 52
  • Thumbnail Image
    Item
    ChartSumm: A large scale benchmark for Chart to Text Summarization
    (Department of Computer Science and Engineering(CSE), Islamic University of Technology(IUT), Board Bazar, Gazipur, Bangladesh, 2022-05-30) Rahman, Raian; Hasan, Rizvi; Farhad, Abdullah Al
    Information visualization such as bar- and line-charts are quite popular for understanding large tabular data. But, interpreting information solely with different visualization techniques can also be difficult due to different reasons like visual impairment or the requirement of prior domain knowledge to understand the chart. Automatic "chart to text summarization" can be promising and effective tool for providing accessibility as well as precised insights of chart data in natural language. In spite of having a good potential, there have not been a lot of works on chart to text summarization making it a low resource task. Scarcity of large scale datasets for chart to text summarization is one of the reason behind this. The human written descriptions in the available dataset also contains information beyond the knowledge of the chart making it difficult for us to have an unbiased evaluation. In our thesis, we propose ChartSumm a large scale dataset for chart to text summarization consisting of 84,363 charts along with their metadata and descriptions. We also propose two test sets: test-e and test-h for evaluating the performance of the trained models available in this domain. Our experiment shows that a T5 model trained on our dataset has achieved BLEU score of 75.72 in test-e set and 64.78 in test-h set. From our analysis we can conclude that large language models like T5 and BART can generate short precised deception from given chart metadata.
  • Thumbnail Image
    Item
    Bioinformatics Analysis of Differentially Expressed Gene's in Breast Cancer Using DESeq2
    (Department of Computer Science and Engineering(CSE), Islamic University of Technology(IUT), Board Bazar, Gazipur, Bangladesh, 2022-05-30) Malick, Sow Bocar Amadou; Conteh, Fatoumatta; Sawo, Muhammed
    Differential Gene Expression Analysis is a strong tool for determining if genes in two or more sample groups are expressed at significantly different levels. To estimate gene counts and identify deferentially expressed genes, we’ll utilize the DESeq2 software. Also, while determining whether genes are deferentially expressed, we must account for variation in the data. The purpose is to see if differences between groups are substantial for each gene, given the biological differences between biological replicates. Using Normalized to Read Count Data (NRCD) and statistical analysis, DEG analysis was used to find quantitative differences in expression levels between experimental groups. For example; statistical testing is used to decide whether for a given gene and observed difference in read counts is significant. I.e., whether it is greater than what would be expected just due to natural random variation. The analysis requires gene expression values to be compared between sample group types. The goal is to determine which genes are expressed at different levels between conditions. It has become a widely used technology that allows for effective genome-wide relative gene expression quantification, and it is the method of choice for identifying deferentially expressed genes between two or more biological situations of interest. The primary challenges surrounding such DE analysis have been highlighted from the start, and several methodologies and tools have been offered in the relevant literature. One of the most difficult aspects of this study, as with any other statistical research, has been determining the probabilistic model that best fits the data, as well as the model’s optimal parameter estimates. Another significant challenge was the requirement for data normalization in order to appropriately compare two biological situations by analyzing and removing any potential technological and/or biological biases. Last but not least, several research have emphasized the practical requirement to determine the ideal number of biological replicates per condition and the optimal library size. We’ll go over the use of DeSeq2 method as a utilized methodology and tools for DE analysis in this article. The gene outcomes can offer biological insights into processes affected by the conditions. greater than what would be expected just due to natural random variation.
  • Thumbnail Image
    Item
    Detection of Lung Adenocarcinoma Cancer based on RNA-seq gene expression data using LIMMA and TabNet
    (Department of Computer Science and Engineering(CSE), Islamic University of Technology(IUT), Board Bazar, Gazipur, Bangladesh, 2022-05-30) Rahman, Faysal Bin; Anjum, Farhan; Khan, Musaddiq Hasan Fatin
    Lung cancer is one of the deadliest diseases of the world to this date with the highest mortality rate amidst all other forms of cancer. Detection of cancer in early stages is crucial for cancer treatment. Progress in cancer detection has been increasingly made based on gene expression levels, giving insight into making correct and successful treatment decisions, thanks to recent advances in high-throughput sequencing technology such as RNA-seq and the use of several machine learning approaches. However, most of the work on cancer detection uses micro-array data and machine learning models. This paper presents a new methodology based on RNA-seq data which is better at detecting transcripts than micro-array along with Deep Neural Network (Tabnet) to classify human lung cancer.
  • Thumbnail Image
    Item
    Blockchain for Electronic Health Record
    (Department of Computer Science and Engineering(CSE), Islamic University of Technology(IUT), Board Bazar, Gazipur, Bangladesh, 2022-05-30) Farah, Radwan Mohamed; Atiku, Ali Umar
    Blockchain may additionally reinvent the technique victim’s electronic fitness statistics are distributed and kept via imparting more secure mechanisms for reallocate peer-to-peer (P2P). That means to help and simplicity the recognize of this give out ledger era, a strong systematic literature overview changed into manage, attending to discover the latest literature on blockchain and care realm and establish current question and open queries, coached way of the enhance of take a look at queries concerning EHR for the duration of a blockchain. Pretty 300 scientific research revealed within the last 10 years have been researched, culminating in the construction of an up-to-date classification, question and unlocked queries recognized, and additionally the foremost important tactics, statistics types, requirements and architectures concerning using blockchain for EHR had been evaluated and referred to.
  • Thumbnail Image
    Item
    Emotion Detection in Online Social networks: Using Deep Learning Approach
    (Department of Computer Science and Engineering(CSE), Islamic University of Technology(IUT), Board Bazar, Gazipur, Bangladesh, 2022-05-30) Khalid, Tarik; Djibrine, Mahamat; Mohamed, Hafso
    Emotion recognition is one of the most difficult jobs in the Natural Language Processing (NLP) sector since it relies significantly on contextual information and mixed emotions in a sentence during the emotion detection process. Therefore, we propose two deep learning approaches CNN and Bi-LSTM, we built these two models on a dataset that contains six levels of emotions. The two models have proven to give good accuracy above 90% on this dataset. From that, we have decided to try them out on thirteen levels of emotions to see if we can still achieve reasonable performance on a high level of emotions
  • Thumbnail Image
    Item
    An Indoor Object Dataset for Mobile-based Detection and Recognition Systems for the Visually Impaired
    (Department of Computer Science and Engineering(CSE), Islamic University of Technology(IUT), Board Bazar, Gazipur, Bangladesh, 2022-05-30) Azad, Shehreen; Sayed, Abdullah Abu; Faiyrooz, Noshin
    Indoor object detection is a challenging area of computer vision where comparatively lesser work has been done compared to its outdoor counterpart. Surely, such a task requires huge amount of training data to make any classifier detect indoor objects with high precision. This indoor object detection can become way more challenging when it has to be specifically tailored for visually impaired people’s mobility and interaction with everyday use objects. This report presents a novel indoor object dataset containing 5196 unique images of 8 everyday use indoor object category relevant to daily interaction of visually impaired people. The uniqueness of this dataset compared to existing indoor objects dataset is this dataset deals with everyday use objects and presents them with more contextual information than that is available in existing literature. Moreover, the varying lighting condition, non-canonical viewpoints, occlusion and complex background makes the dataset more robust while being trained on any object detection algorithm. Instead of going for higher accuracy we aim to find a trade-off between accuracy and speed as if this dataset is to be used to build a system for visually impaired people’s navigation needs, that system has to be deployed on mobile or sensor-based hand-held device which requires lightweight models. Hence our proposed dataset is tested on two light-weight model, namely, SSD MobileNet V2 FPNLite and EfficientDet D0. It has achieved a mean average precision (mAP) of 29.5 and 39.4 respectively on both the models which is better than the original mAP values achieved by these models. Our proposed dataset can be extended with other indoor object detection dataset, as well as it can be used to build a system for visually impaired people’s navigation.
  • Thumbnail Image
    Item
    SDN-based Time Series Traffic Flow Forecasting in VANET
    (Department of Computer Science and Engineering(CSE), Islamic University of Technology(IUT), Board Bazar, Gazipur, Bangladesh, 2022-05-30) Shuvro, Ali Abir; Khan, Mohammad Shian; Rahman, Monzur
    Intelligent Transportation Systems(ITS) provides services for proper traffic assistance. Vehicular Ad-hoc Network(VANET) provides internet connectivity to vehicles and helps in traffic guidance. In this paper, traffic flow prediction is done using a modified transformer architecture for time-series vehicular data. Sequences are generated from the dataset for capturing temporal dependencies. The transformer model has been engineered to capture inter-feature correlations along with inter-sample correlations. Our transformer model has performed much better than other models like LSTM. We also propose a holistic networking model where the vehicles will be connected to Road-side Units(RSUs) and the backbone network will be Software Defined Network(SDN). The traditional design principles, that incorporates data, control and management planes together in a network device, are incapable to adapt with this much data growth, bandwidth, speed, security, scalability compared to SDN as it provides with centralized programmable mechanism reliably. The trained parameters learned using the transformer model will be passed throughout the network for traffic guidance. Similar sized packets are passed using a simulator to demonstrate the time required for the propagation of the parameters.
  • Thumbnail Image
    Item
    A Semi-Automated Approach to Generate Bangla Dataset for Question-Answering and Query-Based Text Summarization
    (Department of Computer Science and Engineering(CSE), Islamic University of Technology(IUT), Board Bazar, Gazipur, Bangladesh, 2022-05-30) Mushabbir, Mueeze Al; Alamgir, Refaat Mohammad; Humdoon, Ahmed Azaz
    With the vast amount of information available on the Internet, finding answers to questions is as important as ever in today’s day and age. In Natural Language Processing Research, Question Answering (QA) and Query-based Text Summarization (QBSUM) are there to tackle this challenge. However, most of the work being done neglects low resource languages such as Bangla, resulting in the small number of quality datasets available in the literature. Therefore to address this research gap, in this work, we propose a semi-automated methodology for generating a Bangla dataset with Natural Questions for three tasks - Question Answering (QA), Query-based Single Document Text Summarization (SD-QBSUM) and Query-based Multi-Document Text Summarization (MD-QBSUM). We then provide baselines for this dataset on those tasks and also compare our dataset with existing ones on various metrics.
  • Thumbnail Image
    Item
    An End to End System for Online Handwritten Bangla Character Recognition
    (Department of Computer Science and Engineering(CSE), Islamic University of Technology(IUT), Board Bazar, Gazipur, Bangladesh, 2022-05-30) Nahin, Shahriar Nur; Imam, Kazi Fahim; Rahman, Nabil; Tasnim, Anika
    This report summarizes the attempt to find the way towards building an Optical Character Recognition System for handwritten Bangla characters. The complex and unique structure of scripts like Bangla and ever challenging nature of hand- written texts combined makes it really difficult to complete a perfect system to approach to convert the scanned handwritten Bangla scripts to machine editable digital counterpart format of it- as segmentation of the whole image into char- acters and then classification of the segmented characters is difficult enough to make the task challenging. In our work, we propose to approach the segmentation process (directly segment to words) with Distance Transform and morphological operations for error correction later. Then two zone approach (either side of matra- upper and lower zone) and apply connected component analysis on both zones. We handled or adjusted the failed and not directly successful cases by experimenting with the characteristics of handwritten characters. Then for clas- sification process, we proposed to classify the segmented characters using neural networks trained on the relatively newly available datasets. Multiple column, Mixed characters (Bangla- other languages) and Scene Text Recognition is out of the scope of our study so far. And we could not include the post-processing part for our work for lack of work or mention in existing literature, which might be a great addition in the way of building a complete OCR system.
  • Thumbnail Image
    Item
    Nuclei Instance Segmentation of Cryosectioned H&E Stained Histological Images using Deep Learning
    (Department of Computer Science and Engineering(CSE), Islamic University of Technology(IUT), Board Bazar, Gazipur, Bangladesh, 2022-04-30) Ahmed, Zarif; Siddiqi, Chowdhury Nur e Alam; Alam, Fardifa Fathmiul
    Nuclei instance segmentation is an important step for oncological diagnosis and pathology research of cancer. HE stained images are considered the gold standard for medical diagnosis. But before being used for segmentation, it is required to pre process them. There are two principle methods to preprocess them formalin-fixed paraffin-embedded samples (FFPE) and frozen tissue samples (FS). Even though FFPE is widely used, it is a time consuming process whereas FS samples can be processed very quickly. But analysis of FS-derived HE stained images can be more difficult as rapid preparation, staining, and scanning of FS sections results in degradation of image quality. Therefore, in this thesis, we explored various state of the art segmentation architectures to create a model that will segment nuclei of FS-derived HE stained images with a high quality feature extraction. Here, we have been working on a novel dataset called CryoNuSeg that contains 30 FS-sectioned images of 10 human organs. It has a benchline score of DICE 80.3 ±4.3, AJI 52.5 5.0, PQ 47.7 6.1. U-Net is the first and most prominent architecture for biomedical image segmentation. We are exploring various U-Net architectures. We have trained Triple U-NET on the dataset using binary masks in place of U-NET keeping all other parts of the instance segmentation algorithm same such as Gaussian Filtering and Watershed Post processing. The results using Triple U-NET crossed all the benchline scores. The triple U-Net architecture gives a score of DICE 80.33, AJI 67.41 and PQ 50.56. We have developed a deep learning model that performs highly accurate nuclei segmentation of FS sections despite degraded image quality for fast oncological diagnosis.