Browsing by Author "Alam, Firoj"
Now showing 1 - 15 of 15
- Results Per Page
- Sort Options
Item Acoustic analysis of Bangla consonants(BRAC University, 2008) Alam, Firoj; Habib, S. M. Murtoza; Khan, MumitThis paper describes the acoustic characteristics of Bangla consonants, obtained by analyzing the recordings of male and female voices. First, the duration of each phoneme was identified by averaging both the male and female voice data; then, formant were measured and formant comparison was made for controversial phonemes, which also served to resolve the controversies in the existing phoneme inventories; and finally, a consonant phoneme inventory was designed.Item Acoustic analysis of Bangla vowel inventory(BRAC University, 2008) Alam, Firoj; Habib, S. M. Murtoza; Khan, MumitThis paper describes the acoustic characteristics of Bangla vowels, obtained by analyzing the recordings of male and female voices. First, the duration of each phoneme was identified by averaging both the male and female voice data; then, formants were analyzed for all the phonemes and finally vowel phoneme inventory was designed and presented in this paper.Item Acoutstic analysis of Bangla consonants(BRAC University, 2008) Alam, Firoj; Habib, S. M. Murtoza; Khan, MumitThis paper describes the acoustic characteristics of Bangla consonants, obtained by analyzing the recordings of male and female voices. First, the duration of each phoneme was identified by averaging both the male and female voice data; then, formant were measured and formant comparison was made for controversial phonemes, which also served to resolve the controversies in the existing phoneme inventories; and finally, a consonant phoneme inventory was designed.Item Development of annotated Bangla speech corpora(BRAC University, 2010) Alam, Firoj; Habib, S. M. Murtoza; Sultana, Dil Afroza; Khan, MumitThis paper describes the development procedure of three different Bangla read speech corpora which can be used for phonetic research and developing speech applications. Several criteria were maintained in the corpora development process that includes considering the phonetic and prosodic features during text selection. On the other hand, a specification was maintained in the recording phase as the speaking style is a vital part in speech applications. We also concentrated on proper text normalization, pronunciation, aligning, and labeling. The labeling was done manually – in the present endeavor sentence level labeling (annotation) was completed by maintaining a specification so that it could be expanded in future.Item Kotha : the first to speech synthesis for Bangla language(BRAC University, 2006) Alam, Firoj; Khan, MumitIn this paper, we present Text to Speech (TTS) synthesis system for Bangla language. Here the system developed using phonology, G2P conversion and prosodic information in the festival [1] framework. Since Festival does not provide complete language processing support specific to various languages, so it is augmented with linguistic resources to facilitate the development of TTS systems. We propose how various language-processing modules such as text normalization, grapheme-to-phoneme (G2P), intonation, and duration models can be develop and integrate within Festival to develop Bangla TTS system.Item MEDIC: A Multi-Task Learning Dataset for Disaster Image Classification(Springer Nature, 2022-09-03) Alam, Firoj; Alam, Tanvirul; Hasan, Md. Arid; Hasnat, Abul; Imran, Muhammad; Ofli, FerdaRecent research in disaster informatics demonstrates a practical and important use case of artificial intelligence to save human lives and suffering during natural disasters based on social media contents (text and images). While notable progress has been made using texts, research on exploiting the images remains relatively under-explored. To advance image-based approaches, we propose MEDIC (https://crisisnlp.qcri.org/medic/index.html), which is the largest social media image classification dataset for humanitarian response consisting of 71,198 images to address four different tasks in a multi-task learning setup. This is the first dataset of its kind: social media images, disaster response, and multi-task learning research. An important property of this dataset is its high potential to facilitate research on multi-task learning, which recently receives much interest from the machine learning community and has shown remarkable results in terms of memory, inference speed, performance, and generalization capability. Therefore, the proposed dataset is an important resource for advancing image-based disaster management and multi-task machine learning research. We experiment with different deep learning architectures and report promising results, which are above the majority baselines for all tasks. Along with the dataset, we also release all relevant scriptsItem Research report on Bengali NLP engine for TTS(BRAC University, 2008-04-07) Alam, FirojThis report describes the Bengali NLP processor for TTS, along with the challenges faced in developing the NLP processor.Item Research report on Translations of gTLDs and ccTLDs in Bangla(BRAC University, 2007-10-08) Alam, Firoj; Habib, Murtoza; Hayder, Kamrul; Khan, Mumit; Khan, MumitThis report describes the initial translations of gTLDs and ccTLDs in Bengali, along with the challenges faced in creating the translations.Item Text normalization system for Bangla(BRAC University, 2008) Alam, Firoj; Habib, S. M. Murtoza; Khan, MumitThis paper describes a process of text normalization system of Bangla language (exonym: Bengali) by identifying the semiotic classes from Bangla text corpus. After identifying the semiotic classes a set of rules were written for tokenization and verbalization. This study is important for Text-To-Speech (TTS) system and as well as in language model for speech recognition.Item Text to speech for Bangla language using festival(BRAC University, 2007) Alam, Firoj; Nath, Promila Kanti; Khan, MumitIn this paper, we present a Text to Speech (TTS) synthesis system for Bangla language using the opensource Festival TTS engine. Festival is a complete TTS synthesis system, with components supporting front-end processing of the input text, language modeling, and speech synthesis using its signal processing module. The Bangla TTS system proposed here, creates the voice data for festival, and additionally extends festival using its embedded scheme scripting interface to incorporate Bangla language support. Festival is a concatenative TTS system using diphone or other unit selection speech units. Our TTS implementation uses two different kinds of these concatenative methods supported in Festival: unit selection and multisyn unit selection. The modules of such a TTS system are described in this paper, followed by an evaluation of the quality of synthesized speech for acceptability and intelligibility.Item Text to speech for Bangla language using festival(BRAC University, 2007) Alam, Firoj; Nath, Promila Kanti; Khan, MumitIn this paper, we present a Text to Speech (TTS) synthesis system for Bangla language using the open-source Festival TTS engine. Festival is a complete TTS synthesis system, with components supporting front-end processing of the input text, language modeling, and speech synthesis using its signal processing module. The Bangla TTS system proposed here, creates the voice data for festival, and additionally extends festival using its embedded scheme scripting interface to incorporate Bangla language support. Festival is a oncatenative TTS system using diphone or other unit selection speech units. Our TTS implementation uses two different kinds of these concatenative methods supported in Festival: unit selection and multisyn unit selection. The function of a Text-to-Speech system is to convert some language text into its spoken equivalent by a series of modules. These modules, constituting the TTS system are described in detail which is very much helpful for future development. Finally, the quality of synthesized speech is assessed in terms of acceptability and intelligibility.Item Z-Index at CheckThat! 2023:(Daffodil International University, 2023-08) Tarannum, Prerona; Hasan, Md. Arid; Alam, Firoj; Noori, Sheak Rashed Haider"In this study, we report our participation in CheckThat! lab’s Task 1. The aim is to determine whether a claim made in either unimodal or multimodal content is worth fact-checking. We implemented standard preprocessing and fine-tuned the XLM-RoBERTa-large model. Additionally, we applied zero-shot learning and utilized a feed-forward network with embeddings for unimodal content. For subtask 1A submission, we used combined BERT-based models (BERT and BERT multilingual), ResNet50, and Feed Forward network and we ranked as 3rd (Arabic) and 5th (English). We used feed forward network with embeddings for subtask 1B submission and ranked as 3rd in Arabic and 6th in both English and Spanish. In further experiments, our evaluation shows that XLM-RoBERTa-large model outperforms the other models"Item Zero- and Few-Shot Prompting with LLMs:(Daffodil International University, 2024-05-30) Hasan, Md. Arid; Das, Shudipta; Anjum, Afiyat; Alam, Firoj; Anjum, Anika; Sarke, Avijit; Noori, Sheak Rashed HaiderThe rapid expansion of the digital world has propelled sentiment analysis into a critical tool across diverse sectors such as marketing, politics, customer service, and healthcare. While there have been significant advancements in sentiment analysis for widely spoken languages, low-resource languages, such as Bangla, remain largely under-researched due to resource constraints. Furthermore, the recent unprecedented performance of Large Language Models (LLMs) in various applications highlights the need to evaluate them in the context of low-resource languages. In this study, we present a sizeable manually annotated dataset encompassing 33,606 Bangla news tweets and Facebook comments. We also investigate zero- and few-shot in-context learning with several language models, including Flan-T5, GPT-4, and Bloomz, offering a comparative analysis against fine-tuned models. Our findings suggest that monolingual transformer-based models consistently outperform other models, even in zero and few-shot scenarios. To foster continued exploration, we intend to make this dataset and our research tools publicly available to the broader research community.Item Zero- and Few-Shot Prompting with LLMs: A Comparative Study with Fine-tuned Models for Bangla Sentiment Analysis(Scopus, 2024) Hasan, Md. Arid; Das, Shudipta; Anjum, Afiyat; Alam, Firoj; Anjum, Anika; Sarker, Avijit; Noori, Sheak Rashed HaiderThe rapid expansion of the digital world has propelled sentiment analysis into a critical tool across diverse sectors such as marketing, politics, customer service, and healthcare. While there have been significant advancements in sentiment analysis for widely spoken languages, low-resource languages, such as Bangla, remain largely under-researched due to resource constraints. Furthermore, the recent unprecedented performance of Large Language Models (LLMs) in various applications highlights the need to evaluate them in the context of low-resource languages. In this study, we present a sizeable manually annotated dataset encompassing 33,606 Bangla news tweets and Facebook comments. We also investigate zero- and few-shot in-context learning with several language models, including Flan-T5, GPT-4, and Bloomz, offering a comparative analysis against fine-tuned models. Our findings suggest that monolingual transformer-based models consistently outperform other models, even in zero and few-shot scenarios. To foster continued exploration, we intend to make this dataset and our research tools publicly available to the broader research community.Item Zero- and Few-Shot Prompting with LLMs: A Comparative Study with Fine-tuned Models for Bangla Sentiment Analysis(Elsevier, 2024-05-25) Hasan, Md. Arid; Das, Shudipta; Anjum, Afiyat; Alam, Firoj; Anjum, Anika; Sarker, Avijit; Sheak Rashed Haider; Noori, Sheak Rashed HaiderThe rapid expansion of the digital world has propelled sentiment analysis into a critical tool across diverse sectors such as marketing, politics, customer service, and healthcare. While there have been significant advancements in sentiment analysis for widely spoken languages, low-resource languages, such as Bangla, remain largely under-researched due to resource constraints. Furthermore, the recent unprecedented performance of Large Language Models (LLMs) in various applications highlights the need to evaluate them in the context of low-resource languages. In this study, we present a sizeable manually annotated dataset encompassing 33,606 Bangla news tweets and Facebook comments. We also investigate zero- and few-shot in-context learning with several language models, including Flan-T5, GPT-4, and Bloomz, offering a comparative analysis against fine-tuned models. Our findings suggest that monolingual transformer-based models consistently outperform other models, even in zero and few-shot scenarios. To foster continued exploration, we intend to make this dataset and our research tools publicly available to the broader research community.
