17 papers · ranked by Valyu relevance
Sagor Sarker
BNLP is an open-source language processing toolkit for Bengali consisting of tokenization, word embedding, part of speech(POS) tagging, name entity recognition(NER) facilities. BNLP provides pre-trained model with high accuracy to do model-based tokenization, embedding, POS, NER tasks for Bengali. BNLP pre-trained…
Arpita Bose, Niladri S. Dash, Samrah Ahmed, Manaswita Dutta + 4 more
'Aparna Dutt' 'Ranita Nandi' 'Yesi Cheng' 'Tina M. D. Mello'] Background and aim: Speech and language characteristics of connected speech provide a valuable tool for identifying, diagnosing and monitoring progression in Alzheimer's Disease (AD). Our knowledge of linguistic features of connected speech in AD is…
Haque, Farah Binta, Mohammed Nasheed Yasin, Saha + 3 more
Despite the growing progress in Natural Language Inference (NLI) research, resources for the Bengali language remain extremely limited. Existing Bengali NLI datasets exhibit several inconsistencies, including annotation errors, ambiguous sentence pairs, and inadequate linguistic diversity, which hinder effective model…
Firoj Ahmmed Patwary, Abdullah Al Noman
Tokenization is an important first step in Natural Language Processing (NLP) pipelines because it decides how models learn and represent linguistic information. But current subword tokenizers like SentencePiece or HuggingFace BPE are mostly made for Latin or multilingual corpora and don't work well on languages with a…
Abhayjeet Singh, Arjun Singh Mehta, Ashish Khuraishi K S, G Deekshitha + 16 more
'G Deekshitha' 'Gauri Date' 'Jai Nanavati' 'Jesuraja Bandekar' 'Karnalius Basumatary' 'P. Karthika' 'Sandhya Badiger' 'Sathvik Udupa' 'Saurabh Kumar' 'Savitha' 'Prasanta Kumar Ghosh' 'Vempaty Prashanthi' 'Priyanka Pai' 'Raoul Nanavati' 'Rohan Saxena' 'Sai Praneeth Reddy Mora' 'Srinivasa R. Raghavan'] Automatic speech…
Umme Ayman, Md. Nahid Hasan, Ms. Nusrat Khan, Ms. Chayti Saha + 2 more
'Md. Fayejullah' 'Saman Kasmaiee'] Tense classification in Bengali sentences is a fundamental yet unsolved problem of Bangla natural language processing (NLP) which is essential for tasks like machine translation, sentiment analysis, grammar correction, writing assistance and sentence generation. This study addresses…
Abhik Bhattacharjee, Tahmid Hasan, Wasi Uddin Ahmad, Kazi Samin + 4 more
'Md. Saiful Islam' 'Anindya Iqbal' 'M. Sohel Rahman' 'Rifat Shahriyar'] In this work, we introduce BanglaBERT, a BERT-based Natural Language Understanding (NLU) model pretrained in Bangla, a widely spoken yet low-resource language in the NLP literature. To pretrain BanglaBERT, we collect 27.5 GB of Bangla pretraining…
Samiul Alam, Asif Shahriyar Sushmit, Zaowad Abdullah, Shahrin Nakkhatra + 5 more
'Shahrin Nakkhatra' 'MD. Nazmuddoha Ansary' 'Syed Mobassir Hossen' 'Sazia Mehnaz' 'Tahsin Reasat' 'Ahmed Imtiaz Humayun'] Bengali is one of the most spoken languages in the world with over 300 million speakers globally. Despite its popularity, research into the development of Bengali speech recognition systems is…
Md. Julkar Naeen, Sourav Kumar Das, Sakib Alam Jisan, Sharun Akter Khushbu + 2 more
Classifying scattered Bengali text is the primary focus of this study, with an emphasis on explainability in Natural Language Processing (NLP) for low-resource languages. We employed supervised Machine Learning (ML) models as a baseline and compared their performance with Long Short-Term Memory (LSTM) networks from the…
Sanchita Mondal, Debnarayan Khatua, Sourav Mandal, Dilip K. Prasad + 1 more
'Arif Ahmed Sekh'] Solving math word problems of varying complexities is one of the most challenging and exciting research questions in artificial intelligence (AI), particularly in natural language processing (NLP) and machine learning (ML). Foundational language models such as GPT must be evaluated for intelligence…
Salim Sazzed, Kathiravan Srinivasan
Bengali is a low-resource language that lacks tools and resources for various natural language processing (NLP) tasks, such as sentiment analysis or profanity identification. In Bengali, only the translated versions of English sentiment lexicons are available. Moreover, no dictionary exists for detecting profanity in…
Khadija Akter Lima, Khan Md Hasib, Sami Azam, Asif Karim + 4 more
'Sidratul Montaha' 'Sheak Rashed Haider Noori' 'Mirjam Jonkman' 'Valerie L. Shalin'] Named Entity Recognition (NER) plays a significant role in enhancing the performance of all types of domain specific applications in Natural Language Processing (NLP). According to the type of application, the goal of NER is to…
Pramit Bhattacharyya, Arnab Bhattacharya
Bangla is still in its nascent stage. This is mostly due to the need for a large corpus of grammatically incorrect sentences, with their corresponding correct counterparts. The present state-of-the-art techniques to curate a corpus for grammatically wrong sentences involve random swapping, insertion and deletion of…
Ranajit Das, Priyanka Upadhyai
The Indian subcontinent includes India, Bangladesh, Pakistan, Nepal, Bhutan, and Sri Lanka that collectively share common anthropological and cultural roots. Given the enigmatic population structure, complex history and genetic heterogeneity of populations from this region, their biogeographical origin and history…
Ishtiaque Ahammad, Anisur Rahman, Zeshan Mahmud Chowdhury, Arittra Bhattacharjee + 6 more
Human gut microbiome is influenced by ethnicity and other factors. In this study, we have explored the gut microbiome of Bengali population (n=13) and four Tibeto-Burman indigenous communities-Chakma (n=15), Marma (n=6), Khyang (n=10), and Tripura (n=11) using 16S rRNA amplicon sequencing. A total of 19 characterized…
Arittra Bhattacharjee, Tabassum Binte Jamal, Ishtiaque Ahammad, Zeshan Mahmud Chowdhury + 7 more
Antibiotic resistance management is a challenging task in Low and Middle-Income Countries (LMICs) such as Bangladesh. Improper regulation and uncontrolled spreading of Antibiotic Resistant Genes (ARGs) from LIMCs pose a great threat to global public health. The human gut microbiome is a massive reservoir of Antibiotic…
R. Kapuri, A. K. Sinha, P. de, R. Roy + 1 more
Nandus banshlaii, sp. nov. is described from the Banshlai River of West Bengal. This species is distinguished from all its congeners in having a golden brown body in live and a combination of characters like longest head and snout length (44.28% SL and 35.58% HL respectively) and from its two Indian congeners in…