English
Related papers

Related papers: Bangla Grammatical Error Detection Using T5 Transf…

200 papers

Segmentation of handwritten document images into text lines and words is one of the most significant and challenging tasks in the development of a complete Optical Character Recognition (OCR) system. This paper addresses the automatic…

Computer Vision and Pattern Recognition · Computer Science 2020-09-18 Pawan Kumar Singh , Shubham Sinha , Sagnik Pal Chowdhury , Ram Sarkar , Mita Nasipuri

Automatic Speech Recognition (ASR) transcripts, especially in low-resource languages like Bangla, contain a critical ambiguity: word-word repetitions can be either Repetition Disfluency (unintentional ASR error/hesitation) or Morphological…

Computation and Language · Computer Science 2025-11-18 Zaara Zabeen Arpa , Sadnam Sakib Apurbo , Nazia Karim Khan Oishee , Ajwad Abrar

In this paper, we present HS-BAN, a binary class hate speech (HS) dataset in Bangla language consisting of more than 50,000 labeled comments, including 40.17% hate and rest are non hate speech. While preparing the dataset a strict and…

Computation and Language · Computer Science 2021-12-06 Nauros Romim , Mosahed Ahmed , Md Saiful Islam , Arnab Sen Sharma , Hriteshwar Talukder , Mohammad Ruhul Amin

In the field of natural language processing and human-computer interaction, human attitudes and sentiments have attracted the researchers. However, in the field of human-computer interaction, human abnormality detection has not been…

Computation and Language · Computer Science 2020-07-22 M. F. Mridha , Md. Saifur Rahman , Abu Quwsar Ohi

Authorship Attribution is the task of creating an appropriate characterization of text that captures the authors' writing style to identify the original author of a given piece of text. With increased anonymity on the internet, this task…

Computation and Language · Computer Science 2024-03-11 Aisha Khatun , Anisur Rahman , Md Saiful Islam , Hemayet Ahmed Chowdhury , Ayesha Tasnim

This paper introduces a novel approach for identifying the possible large language models (LLMs) involved in text generation. Instead of adding an additional classification layer to a base LM, we reframe the classification task as a…

Computation and Language · Computer Science 2024-02-08 Yutian Chen , Hao Kang , Vivian Zhai , Liangze Li , Rita Singh , Bhiksha Raj

Most of previous work on learning diacritization of the Arabic language relied on training models from scratch. In this paper, we investigate how to leverage pre-trained language models to learn diacritization. We finetune token-free…

Computation and Language · Computer Science 2023-03-28 Bashar Al-Rfooh , Gheith Abandah , Rami Al-Rfou

Multilingual sequence-to-sequence models perform poorly with increased language coverage and fail to consistently generate text in the correct target language in few-shot settings. To address these challenges, we propose mmT5, a modular…

Computation and Language · Computer Science 2023-05-24 Jonas Pfeiffer , Francesco Piccinno , Massimo Nicosia , Xinyi Wang , Machel Reid , Sebastian Ruder

Error correction models form an important part of Automatic Speech Recognition (ASR) post-processing to improve the readability and quality of transcriptions. Most prior works use the 1-best ASR hypothesis as input and therefore can only…

Computation and Language · Computer Science 2023-10-11 Rao Ma , Mark J. F. Gales , Kate M. Knill , Mengjie Qian

Large Language Models (LLMs) have tremendous potential to play a key role in supporting mathematical reasoning, with growing use in education and AI research. However, most existing benchmarks are limited to English, creating a significant…

Computers and Society · Computer Science 2025-10-16 Tabia Tanzin Prama , Christopher M. Danforth , Peter Sheridan Dodds

Speech synthesis is one of the challenging tasks to automate by deep learning, also being a low-resource language there are very few attempts at Bangla speech synthesis. Most of the existing works can't work with anything other than simple…

Sound · Computer Science 2021-06-09 Zabir Al Nazi , Sayed Mohammed Tasmimul Huda

Sentiment analysis has been widely used to understand our views on social and political agendas or user experiences over a product. It is one of the cores and well-researched areas in NLP. However, for low-resource languages, like Bangla,…

Computation and Language · Computer Science 2020-11-23 Md. Arid Hasan , Jannatul Tajrin , Shammur Absar Chowdhury , Firoj Alam

Bangla or Bengali is the national language of Bangladesh, people from different regions don't talk in proper Bangla. Every division of Bangladesh has its own local language like Sylheti, Chittagong etc. In recent years some papers were…

Pretrained language models inherently exhibit various social biases, prompting a crucial examination of their social impact across various linguistic contexts due to their widespread usage. Previous studies have provided numerous methods…

Computation and Language · Computer Science 2024-06-26 Jayanta Sadhu , Ayan Antik Khan , Abhik Bhattacharjee , Rifat Shahriyar

The rapid spread of fake news presents a significant global challenge, particularly in low-resource languages like Bangla, which lack adequate datasets and detection tools. Although manual fact-checking is accurate, it is expensive and slow…

Computation and Language · Computer Science 2025-02-18 Hrithik Majumdar Shibu , Shrestha Datta , Md. Sumon Miah , Nasrullah Sami , Mahruba Sharmin Chowdhury , Md. Saiful Islam

In the current digital landscape, misinformation circulates rapidly, shaping public perception and causing societal divisions. It is difficult to identify hyperpartisan news in Bangla since there aren't many sophisticated natural language…

Computation and Language · Computer Science 2025-07-30 Mohammad Mehadi Hasan , Fatema Binte Hassan , Md Al Jubair , Zobayer Ahmed , Sazzatul Yeakin , Md Masum Billah

Recognition of handwritten Bangla compound characters remains a challenging problem due to complex character structures, large intra-class variation, and limited availability of high-quality annotated data. Existing Bangla handwritten…

Computer Vision and Pattern Recognition · Computer Science 2026-05-13 Md. Sultan Al Rayhan

Despite being the 5th most spoken language, Bangla remains underrepresented in Large Language Models (LLMs), particularly for code generation. This primarily stems from the scarcity of high-quality data to pre-train and/or finetune such…

Computation and Language · Computer Science 2025-09-12 Nishat Raihan , Antonios Anastasopoulos , Marcos Zampieri

This research presents a comprehensive investigation into Bangla authorship attribution, introducing a new balanced benchmark corpus BARD10 (Bangla Authorship Recognition Dataset of 10 authors) and systematically analyzing the impact of…

Computation and Language · Computer Science 2025-11-12 Abdullah Muhammad Moosa , Nusrat Sultana , Mahdi Muhammad Moosa , Md. Miraiz Hossain

In this paper, we present TituLLMs, the first large pretrained Bangla LLMs, available in 1b and 3b parameter sizes. Due to computational constraints during both training and inference, we focused on smaller models. To train TituLLMs, we…