English
Related papers

Related papers: Stemming -- The Evolution and Current State with a…

200 papers

We present BIG-C (Bemba Image Grounded Conversations), a large multimodal dataset for Bemba. While Bemba is the most populous language of Zambia, it exhibits a dearth of resources which render the development of language technologies or…

Computation and Language · Computer Science 2023-05-30 Claytone Sikasote , Eunice Mukonde , Md Mahfuz Ibn Alam , Antonios Anastasopoulos

This chapter focuses on gender-related errors in machine translation (MT) in the context of low-resource languages. We begin by explaining what low-resource languages are, examining the inseparable social and computational factors that…

Computation and Language · Computer Science 2024-03-13 Sourojit Ghosh , Srishti Chatterjee

Sentiment analysis is an essential part of text analysis, which is a larger field that includes determining and evaluating the author's emotional state. This method is essential since it makes it easier to comprehend consumers' feelings,…

Computation and Language · Computer Science 2025-10-03 Sumaiya Tabassum

Urdu is a combination of several languages like Arabic, Hindi, English, Turkish, Sanskrit etc. It has a complex and rich morphology. This is the reason why not much work has been done in Urdu language processing. Stemming is used to convert…

Computation and Language · Computer Science 2013-10-03 Vaishali Gupta , Nisheeth Joshi , Iti Mathur

In this work, we introduce BanglaBERT, a BERT-based Natural Language Understanding (NLU) model pretrained in Bangla, a widely spoken yet low-resource language in the NLP literature. To pretrain BanglaBERT, we collect 27.5 GB of Bangla…

Computation and Language · Computer Science 2022-05-11 Abhik Bhattacharjee , Tahmid Hasan , Wasi Uddin Ahmad , Kazi Samin , Md Saiful Islam , Anindya Iqbal , M. Sohel Rahman , Rifat Shahriyar

Stemming is the process of extracting root word from the given inflection word and also plays significant role in numerous application of Natural Language Processing (NLP). Tamil Language raises several challenges to NLP, since it has rich…

Computation and Language · Computer Science 2013-10-03 M. Thangarasu , R. Manavalan

Pretrained language models inherently exhibit various social biases, prompting a crucial examination of their social impact across various linguistic contexts due to their widespread usage. Previous studies have provided numerous methods…

Computation and Language · Computer Science 2024-06-26 Jayanta Sadhu , Ayan Antik Khan , Abhik Bhattacharjee , Rifat Shahriyar

Bengali is spoken by over 230 million people yet remains severely under-served in automatic speech recognition (ASR) and speaker diarization research. In this paper, we present our system for the DL Sprint 4.0 Bengali Long-Form Speech…

Computation and Language · Computer Science 2026-03-23 Md. Nazmus Sakib , Shafiul Tanvir , Mesbah Uddin Ahamed , H. M. Aktaruzzaman Mukdho

Spoken term detection (STD) is often hindered by reliance on frame-level features and the computationally intensive DTW-based template matching, limiting its practicality. To address these challenges, we propose a novel approach that…

Audio and Speech Processing · Electrical Eng. & Systems 2024-12-24 Anup Singh , Kris Demuynck , Vipul Arora

One of the major challenges for developing automatic speech recognition (ASR) for low-resource languages is the limited access to labeled data with domain-specific variations. In this study, we propose a pseudo-labeling approach to develop…

One of the most alarming issues in digital society is hate speech (HS) on social media. The severity is so high that researchers across the globe are captivated by this domain. A notable amount of work has been conducted to address the…

Machine Learning · Computer Science 2025-04-14 Md Abdullah Al Kafi , Sumit Kumar Banshal , Md Sadman Shakib , Showrov Azam , Tamanna Alam Tabashom

A formulation of bit-string models of language evolution, based on differential equations for the population speaking each language, is introduced and preliminarily studied. Connections with replicator dynamics and diffusion processes are…

Physics and Society · Physics 2009-11-13 Damian H. Zanette

This study developed a new Bangla abstractive summarization dataset to generate concise summaries of Bangla articles from diverse sources. Most existing studies in this field have concentrated on news articles, where journalists usually…

Computation and Language · Computer Science 2025-12-17 Md. Tanzim Ferdous , Naeem Ahsan Chowdhury , Prithwiraj Bhattacharjee

Natural language processing area is still under research. But now a day it is on platform for worldwide researchers. Natural language processing includes analyzing the language based on its structure and then tagging of each word…

Computation and Language · Computer Science 2013-07-23 Miral Patel , Prem Balani

This paper introduces BnTTS (Bangla Text-To-Speech), the first framework for Bangla speaker adaptation-based TTS, designed to bridge the gap in Bangla speech synthesis using minimal training data. Building upon the XTTS architecture, our…

We present a survey covering the state of the art in low-resource machine translation research. There are currently around 7000 languages spoken in the world and almost all language pairs lack significant resources for training machine…

Computation and Language · Computer Science 2022-02-08 Barry Haddow , Rachel Bawden , Antonio Valerio Miceli Barone , Jindřich Helcl , Alexandra Birch

The selection of features for text classification is a fundamental task in text mining and information retrieval. Despite being the sixth most widely spoken language in the world, Bangla has received little attention due to the scarcity of…

Information Retrieval · Computer Science 2023-08-29 Md. Rafi-Ur-Rashid , Sami Azam , Mirjam Jonkman

Visual Question Answer (VQA) poses the problem of answering a natural language question about a visual context. Bangla, despite being a widely spoken language, is considered low-resource in the realm of VQA due to the lack of proper…

In linguistic morphology and information retrieval, stemming is the process for reducing inflected (or sometimes derived) words to their stem, base or root form; generally a written word form. In this work is presented suffix stripping…

Computation and Language · Computer Science 2012-09-21 Nikola Milošević

Recent advances in word embeddings and language models use large-scale, unlabelled data and self-supervised learning to boost NLP performance. Multilingual models, often trained on web-sourced data like Wikipedia, face challenges: few…

Computation and Language · Computer Science 2025-07-02 David Ifeoluwa Adelani
‹ Prev 1 3 4 5 6 7 10 Next ›