中文
相关论文

相关论文: Uddessho: An Extensive Benchmark Dataset for Multi…

200 篇论文

Annotation automation via Large Language Models (LLMs) is the core approach for scaling NLP datasets; however, LLM behavior with respect to closed-set instructions in low-resource languages has not been well studied. We present MultiSoc-4D,…

In recent years, social media platforms have become prominent spaces for individuals to express their opinions on ongoing events, including criminal incidents. As a result, public sentiment can shift dynamically over time. This study…

The Bangla linguistic variety is a fascinating mix of regional dialects that contributes to the cultural diversity of the Bangla-speaking community. Despite extensive study into translating Bangla to English, English to Bangla, and Banglish…

Sociotechnical systems, such as language technologies, frequently exhibit identity-based biases. These biases exacerbate the experiences of historically marginalized communities and remain understudied in low-resource contexts. While models…

计算与语言 · 计算机科学 2026-05-08 Dipto Das , Shion Guha , Bryan Semaan

In this research, we propose a complete set of approaches for identifying and extracting emotions from Bangla texts. We provide a Bangla emotion classifier for six classes: anger, disgust, fear, joy, sadness, and surprise, from Bangla words…

计算与语言 · 计算机科学 2023-06-14 Md Sakib Ullah Sourav , Huidong Wang , Mohammad Sultan Mahmud , Hua Zheng

As voice assistants cement their place in our technologically advanced society, there remains a need to cater to the diverse linguistic landscape, including colloquial forms of low-resource languages. Our study introduces the first-ever…

计算与语言 · 计算机科学 2023-10-18 Fardin Ahsan Sakib , A H M Rezaul Karim , Saadat Hasan Khan , Md Mushfiqur Rahman

Sign language discourse is an essential mode of daily communication for the deaf and hard-of-hearing people. However, research on Bangla Sign Language (BdSL) faces notable limitations, primarily due to the lack of datasets. Recognizing…

计算机视觉与模式识别 · 计算机科学 2025-12-23 Husne Ara Rubaiyeat , Hasan Mahmud , Ahsan Habib , Md. Kamrul Hasan

The Bangla language is the seventh most spoken language, with 265 million native and non-native speakers worldwide. However, English is the predominant language for online resources and technical knowledge, journals, and documentation.…

The exponential growths of social media and micro-blogging sites not only provide platforms for empowering freedom of expressions and individual voices, but also enables people to express anti-social behaviour like online harassment,…

Offensive content is pervasive in social media and a reason for concern to companies and government organizations. Several studies have been recently published investigating methods to detect the various forms of such content (e.g. hate…

计算与语言 · 计算机科学 2021-05-21 Tharindu Ranasinghe , Marcos Zampieri

This study presents a large multi-modal Bangla YouTube clickbait dataset consisting of 253,070 data points collected through an automated process using the YouTube API and Python web automation frameworks. The dataset contains 18 diverse…

机器学习 · 计算机科学 2023-10-19 Abdullah Al Imran , Md Sakib Hossain Shovon , M. F. Mridha

An robust sign language recognition system can greatly alleviate communication barriers, particularly for people who struggle with verbal communication. This is crucial for human growth and progress as it enables the expression of thoughts,…

计算机视觉与模式识别 · 计算机科学 2023-04-11 Md Shamimul Islam , A. J. M. Akhtarujjaman Joha , Md Nur Hossain , Sohaib Abdullah , Ibrahim Elwarfalli , Md Mahedi Hasan

Recognition of handwritten Bangla compound characters remains a challenging problem due to complex character structures, large intra-class variation, and limited availability of high-quality annotated data. Existing Bangla handwritten…

计算机视觉与模式识别 · 计算机科学 2026-05-13 Md. Sultan Al Rayhan

The widespread presence of offensive language on social media motivated the development of systems capable of recognizing such content automatically. Apart from a few notable exceptions, most research on automatic offensive language…

计算与语言 · 计算机科学 2021-09-09 Saurabh Gaikwad , Tharindu Ranasinghe , Marcos Zampieri , Christopher M. Homan

Human communication is inherently multimodal and asynchronous. Analyzing human emotions and sentiment is an emerging field of artificial intelligence. We are witnessing an increasing amount of multimodal content in local languages on social…

This paper introduces the approach of "Gradient Masters" for BLP-2025 Task 1: "Bangla Multitask Hate Speech Identification Shared Task". We present an ensemble-based fine-tuning strategy for addressing subtasks 1A (hate-type classification)…

计算与语言 · 计算机科学 2025-11-25 Syed Mohaiminul Hoque , Naimur Rahman , Md Sakhawat Hossain

This paper introduces \textit{Bangla Key2Text}, a large-scale dataset of $2.6$ million Bangla keyword--text pairs designed for keyword-driven text generation in a low-resource language. The dataset is constructed using a BERT-based keyword…

计算与语言 · 计算机科学 2026-04-22 Tonmoy Talukder , G M Shahariar

Figurative language understanding remains a significant challenge for Large Language Models (LLMs), especially for low-resource languages. To address this, we introduce a new idiom dataset, a large-scale, culturally-grounded corpus of…

计算与语言 · 计算机科学 2026-02-16 Adib Sakhawat , Shamim Ara Parveen , Md Ruhul Amin , Shamim Al Mahmud , Md Saiful Islam , Tahera Khatun

Large Language Models have demonstrated strong multilingual fluency, yet fluency alone does not guarantee socially appropriate language use. In high-context languages, communicative competence requires sensitivity to social hierarchy,…

Social media platforms are hubs for multimodal information exchange, encompassing text, images, and videos, making it challenging for machines to comprehend the information or emotions associated with interactions in online spaces.…

计算与语言 · 计算机科学 2024-09-04 Yiqiao Jin , Minje Choi , Gaurav Verma , Jindong Wang , Srijan Kumar