English
Related papers

Related papers: BanglaLlama: LLaMA for Bangla Language

200 papers

In this work, we explain our approach employed in the BabyLM Challenge, which uses various methods of training language models (LMs) with significantly less data compared to traditional large language models (LLMs) and are inspired by how…

Computation and Language · Computer Science 2025-03-07 Mohammad Amin Ghanizadeh , Mohammad Javad Dousti

As large language models (LLMs) see increasing adoption across the globe, it is imperative for LLMs to be representative of the linguistic diversity of the world. India is a linguistically diverse country of 1.4 Billion people. To…

Computation and Language · Computer Science 2024-08-09 Harman Singh , Nitish Gupta , Shikhar Bharadwaj , Dinesh Tewari , Partha Talukdar

Recent NLP advances focus primarily on standardized languages, leaving most low-resource dialects under-served especially in Indian scenarios. In India, the issue is particularly important: despite Hindi being the third most spoken language…

Computation and Language · Computer Science 2026-01-16 Tarun Sharma , Manikandan Ravikiran , Sourava Kumar Behera , Pramit Bhattacharya , Arnab Bhattacharya , Rohit Saluja

High-quality data resources play a crucial role in learning large language models (LLMs), particularly for low-resource languages like Cantonese. Despite having more than 85 million native speakers, Cantonese is still considered a…

Computation and Language · Computer Science 2025-03-06 Jiyue Jiang , Alfred Kar Yin Truong , Yanyu Chen , Qinghang Bao , Sheng Wang , Pengan Chen , Jiuming Wang , Lingpeng Kong , Yu Li , Chuan Wu

Bangla is a low-resource language for code generation, lacking large-scale annotated datasets and tools to transform natural language specifications into executable programs. This makes Bangla-to-code generation a challenging task requiring…

Software Engineering · Computer Science 2025-12-23 Mahir Labib Dihan , Sadif Ahmed , Md Nafiu Rahman

People commonly communicate in English, Arabic, and Bengali spoken languages through various mediums. However, deaf and hard-of-hearing individuals primarily use body language and sign language to express their needs and achieve…

Computer Vision and Pattern Recognition · Computer Science 2024-08-21 Md Hadiuzzaman , Mohammed Sowket Ali , Tamanna Sultana , Abdur Raj Shafi , Abu Saleh Musa Miah , Jungpil Shin

This study presents BanStereoSet, a dataset designed to evaluate stereotypical social biases in multilingual LLMs for the Bangla language. In an effort to extend the focus of bias research beyond English-centric datasets, we have localized…

Computation and Language · Computer Science 2025-06-02 Mahammed Kamruzzaman , Abdullah Al Monsur , Shrabon Das , Enamul Hassan , Gene Louis Kim

Audio large language models (AudioLLMs) enable instruction-following over speech and general audio, but progress is increasingly limited by the lack of diverse, conversational, instruction-aligned speech-text data. This bottleneck is…

Urdu, spoken by 230 million people worldwide, lacks dedicated transformer-based language models and curated corpora. While multilingual models provide limited Urdu support, they suffer from poor performance, high computational costs, and…

Computation and Language · Computer Science 2026-01-27 Syed Muhammad Ali , Hammad Sajid , Zainab Haider , Ali Muhammad Asad , Haya Fatima , Abdul Samad

The proliferation of transliterated texts in digital spaces has emphasized the need for detecting and classifying hate speech in languages beyond English, particularly in low-resource languages. As online discourse can perpetuate…

Although the Indonesian language is spoken by almost 200 million people and the 10th most spoken language in the world, it is under-represented in NLP research. Previous work on Indonesian has been hampered by a lack of annotated datasets,…

Computation and Language · Computer Science 2020-11-03 Fajri Koto , Afshin Rahimi , Jey Han Lau , Timothy Baldwin

Large Language Models (LLMs) have exhibited remarkable capabilities in understanding and interacting with natural language across various sectors. However, their effectiveness is limited in specialized areas requiring high accuracy, such as…

Computation and Language · Computer Science 2024-01-04 Xianjun Yang , Junfeng Gao , Wenxin Xue , Erik Alexandersson

Question-Answering (QA) models for low-resource languages like Bangla face challenges due to limited annotated data and linguistic complexity. A key issue is determining whether models rely more on pre-encoded (parametric) knowledge or…

Computation and Language · Computer Science 2026-02-03 Umme Abira Azmary , MD Ikramul Kayes , Swakkhar Shatabda , Farig Yousuf Sadeque

In pursuit of more inclusive Vision-Language Models (VLMs), this study introduces a Large Multilingual Multimodal Model called PALO. PALO offers visual reasoning capabilities in 10 major languages, including English, Chinese, Hindi,…

Large Language Models (LLMs) play a critical role in how humans access information. While their core use relies on comprehending written requests, our understanding of this ability is currently limited, because most benchmarks evaluate LLMs…

Bangla handwriting recognition is becoming a very important issue nowadays. It is potentially a very important task specially for Bangla speaking population of Bangladesh and West Bengal. By keeping that in our mind we are introducing a…

Computation and Language · Computer Science 2017-04-03 Mithun Biswas , Rafiqul Islam , Gautam Kumar Shom , Md Shopon , Nabeel Mohammed , Sifat Momen , Md Anowarul Abedin

Bangla Sign Language Translation (BdSLT) has been severely constrained so far as the language itself is very low resource. Standard sentence level dataset creation for BdSLT is of immense importance for developing AI based assistive tools…

Computation and Language · Computer Science 2025-11-27 Husne Ara Rubaiyeat , Hasan Mahmud , Md Kamrul Hasan

Large Language Models (LLMs) are increasingly used as automated annotators to scale dataset creation, yet their reliability as unbiased annotators--especially for low-resource and identity-sensitive settings--remains poorly understood. In…

Computation and Language · Computer Science 2026-03-03 Md. Najib Hasan , Touseef Hasan , Souvika Sarkar

Being the seventh most spoken language in the world, the use of the Bangla language online has increased in recent times. Hence, it has become very important to analyze Bangla text data to maintain a safe and harassment-free online place.…

Computation and Language · Computer Science 2021-02-05 Md Faisal Ahmed , Zalish Mahmud , Zarin Tasnim Biash , Ahmed Ann Noor Ryen , Arman Hossain , Faisal Bin Ashraf

Low-resource languages present unique challenges to (neural) machine translation. We discuss the case of Bambara, a Mande language for which training data is scarce and requires significant amounts of pre-processing. More than the…