English
Related papers

Related papers: Image Pre-processing on NumtaDB for Bengali Handwr…

200 papers

Handwritten document image binarization is challenging due to high variability in the written content and complex background attributes such as page style, paper quality, stains, shadow gradients, and non-uniform illumination. While the…

Computer Vision and Pattern Recognition · Computer Science 2021-11-04 Kaustubh Sadekar , Ashish Tiwari , Prajwal Singh , Shanmuganathan Raman

This paper presents a high-quality dataset for evaluating the quality of Bangla word embeddings, which is a fundamental task in the field of Natural Language Processing (NLP). Despite being the 7th most-spoken language in the world, Bangla…

Computation and Language · Computer Science 2023-04-11 Mousumi Akter , Souvika Sarkar , Shubhra Kanti Karmaker Santu

The contributions in this article are two-fold. First, we introduce a new hand-written digit data set that we collected. It contains high-resolution images of hand-written The contributions in this article are two-fold. First, we introduce…

Computer Vision and Pattern Recognition · Computer Science 2022-11-29 Cédric Beaulac , Jeffrey S. Rosenthal

In dealing with the problem of recognition of handwritten character patterns of varying shapes and sizes, selection of a proper feature set is important to achieve high recognition performance. The current research aims to evaluate the…

Computer Vision and Pattern Recognition · Computer Science 2014-10-03 Nibaran Das , Sandip Pramanik , Subhadip Basu , Punam Kumar Saha , Ram Sarkar , Mahantapas Kundu , Mita Nasipuri

At present, recognition of the Bangla handwriting compound character has been an essential issue for many years. In recent years there have been application-based researches in machine learning, and deep learning, which is gained interest,…

Computer Vision and Pattern Recognition · Computer Science 2024-12-05 Jannatul Ferdous , Suvrajit Karmaker , A K M Shahariar Azad Rabby , Syed Akhter Hossain

Bangla, the seventh most widely spoken language worldwide with 300 million native speakers, faces digital under-representation due to limited resources and lack of annotated datasets. Stemming, a critical preprocessing step in language…

Computation and Language · Computer Science 2025-08-22 Abhijit Paul , Mashiat Amin Farin , Sharif Md. Abdullah , Ahmedul Kabir , Zarif Masud , Shebuti Rayana

The importance of Scene Text Recognition (STR) in today's increasingly digital world cannot be overstated. Given the significance of STR, data intensive deep learning approaches that auto-learn feature mappings have primarily driven the…

Computer Vision and Pattern Recognition · Computer Science 2024-03-18 Harsh Lunia , Ajoy Mondal , C V Jawahar

Understanding political discourse in online spaces is crucial for analyzing public opinion and ideological polarization. While social computing and computational linguistics have explored such discussions in English, such research efforts…

Computation and Language · Computer Science 2025-06-10 Dipto Das , Syed Ishtiaque Ahmed , Shion Guha

This article includes a comprehensive collection of over 800 high-resolution streetlight images taken systematically from India's major streets, primarily in the Chennai region. The images were methodically collected following standardized…

Computer Vision and Pattern Recognition · Computer Science 2024-07-02 Eliza Femi Sherley S , Sanjay T , Shri Kaanth P , Jeffrey Samuel S

Large language models (LLMs) have achieved remarkable success across various natural language processing tasks. However, most LLM models use traditional tokenizers like BPE and SentencePiece, which fail to capture the finer nuances of a…

Computation and Language · Computer Science 2025-05-26 Pramit Bhattacharyya , Arnab Bhattacharya

One of the most popular downstream tasks in the field of Natural Language Processing is text classification. Text classification tasks have become more daunting when the texts are code-mixed. Though they are not exposed to such text during…

Computation and Language · Computer Science 2024-03-15 Md Nishat Raihan , Dhiman Goswami , Antara Mahmud

Code-mixing is a well-studied linguistic phenomenon when two or more languages are mixed in text or speech. Several datasets have been build with the goal of training computational models for code-mixing. Although it is very common to…

Computation and Language · Computer Science 2023-11-30 Md Nishat Raihan , Dhiman Goswami , Antara Mahmud , Antonios Anastasopoulos , Marcos Zampieri

We report generation of a MNIST [4] compatible data set [1] for Tamil vowels to enable building a classification DNN or other such ML/AI deep learning [2] models for Tamil OCR/Handwriting applications. We report the capability of the 60,000…

Computer Vision and Pattern Recognition · Computer Science 2020-06-18 Muthiah Annamalai

With the rise of "Metaverse" and "Web 3.0", Non-Fungible Token (NFT) has emerged as a kind of pivotal digital asset, garnering significant attention. By the end of March 2024, more than 1.7 billion NFTs have been minted across various…

Information Retrieval · Computer Science 2024-10-18 Shuxun Wang , Yunfei Lei , Ziqi Zhang , Wei Liu , Haowei Liu , Li Yang , Wenjuan Li , Bing Li , Weiming Hu

Bengali is spoken by over 230 million people yet remains severely under-served in automatic speech recognition (ASR) and speaker diarization research. In this paper, we present our system for the DL Sprint 4.0 Bengali Long-Form Speech…

Computation and Language · Computer Science 2026-03-23 Md. Nazmus Sakib , Shafiul Tanvir , Mesbah Uddin Ahamed , H. M. Aktaruzzaman Mukdho

The deep neural networks used in modern computer vision systems require enormous image datasets to train them. These carefully-curated datasets typically have a million or more images, across a thousand or more distinct categories. The…

Computer Vision and Pattern Recognition · Computer Science 2021-12-20 Connor Anderson , Ryan Farrell

Handwritten character recognition is a challenging research in the field of document image analysis over many decades due to numerous reasons such as large writing styles variation, inherent noise in data, expansive applications it offers,…

Computer Vision and Pattern Recognition · Computer Science 2021-07-21 Noushath Shaffi , Faizal Hajamohideen

Finding local invariant patterns in handwrit-ten characters and/or digits for optical character recognition is a difficult task. Variations in writing styles from one person to another make this task challenging. We have proposed a…

Computer Vision and Pattern Recognition · Computer Science 2020-04-28 Animesh Singh , Ritesh Sarkhel , Nibaran Das , Mahantapas Kundu , Mita Nasipuri

In this paper, we present HS-BAN, a binary class hate speech (HS) dataset in Bangla language consisting of more than 50,000 labeled comments, including 40.17% hate and rest are non hate speech. While preparing the dataset a strict and…

Computation and Language · Computer Science 2021-12-06 Nauros Romim , Mosahed Ahmed , Md Saiful Islam , Arnab Sen Sharma , Hriteshwar Talukder , Mohammad Ruhul Amin

The dramatic increase in the use of social media platforms for information sharing has also fueled a steep growth in online abuse. A simple yet effective way of abusing individuals or communities is by creating memes, which often integrate…

Computer Vision and Pattern Recognition · Computer Science 2023-10-19 Mithun Das , Animesh Mukherjee
‹ Prev 1 4 5 6 7 8 10 Next ›