English
Related papers

Related papers: DohaScript: A Large-Scale Multi-Writer Dataset for…

200 papers

Dockerfiles are one of the most prevalent kinds of DevOps artifacts used in industry. Despite their prevalence, there is a lack of sophisticated semantics-aware static analysis of Dockerfiles. In this paper, we introduce a dataset of…

Software Engineering · Computer Science 2020-03-31 Jordan Henkel , Christian Bird , Shuvendu K. Lahiri , Thomas Reps

Handwritten text recognition is challenging because of the virtually infinite ways a human can write the same message. Our fully convolutional handwriting model takes in a handwriting sample of unknown length and outputs an arbitrary stream…

Computer Vision and Pattern Recognition · Computer Science 2019-07-12 Felipe Petroski Such , Dheeraj Peri , Frank Brockler , Paul Hutkowski , Raymond Ptucha

To benchmark Bengali digit recognition algorithms, a large publicly available dataset is required which is free from biases originating from geographical location, gender, and age. With this aim in mind, NumtaDB, a dataset consisting of…

Computer Vision and Pattern Recognition · Computer Science 2018-06-08 Samiul Alam , Tahsin Reasat , Rashed Mohammad Doha , Ahmed Imtiaz Humayun

India's linguistic landscape, spanning 22 scheduled languages and hundreds of marginalized dialects, has driven rapid growth in NLP datasets, benchmarks, and pretrained models. However, no dedicated survey consolidates resources developed…

Computation and Language · Computer Science 2026-04-21 Raghvendra Kumar , Devankar Raj , Sriparna Saha

Progress in Automated Handwriting Recognition has been hampered by the lack of large training datasets. Nearly all research uses a set of small datasets that often cause models to overfit. We present CENSUS-HWR, a new dataset consisting of…

Computer Vision and Pattern Recognition · Computer Science 2024-09-02 Chetan Joshi , Lawry Sorenson , Ammon Wolfert , Mark Clement , Joseph Price , Kasey Buckles

The task of headline generation within the realm of Natural Language Processing (NLP) holds immense significance, as it strives to distill the true essence of textual content into concise and attention-grabbing summaries. While noteworthy…

Computation and Language · Computer Science 2023-11-30 Lokesh Madasu , Gopichand Kanumolu , Nirmal Surange , Manish Shrivastava

Dialogue assessment plays a critical role in the development of open-domain dialogue systems. Existing work are uncapable of providing an end-to-end and human-epistemic assessment dataset, while they only provide sub-metrics like coherence…

Computation and Language · Computer Science 2023-10-26 Yukun Zhao , Lingyong Yan , Weiwei Sun , Chong Meng , Shuaiqiang Wang , Zhicong Cheng , Zhaochun Ren , Dawei Yin

Handwritten Text Recognition (HTR) is still a challenging problem because it must deal with two important difficulties: the variability among writing styles, and the scarcity of labelled data. To alleviate such problems, synthetic data…

Computer Vision and Pattern Recognition · Computer Science 2020-05-28 Lei Kang , Marçal Rusiñol , Alicia Fornés , Pau Riba , Mauricio Villegas

In this work, a novel deep learning technique for the recognition of handwritten Bangla isolated compound character is presented and a new benchmark of recognition accuracy on the CMATERdb 3.1.3.3 dataset is reported. Greedy layer wise…

Computer Vision and Pattern Recognition · Computer Science 2018-02-05 Saikat Roy , Nibaran Das , Mahantapas Kundu , Mita Nasipuri

Determining the readability of a text is the first step to its simplification. In this paper, we present a readability analysis tool capable of analyzing text written in the Bengali language to provide in-depth information on its…

Computation and Language · Computer Science 2020-12-15 Susmoy Chakraborty , Mir Tafseer Nayeem , Wasi Uddin Ahmad

In this paper, we use statistical texture features for handwritten and printed text classification. We primarily aim for word level classification in south Indian scripts. Words are first extracted from the scanned document. For each…

Computer Vision and Pattern Recognition · Computer Science 2013-04-11 Mallikarjun Hangarge , K. C. Santosh , Srikanth Doddamani , Rajmohan Pardeshi

This study aims to develop a semi-automatically labelled prosody database for Hindi, for enhancing the intonation component in ASR and TTS systems, which is also helpful for building Speech to Speech Machine Translation systems. Although no…

Computation and Language · Computer Science 2021-12-14 Esha Banerjee , Atul Kr. Ojha , Girish Nath Jha

We describe a method for classification of handwritten Kannada characters using Hidden Markov Models (HMMs). Kannada script is agglutinative, where simple shapes are concatenated horizontally to form a character. This results in a large…

Machine Learning · Computer Science 2014-10-17 Manasij Venkatesh , Vikas Majjagi , Deepu Vijayasenan

Recent advancements in the field of computer vision with the help of deep neural networks have led us to explore and develop many existing challenges that were once unattended due to the lack of necessary technologies. Hand Sign/Gesture…

Computer Vision and Pattern Recognition · Computer Science 2023-05-12 Shahjalal Ahmed , Md. Rafiqul Islam , Jahid Hassan , Minhaz Uddin Ahmed , Bilkis Jamal Ferdosi , Sanjay Saha , Md. Shopon

While natural language processing tools have been developed extensively for some of the world's languages, a significant portion of the world's over 7000 languages are still neglected. One reason for this is that evaluation datasets do not…

Computation and Language · Computer Science 2024-06-05 Chunlan Ma , Ayyoob ImaniGooghari , Haotian Ye , Renhao Pei , Ehsaneddin Asgari , Hinrich Schütze

The analysis of consumer sentiment, as expressed through reviews, can provide a wealth of insight regarding the quality of a product. While the study of sentiment analysis has been widely explored in many popular languages, relatively less…

Computation and Language · Computer Science 2023-06-09 Mohsinul Kabir , Obayed Bin Mahfuz , Syed Rifat Raiyan , Hasan Mahmud , Md Kamrul Hasan

Fact-checking in code-mixed, low-resource languages such as Hinglish remains an underexplored challenge in natural language processing. Existing fact-verification systems largely focus on high-resource, monolingual settings and fail to…

Computation and Language · Computer Science 2025-08-15 Rakesh Thakur , Sneha Sharma , Gauri Chopra

Sentiment analysis, the automated process of determining emotions or opinions expressed in text, has seen extensive exploration in the field of natural language processing. However, one aspect that has remained underrepresented is the…

Computation and Language · Computer Science 2024-09-16 Mouad Jbel , Mourad Jabrane , Imad Hafidi , Abdulmutallib Metrane

Large, openly licensed speech datasets are essential for building automatic speech recognition (ASR) systems, yet many widely spoken languages remain underrepresented in public resources. Pashto, spoken by more than 60 million people, has…

Computation and Language · Computer Science 2026-02-17 Jandad Jahani , Mursal Dawodi , Jawid Ahmad Baktash

Human Interactive Proofs (HIPs) are automatic reverse Turing tests designed to distinguish between various groups of users. Completely Automatic Public Turing test to tell Computers and Humans Apart (CAPTCHA) is a HIP system that…

Cryptography and Security · Computer Science 2011-09-02 Sushma Yalamanchili , M. Kameswara Rao