English
Related papers

Related papers: CENSUS-HWR: a large training dataset for offline h…

200 papers

Handwritten mathematical expression recognition (HMER) is a challenging task that has many potential applications. Recent methods for HMER have achieved outstanding performance with an encoder-decoder architecture. However, these methods…

Computer Vision and Pattern Recognition · Computer Science 2022-06-08 Ye Yuan , Xiao Liu , Wondimu Dikubab , Hui Liu , Zhilong Ji , Zhongqin Wu , Xiang Bai

We present a new dataset for form understanding in noisy scanned documents (FUNSD) that aims at extracting and structuring the textual content of forms. The dataset comprises 199 real, fully annotated, scanned forms. The documents are noisy…

Information Retrieval · Computer Science 2019-10-30 Guillaume Jaume , Hazim Kemal Ekenel , Jean-Philippe Thiran

Convolutional Neural Networks have made their mark in various fields of computer vision in recent years. They have achieved state-of-the-art performance in the field of document analysis as well. However, CNNs require a large amount of…

Computer Vision and Pattern Recognition · Computer Science 2018-01-29 Neha Gurjar , Sebastian Sudholt , Gernot A. Fink

Automatic Speech Recognition (ASR) is traditionally evaluated using Word Error Rate (WER), a metric that is insensitive to meaning. Embedding-based semantic metrics are better correlated with human perception, but decoder-based Large…

Offline handwritten text recognition from images is an important problem for enterprises attempting to digitize large volumes of handmarked scanned documents/reports. Deep recurrent models such as Multi-dimensional LSTMs have been shown to…

Computation and Language · Computer Science 2018-07-27 Arindam Chowdhury , Lovekesh Vig

Most existing work on adversarial data generation focuses on English. For example, PAWS (Paraphrase Adversaries from Word Scrambling) consists of challenging English paraphrase identification pairs from Wikipedia and Quora. We remedy this…

Computation and Language · Computer Science 2019-09-02 Yinfei Yang , Yuan Zhang , Chris Tar , Jason Baldridge

As language models grow in popularity, it becomes increasingly important to clearly measure all possible markers of demographic identity in order to avoid perpetuating existing societal harms. Many datasets for measuring bias currently…

Computation and Language · Computer Science 2022-10-31 Eric Michael Smith , Melissa Hall , Melanie Kambadur , Eleonora Presani , Adina Williams

Scene Text Image Super-resolution (STISR) aims to recover high-resolution (HR) scene text images with visually pleasant and readable text content from the given low-resolution (LR) input. Most existing works focus on recovering English…

Computer Vision and Pattern Recognition · Computer Science 2023-08-08 Jianqi Ma , Zhetong Liang , Wangmeng Xiang , Xi Yang , Lei Zhang

Handwritten text recognition has been developed rapidly in the recent years, following the rise of deep learning and its applications. Though deep learning methods provide notable boost in performance concerning text recognition,…

Computer Vision and Pattern Recognition · Computer Science 2024-04-18 George Retsinas , Giorgos Sfikas , Basilis Gatos , Christophoros Nikou

In this paper, we investigate the problem of training neural machine translation (NMT) systems with a dataset of more than 40 billion bilingual sentence pairs, which is larger than the largest dataset to date by orders of magnitude.…

Computation and Language · Computer Science 2019-10-08 Yuxian Meng , Xiangyuan Ren , Zijun Sun , Xiaoya Li , Arianna Yuan , Fei Wu , Jiwei Li

Words embedding (distributed word vector representations) have become an essential component of many natural language processing (NLP) tasks such as machine translation, sentiment analysis, word analogy, named entity recognition and word…

Computation and Language · Computer Science 2020-01-08 Idris Abdulmumin , Bashir Shehu Galadanci

In today's digital world, there is an increasing focus on soft skills. On the one hand, they facilitate innovation at companies, but on the other, they are unlikely to be automated soon. Researchers struggle with accurately approaching…

Computation and Language · Computer Science 2021-07-14 Silvia Fareri , Nicola Melluso , Filippo Chiarello , Gualtiero Fantoni

Automatic chart to text summarization is an effective tool for the visually impaired people along with providing precise insights of tabular data in natural language to the user. A large and well-structured dataset is always a key part for…

Computation and Language · Computer Science 2023-06-13 Raian Rahman , Rizvi Hasan , Abdullah Al Farhad , Md Tahmid Rahman Laskar , Md. Hamjajul Ashmafee , Abu Raihan Mostofa Kamal

As handwriting input becomes more prevalent, the large symbol inventory required to support Chinese handwriting recognition poses unique challenges. This paper describes how the Apple deep learning recognition system can accurately handle…

Computer Vision and Pattern Recognition · Computer Science 2020-05-19 Youssouf Chherawala , Hans J. G. A. Dolfing , Ryan S. Dixon , Jerome R. Bellegarda

Even today in Twenty First Century Handwritten communication has its own stand and most of the times, in daily life it is globally using as means of communication and recording the information like to be shared with others. Challenges in…

Artificial Intelligence · Computer Science 2012-05-18 Yusuf Perwej , Ashish Chaturvedi

The lack of large-scale datasets has been a major hindrance to the development of NLP tasks such as spelling correction and grammatical error correction (GEC). As a complementary new resource for these tasks, we present the GitHub Typo…

Computation and Language · Computer Science 2019-12-02 Masato Hagiwara , Masato Mita

State-of-the-art methods for handwriting recognition are based on Long Short Term Memory (LSTM) recurrent neural networks (RNN), which now provides very impressive character recognition performance. The character recognition is generally…

Computer Vision and Pattern Recognition · Computer Science 2017-09-26 Bruno Stuner , Clément Chatelain , Thierry Paquet

Detecting harmful content on social media, such as Twitter, is made difficult by the fact that the seemingly simple yes/no classification conceals a significant amount of complexity. Unfortunately, while several datasets have been collected…

Computation and Language · Computer Science 2023-11-14 Saad Almohaimeed , Saleh Almohaimeed , Ashfaq Ali Shafin , Bogdan Carbunar , Ladislau Bölöni

In this paper, we focus on Whisper, a recent automatic speech recognition model trained with a massive 680k hour labeled speech corpus recorded in diverse conditions. We first show an interesting finding that while Whisper is very robust…

Sound · Computer Science 2023-10-10 Yuan Gong , Sameer Khurana , Leonid Karlinsky , James Glass

This work presents a suite of fine-tuned Whisper models for Swedish, trained on a dataset of unprecedented size and variability for this mid-resourced language. As languages of smaller sizes are often underrepresented in multilingual…

Computation and Language · Computer Science 2025-08-15 Leonora Vesterbacka , Faton Rekathati , Robin Kurtz , Justyna Sikora , Agnes Toftgård
‹ Prev 1 8 9 10 Next ›