English
Related papers

Related papers: IPO-Mine: A Toolkit and Dataset for Section-Struct…

200 papers

We present a new dataset for form understanding in noisy scanned documents (FUNSD) that aims at extracting and structuring the textual content of forms. The dataset comprises 199 real, fully annotated, scanned forms. The documents are noisy…

Information Retrieval · Computer Science 2019-10-30 Guillaume Jaume , Hazim Kemal Ekenel , Jean-Philippe Thiran

Audit transaction testing validates accuracy and completeness of customer-facing statements against internal systems of record. Traditional manual, sample-based review of unstructured PDF statements is labor-intensive and does not scale to…

Software Engineering · Computer Science 2026-05-08 Santosh Vasudevan , Velu Natarajan

Developing document understanding models at enterprise scale requires large, diverse, and well-annotated datasets spanning a wide range of document types. However, collecting such data is prohibitively expensive due to privacy constraints,…

AI agents and business automation tools interacting with external web services require standardized, machine-readable information about their APIs in the form of API specifications. However, the information about APIs available online is…

Table extraction (TE) is a key challenge in visual document understanding. Traditional approaches detect tables first, then recognize their structure. Recently, interest has surged in developing methods, such as vision-language models…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Brandon Smock , Valerie Faucon-Morin , Max Sokolov , Libin Liang , Tayyibah Khanam , Amrit Ramesh , Maury Courtland

Automating information extraction from form-like documents at scale is a pressing need due to its potential impact on automating business workflows across many industries like financial services, insurance, and healthcare. The key challenge…

Machine Learning · Computer Science 2022-01-14 Beliz Gunel , Navneet Potti , Sandeep Tata , James B. Wendt , Marc Najork , Jing Xie

Due to the data-driven nature of current face identity (FaceID) customization methods, all state-of-the-art models rely on large-scale datasets containing millions of high-quality text-image pairs for training. However, none of these…

Computer Vision and Pattern Recognition · Computer Science 2025-03-28 Shuhe Wang , Xiaoya Li , Jiwei Li , Guoyin Wang , Xiaofei Sun , Bob Zhu , Han Qiu , Mo Yu , Shengjie Shen , Tianwei Zhang , Eduard Hovy

Human preferences are diverse and dynamic, shaped by regional, cultural, and social factors. Existing alignment methods like Direct Preference Optimization (DPO) and its variants often default to majority views, overlooking minority…

Computation and Language · Computer Science 2026-01-29 Wenqing Wang , Muhammad Asif Ali , Ali Shoker , Ruohan Yang , Junyang Chen , Ying Sha , Huan Wang

AI application developers typically begin with a dataset of interest and a vision of the end analytic or insight they wish to gain from the data at hand. Although these are two very important components of an AI workflow, one often spends…

Databases · Computer Science 2021-03-04 El Kindi Rezig , Michael Cafarella , Vijay Gadepally

Foundation models have shown remarkable success across scientific domains, yet their impact in chemistry remains limited due to the absence of diverse, large-scale, high-quality datasets that reflect the field's multifaceted nature. We…

Process mining techniques enable the analysis of a wide variety of processes using event data. Among the available process mining techniques, most consider a single process perspective at a time-in the shape of a model or log. In this…

Human parsing has recently attracted a lot of research interests due to its huge application potentials. However existing datasets have limited number of images and annotations, and lack the variety of human appearances and the coverage of…

Computer Vision and Pattern Recognition · Computer Science 2017-07-31 Ke Gong , Xiaodan Liang , Dongyu Zhang , Xiaohui Shen , Liang Lin

Increasingly larger number of software systems today are including data science components for descriptive, predictive, and prescriptive analytics. The collection of data science stages from acquisition, to cleaning/curation, to modeling,…

Software Engineering · Computer Science 2022-02-15 Sumon Biswas , Mohammad Wardat , Hridesh Rajan

Dependency analysis is recognized as an important field of software engineering due to a variety of reasons. There exists a large pool of tools providing assistance to software developers and architects. Analysis of inter- and intra-project…

Software Engineering · Computer Science 2021-04-20 V. Repinskiy , V. Kovalenko

IoT data markets in public and private institutions have become increasingly relevant in recent years because of their potential to improve data availability and unlock new business models. However, exchanging data in markets bears…

Cryptography and Security · Computer Science 2022-07-14 Gonzalo Munilla Garrido , Johannes Sedlmeir , Ömer Uludağ , Ilias Soto Alaoui , Andre Luckow , Florian Matthes

Medical imaging papers often focus on methodology, but the quality of the algorithms and the validity of the conclusions are highly dependent on the datasets used. As creating datasets requires a lot of effort, researchers often use…

Automatic document summarization aims to produce a concise summary covering the input document's salient information. Within a report document, the salient information can be scattered in the textual and non-textual content. However,…

Computation and Language · Computer Science 2023-02-09 Shuaiqi Liu , Jiannong Cao , Ruosong Yang , Zhiyuan Wen

Predicting the exit (e.g. bankrupt, acquisition, etc.) of privately held companies is a current and relevant problem for investment firms. The difficulty of the problem stems from the lack of reliable, quantitative and publicly available…

Machine Learning · Computer Science 2019-10-31 Giuseppe Carlo Calafiore , Marisa Hillary Morales , Vittorio Tiozzo , Serge Marquie

Digital and physical footprints are a trail of user activities collected over the use of software applications and systems. As software becomes ubiquitous, protecting user privacy has become challenging. With the increase of user privacy…

Cryptography and Security · Computer Science 2022-03-29 Pattaraporn Sangaroonsilp , Hoa Khanh Dam , Morakot Choetkiertikul , Chaiyong Ragkhitwetsagul , Aditya Ghose

In modern information retrieval (IR). achieving more than just accuracy is essential to sustaining a healthy ecosystem, especially when addressing fairness and diversity considerations. To meet these needs, various datasets, algorithms, and…

Information Retrieval · Computer Science 2025-02-18 Chen Xu , Zhirui Deng , Clara Rus , Xiaopeng Ye , Yuanna Liu , Jun Xu , Zhicheng Dou , Ji-Rong Wen , Maarten de Rijke