English
Related papers

Related papers: SIMARA: a database for key-value information extra…

200 papers

We introduce a novel approach for scanned document representation to perform field extraction. It allows the simultaneous encoding of the textual, visual and layout information in a 3-axis tensor used as an input to a segmentation model. We…

Computer Vision and Pattern Recognition · Computer Science 2021-07-06 Mohamed Kerroumi , Othmane Sayem , Aymen Shabou

Semantic legal metadata provides information that helps with understanding and interpreting legal provisions. Such metadata is therefore important for the systematic analysis of legal requirements. However, manually enhancing a large legal…

Software Engineering · Computer Science 2020-01-31 Amin Sleimi , Nicolas Sannier , Mehrdad Sabetzadeh , Lionel Briand , Marcello Ceci , John Dann

Wikidata has grown to a knowledge graph with an impressive size. To date, it contains more than 17 billion triples collecting information about people, places, films, stars, publications, proteins, and many more. On the other side, most of…

Computation and Language · Computer Science 2024-01-17 Kunpeng Guo , Dennis Diefenbach , Antoine Gourru , Christophe Gravier

To reduce environmental risks and impacts from orphaned wells (abandoned oil and gas wells), it is essential to first locate and then plug these wells. Although some historical documents are available, they are often unstructured, not…

Information Retrieval · Computer Science 2024-05-10 Zhiwei Ma , Javier E. Santo , Greg Lackey , Hari Viswanathan , Daniel O'Malley

We present an effective multifaceted system for exploratory analysis of highly heterogeneous document collections. Our system is based on intelligently tagging individual documents in a purely automated fashion and exploiting these tags in…

Computation and Language · Computer Science 2013-08-13 Arun S. Maiya , John P. Thompson , Francisco Loaiza-Lemos , Robert M. Rolfe

Processing large amounts of data is an essential problem of the big data era. Most of the data exchange is done via direct communication (using APIs) and well-structured file formats (JSON, XML, EDI, etc.), but a significant portion of the…

Information Retrieval · Computer Science 2020-07-17 Vladimir Bernstein , Andrei Afanassenkov

Many models have been proposed for vision and language tasks, especially the image-text retrieval task. All state-of-the-art (SOTA) models in this challenge contained hundreds of millions of parameters. They also were pretrained on a large…

Computer Vision and Pattern Recognition · Computer Science 2023-01-13 Manh-Duy Nguyen , Binh T. Nguyen , Cathal Gurrin

Scientists, governments, and companies increasingly publish datasets on the Web. Google's Dataset Search extracts dataset metadata -- expressed using schema.org and similar vocabularies -- from Web pages in order to make datasets…

Information Retrieval · Computer Science 2020-06-15 Omar Benjelloun , Shiyu Chen , Natasha Noy

Background: Structured information extraction from unstructured histopathology reports facilitates data accessibility for clinical research. Manual extraction by experts is time-consuming and expensive, limiting scalability. Large language…

Saada transforms a set of heterogeneous FITS files or VOTables of various categories (images, tables, spectra ...) in a database without writing code. Databases created with Saada come with a rich Web interface and an Application…

Instrumentation and Methods for Astrophysics · Physics 2014-09-02 Laurent Michel , Christian Motch , Hoan Ngoc Nguyen , François-Xavier Pineau

Four years after the last LISA meeting, the NASA Astrophysics Data System (ADS) finds itself in the middle of major changes to the infrastructure and contents of its database. In this paper we highlight a number of features of great…

In this paper, we demonstrate that information retrieval can be accomplished with a single Transformer, in which all information about the corpus is encoded in the parameters of the model. To this end, we introduce the Differentiable Search…

Computation and Language · Computer Science 2022-10-24 Yi Tay , Vinh Q. Tran , Mostafa Dehghani , Jianmo Ni , Dara Bahri , Harsh Mehta , Zhen Qin , Kai Hui , Zhe Zhao , Jai Gupta , Tal Schuster , William W. Cohen , Donald Metzler

Accessibility to historical documents is mostly limited to scholars. This is due to the language barrier inherent in human language and the linguistic properties of these documents. Given a historical document, modernization aims to…

Computation and Language · Computer Science 2020-03-05 Miguel Domingo , Francisco Casacuberta

The pressing need for digitization of historical documents has led to a strong interest in designing computerised image processing methods for automatic handwritten text recognition. However, not much attention has been paid on studying the…

Computer Vision and Pattern Recognition · Computer Science 2024-01-31 Liang Cheng , Jonas Frankemölle , Adam Axelsson , Ekta Vats

This paper presents a complete workflow designed for extracting information from Quebec handwritten parish registers. The acts in these documents contain individual and family information highly valuable for genetic, demographic and social…

Computer Vision and Pattern Recognition · Computer Science 2023-04-28 Solène Tarride , Martin Maarand , Mélodie Boillet , James McGrath , Eugénie Capel , Hélène Vézina , Christopher Kermorvant

Context: User assistance is generally defined as the guided assistance to a user of a software system in order to help accomplish tasks and enhance user experience. Automated user assistance systems are equipped with online help system that…

Human-Computer Interaction · Computer Science 2020-03-13 Murat Acar , Bedir Tekinerdogan

It is becoming common to archive research datasets that are not only large but also numerous. In addition, their corresponding metadata and the software required to analyse or display them need to be archived. Yet the manual curation of…

Digital Libraries · Computer Science 2011-08-24 Daniel Lemire , Andre Vellino

In this paper, we address the problem of classifying documents available from the global network of (open access) repositories according to their type. We show that the metadata provided by repositories enabling us to distinguish research…

Digital Libraries · Computer Science 2017-07-14 Aristotelis Charalampous , Petr Knoth

Manual annotation of textual documents is a necessary task when constructing benchmark corpora for training and evaluating machine learning algorithms. We created a comprehensive directory of annotation tools that currently includes 93…

Computation and Language · Computer Science 2020-10-15 Mariana Neves , Jurica Seva

Tablets and styluses are increasingly popular for taking notes. To optimize this experience and ensure a smooth and efficient workflow, it's important to develop methods for accurately interpreting and understanding the content of…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Anastasiia Fadeeva , Vincent Coriou , Diego Antognini , Claudiu Musat , Andrii Maksai
‹ Prev 1 4 5 6 7 8 10 Next ›