English
Related papers

Related papers: Explainable Coarse-to-Fine Ancient Manuscript Dupl…

200 papers

Thousands of users consult digital archives daily, but the information they can access is unrepresentative of the diversity of documentary history. The sequence-to-sequence architecture typically used for optical character recognition (OCR)…

Computer Vision and Pattern Recognition · Computer Science 2024-07-29 Jacob Carlson , Tom Bryan , Melissa Dell

We develop and evaluate multilingual scientific documents similarity measurement models in this work. Such models can be used to find related works in different languages, which can help multilingual researchers find and explore papers more…

Computation and Language · Computer Science 2023-09-20 Yang Gao , Ji Ma , Ivan Korotkov , Keith Hall , Dana Alon , Don Metzler

The biomedical literature contains a vast collection of omics studies, yet most published data remain functionally inaccessible for computational reuse. When raw data are deposited in public repositories, essential information for…

Genomics · Quantitative Biology 2026-03-16 Alexandre Hutton , Jesse G. Meyer

The importance of an efficient and scalable document similarity detection system is undeniable nowadays. Search engines need batch text similarity measures to detect duplicated and near-duplicated web pages in their indexes in order to…

Information Retrieval · Computer Science 2018-10-09 Hamid Mohammadi , Amin Nikoukaran

While dense retrieval models, which embed queries and documents into a shared low-dimensional space, have gained widespread popularity, they were shown to exhibit important theoretical limitations and considerably lag behind traditional…

Information Retrieval · Computer Science 2026-04-09 Adrian Bracher , Svitlana Vakulenko

In this work, we focus on the problem of retrieving relevant arguments for a query claim covering diverse aspects. State-of-the-art methods rely on explicit mappings between claims and premises, and thus are unable to utilize large…

Information Retrieval · Computer Science 2021-03-18 Michael Fromm , Max Berrendorf , Sandra Obermeier , Thomas Seidl , Evgeniy Faerman

This paper introduces PreP-OCR, a two-stage pipeline that combines document image restoration with semantic-aware post-OCR correction to enhance both visual clarity and textual consistency, thereby improving text extraction from degraded…

Computation and Language · Computer Science 2025-11-19 Shuhao Guan , Moule Lin , Cheng Xu , Xinyi Liu , Jinman Zhao , Jiexin Fan , Qi Xu , Derek Greene

Oracle Bone Inscription (OBI) is the earliest mature writing system in China, which represents a crucial stage in the development of hieroglyphs. Nevertheless, the substantial quantity of undeciphered OBI characters remains a significant…

Computer Vision and Pattern Recognition · Computer Science 2024-12-30 Zhicong Wu , Qifeng Su , Ke Gu , Xiaodong Shi

Text is ubiquitous in the artificial world and easily attainable when it comes to book title and author names. Using the images from the book cover set from the Stanford Mobile Visual Search dataset and additional book covers and metadata…

Information Retrieval · Computer Science 2014-11-20 Kevin Shih , Wei Di , Vignesh Jagadeesh , Robinson Piramuthu

Determining the chronology of ancient handwritten manuscripts is essential for reconstructing the evolution of ideas. For the Dead Sea Scrolls, this is particularly important. However, there is an almost complete lack of date-bearing…

For language documentation initiatives, transcription is an expensive resource: one minute of audio is estimated to take one hour and a half on average of a linguist's work (Austin and Sallabank, 2013). Recently, collecting aligned…

Computation and Language · Computer Science 2019-10-14 Marcely Zanon Boito , Aline Villavicencio , Laurent Besacier

Identification of fossil species is crucial to evolutionary studies. Recent advances from deep learning have shown promising prospects in fossil image identification. However, the quantity and quality of labeled fossil images are often…

Computer Vision and Pattern Recognition · Computer Science 2024-02-05 Chengbin Hou , Xinyu Lin , Hanhui Huang , Sheng Xu , Junxuan Fan , Yukun Shi , Hairong Lv

Cross-Lingual SynthDocs is a large-scale synthetic corpus designed to address the scarcity of Arabic resources for Optical Character Recognition (OCR) and Document Understanding (DU). The dataset comprises over 2.5 million of samples,…

Computation and Language · Computer Science 2025-11-10 Haneen Al-Homoud , Asma Ibrahim , Murtadha Al-Jubran , Fahad Al-Otaibi , Yazeed Al-Harbi , Daulet Toibazar , Kesen Wang , Pedro J. Moreno

With the recent advancements in information technology there has been a huge surge in amount of data available. But information retrieval technology has not been able to keep up with this pace of information generation resulting in over…

Computation and Language · Computer Science 2017-10-24 Nishant Nikhil , Muktabh Mayank Srivastava

Job descriptions are posted on many online channels, including company websites, job boards or social media platforms. These descriptions are usually published with varying text for the same job, due to the requirements of each platform or…

Computation and Language · Computer Science 2024-06-11 Matthias Engelbach , Dennis Klau , Maximilien Kintz , Alexander Ulrich

Oracle bone script (OBS), as China's earliest mature writing system, present significant challenges in automatic recognition due to their complex pictographic structures and divergence from modern Chinese characters. We introduce…

Computer Vision and Pattern Recognition · Computer Science 2024-11-28 Hanqi Jiang , Yi Pan , Junhao Chen , Zhengliang Liu , Yifan Zhou , Peng Shu , Yiwei Li , Huaqin Zhao , Stephen Mihm , Lewis C Howe , Tianming Liu

Existing ordinal embedding methods usually follow a two-stage routine: outlier detection is first employed to pick out the inconsistent comparisons; then an embedding is learned from the clean data. However, learning in a multi-stage manner…

Machine Learning · Computer Science 2018-12-06 Ke Ma , Qianqian Xu , Xiaochun Cao

Chinese ancient documents, invaluable carriers of millennia of Chinese history and culture, hold rich knowledge across diverse fields but face challenges in digitization and understanding, i.e., traditional methods only scan images, while…

A diversity of tasks use language models trained on semantic similarity data. While there are a variety of datasets that capture semantic similarity, they are either constructed from modern web data or are relatively small datasets created…

Computation and Language · Computer Science 2023-08-25 Emily Silcock , Melissa Dell

With the advent of Open Science, researchers have started to publish their research artefacts (i. e., data, software, and other products of the investigations) in order to allow others to reproduce their investigations. While this…

Digital Libraries · Computer Science 2019-05-02 Max Schröder , Frank Krüger , Sascha Spors
‹ Prev 1 3 4 5 6 7 10 Next ›