English
Related papers

Related papers: BullingerDB: A Dataset for Handwritten Text Recogn…

200 papers

Identifying the type of font (e.g., Roman, Blackletter) used in historical documents can help optical character recognition (OCR) systems produce more accurate text transcriptions. Towards this end, we present an active-learning strategy…

Computer Vision and Pattern Recognition · Computer Science 2016-01-28 Anshul Gupta , Ricardo Gutierrez-Osuna , Matthew Christy , Richard Furuta , Laura Mandell

Alterations in historical manuscripts such as letters represent a promising field of research. On the one hand, they help understand the construction of text. On the other hand, topics that are being considered sensitive at the time of the…

Machine Learning · Computer Science 2020-11-05 David Lassner , Anne Baillot , Sergej Dogadov , Klaus-Robert Müller , Shinichi Nakajima

Albeit Natural Language Processing has seen major breakthroughs in the last few years, transferring such advances into real-world business cases can be challenging. One of the reasons resides in the displacement between popular benchmarks…

Computation and Language · Computer Science 2024-02-16 Andrea Zugarini , Andrew Zamai , Marco Ernandes , Leonardo Rigutini

Disentangled representation learning remains challenging as the underlying factors of variation in the data do not naturally exist. The inherent complexity of real-world data makes it unfeasible to exhaustively enumerate and encapsulate all…

Computation and Language · Computer Science 2024-02-13 Jiawei Zhou , Xiaoguang Li , Lifeng Shang , Xin Jiang , Qun Liu , Lei Chen

We describe a large-scale application of methods for finding plagiarism in research document collections. The methods are applied to a collection of 284,834 documents collected by arXiv.org over a 14 year period, covering a few different…

Databases · Computer Science 2007-05-23 Daria Sorokina , Johannes Gehrke , Simeon Warner , Paul Ginsparg

MOTIVATION: The biological literature is a major repository of knowledge. Many biological databases draw much of their content from a careful curation of this literature. However, as the volume of literature increases, the burden of…

Computation and Language · Computer Science 2011-11-09 Alexander S. Yeh , Lynette Hirschman , Alexander A. Morgan

In the evolving field of Natural Language Processing (NLP), understanding the temporal context of text is increasingly critical for applications requiring advanced temporal reasoning. Traditional pre-trained language models like BERT, which…

Computation and Language · Computer Science 2025-03-06 Jiexin Wang , Adam Jatowt , Yi Cai

Large language models (LLMs) such as GPT, Claude, Gemini, and Grok have been deeply integrated into our daily life. They now support a wide range of tasks -- from dialogue and email drafting to assisting with teaching and coding, serving as…

Computation and Language · Computer Science 2026-01-13 Hongyi Zhou , Jin Zhu , Ying Yang , Chengchun Shi

Existing datasets available for crosslinguistic investigations have tended to focus on large amounts of data for a small group of languages or a small amount of data for a large number of languages. This means that claims based on these…

Computation and Language · Computer Science 2026-01-27 Hiram Ring

The amount of scholarly data has been increasing dramatically over the last years. For newcomers to a particular science domain (e.g., IR, physics, NLP) it is often difficult to spot larger trends and to position the latest research in the…

Digital Libraries · Computer Science 2021-12-08 Naman Paharia , Muhammad Syafiq Mohd Pozi , Adam Jatowt

Large language models (LLMs) have gained significant attention due to their ability to mimic human language. Identifying texts generated by LLMs is crucial for understanding their capabilities and mitigating potential consequences. This…

Computation and Language · Computer Science 2024-07-19 Anjali Rawal , Hui Wang , Youjia Zheng , Yu-Hsuan Lin , Shanu Sushmita

The problem of unveiling the author of a given text document from multiple candidate authors is called authorship attribution. Manifold word-based stylistic markers have been successfully used in deep learning methods to deal with the…

Computation and Language · Computer Science 2023-06-28 Abiodun Modupe , Turgay Celik , Vukosi Marivate , Oludayo O. Olugbara

Digital libraries store images which can be highly degraded and to index this kind of images we resort to word spot- ting as our information retrieval system. Information retrieval for handwritten document images is more challenging due to…

Computer Vision and Pattern Recognition · Computer Science 2016-04-22 Sounak Dey , Anguelos Nicolaou , Josep Llados , Umapada Pal

To prevent misinformation and social issues arising from trustworthy-looking content generated by LLMs, it is crucial to develop efficient and reliable methods for identifying the source of texts. Previous approaches have demonstrated…

Computation and Language · Computer Science 2025-12-03 Fangqi Dai , Xingjian Jiang , Zizhuang Deng

Accurate information retrieval (IR) is critical in the financial domain, where investors must identify relevant information from large collections of documents. Traditional IR methods -- whether sparse or dense -- often fall short in…

Recent strides in Large Language Models (LLMs) have saturated many Natural Language Processing (NLP) benchmarks, emphasizing the need for more challenging ones to properly assess LLM capabilities. However, domain-specific and multilingual…

Temporal awareness is crucial in many information retrieval tasks, particularly in scenarios where the relevance of documents depends on their alignment with the query's temporal context. Traditional approaches such as BM25 and Dense…

Information Retrieval · Computer Science 2025-04-09 Abdelrahman Abdallah , Bhawna Piryani , Jonas Wallat , Avishek Anand , Adam Jatowt

We introduce Strategic Doctrine Language Models (sdLM), a learning-system framework for multi-document strategic reasoning with doctrinal consistency constraints and calibrated uncertainty. The approach combines multi-document attention,…

Machine Learning · Computer Science 2026-01-22 Olaf Yunus Laitinen Imanov , Taner Yilmaz , Derya Umut Kulali

Large text corpora are increasingly important for a wide variety of Natural Language Processing (NLP) tasks, and automatic language identification (LangID) is a core technology needed to collect such datasets in a multilingual context.…

Computation and Language · Computer Science 2020-10-30 Isaac Caswell , Theresa Breiner , Daan van Esch , Ankur Bapna

This competition investigates the performance of large-scale retrieval of historical document images based on writing style. Based on large image data sets provided by cultural heritage institutions and digital libraries, providing a total…

Computer Vision and Pattern Recognition · Computer Science 2019-12-10 Vincent Christlein , Anguelos Nicolaou , Mathias Seuret , Dominique Stutzmann , Andreas Maier
‹ Prev 1 4 5 6 7 8 10 Next ›