中文
相关论文

相关论文: BullingerDB: A Dataset for Handwritten Text Recogn…

200 篇论文

Identifying the type of font (e.g., Roman, Blackletter) used in historical documents can help optical character recognition (OCR) systems produce more accurate text transcriptions. Towards this end, we present an active-learning strategy…

计算机视觉与模式识别 · 计算机科学 2016-01-28 Anshul Gupta , Ricardo Gutierrez-Osuna , Matthew Christy , Richard Furuta , Laura Mandell

Alterations in historical manuscripts such as letters represent a promising field of research. On the one hand, they help understand the construction of text. On the other hand, topics that are being considered sensitive at the time of the…

机器学习 · 计算机科学 2020-11-05 David Lassner , Anne Baillot , Sergej Dogadov , Klaus-Robert Müller , Shinichi Nakajima

Albeit Natural Language Processing has seen major breakthroughs in the last few years, transferring such advances into real-world business cases can be challenging. One of the reasons resides in the displacement between popular benchmarks…

计算与语言 · 计算机科学 2024-02-16 Andrea Zugarini , Andrew Zamai , Marco Ernandes , Leonardo Rigutini

Disentangled representation learning remains challenging as the underlying factors of variation in the data do not naturally exist. The inherent complexity of real-world data makes it unfeasible to exhaustively enumerate and encapsulate all…

计算与语言 · 计算机科学 2024-02-13 Jiawei Zhou , Xiaoguang Li , Lifeng Shang , Xin Jiang , Qun Liu , Lei Chen

We describe a large-scale application of methods for finding plagiarism in research document collections. The methods are applied to a collection of 284,834 documents collected by arXiv.org over a 14 year period, covering a few different…

数据库 · 计算机科学 2007-05-23 Daria Sorokina , Johannes Gehrke , Simeon Warner , Paul Ginsparg

MOTIVATION: The biological literature is a major repository of knowledge. Many biological databases draw much of their content from a careful curation of this literature. However, as the volume of literature increases, the burden of…

计算与语言 · 计算机科学 2011-11-09 Alexander S. Yeh , Lynette Hirschman , Alexander A. Morgan

In the evolving field of Natural Language Processing (NLP), understanding the temporal context of text is increasingly critical for applications requiring advanced temporal reasoning. Traditional pre-trained language models like BERT, which…

计算与语言 · 计算机科学 2025-03-06 Jiexin Wang , Adam Jatowt , Yi Cai

Large language models (LLMs) such as GPT, Claude, Gemini, and Grok have been deeply integrated into our daily life. They now support a wide range of tasks -- from dialogue and email drafting to assisting with teaching and coding, serving as…

计算与语言 · 计算机科学 2026-01-13 Hongyi Zhou , Jin Zhu , Ying Yang , Chengchun Shi

Existing datasets available for crosslinguistic investigations have tended to focus on large amounts of data for a small group of languages or a small amount of data for a large number of languages. This means that claims based on these…

计算与语言 · 计算机科学 2026-01-27 Hiram Ring

The amount of scholarly data has been increasing dramatically over the last years. For newcomers to a particular science domain (e.g., IR, physics, NLP) it is often difficult to spot larger trends and to position the latest research in the…

数字图书馆 · 计算机科学 2021-12-08 Naman Paharia , Muhammad Syafiq Mohd Pozi , Adam Jatowt

Large language models (LLMs) have gained significant attention due to their ability to mimic human language. Identifying texts generated by LLMs is crucial for understanding their capabilities and mitigating potential consequences. This…

计算与语言 · 计算机科学 2024-07-19 Anjali Rawal , Hui Wang , Youjia Zheng , Yu-Hsuan Lin , Shanu Sushmita

The problem of unveiling the author of a given text document from multiple candidate authors is called authorship attribution. Manifold word-based stylistic markers have been successfully used in deep learning methods to deal with the…

计算与语言 · 计算机科学 2023-06-28 Abiodun Modupe , Turgay Celik , Vukosi Marivate , Oludayo O. Olugbara

Digital libraries store images which can be highly degraded and to index this kind of images we resort to word spot- ting as our information retrieval system. Information retrieval for handwritten document images is more challenging due to…

计算机视觉与模式识别 · 计算机科学 2016-04-22 Sounak Dey , Anguelos Nicolaou , Josep Llados , Umapada Pal

To prevent misinformation and social issues arising from trustworthy-looking content generated by LLMs, it is crucial to develop efficient and reliable methods for identifying the source of texts. Previous approaches have demonstrated…

计算与语言 · 计算机科学 2025-12-03 Fangqi Dai , Xingjian Jiang , Zizhuang Deng

Accurate information retrieval (IR) is critical in the financial domain, where investors must identify relevant information from large collections of documents. Traditional IR methods -- whether sparse or dense -- often fall short in…

Recent strides in Large Language Models (LLMs) have saturated many Natural Language Processing (NLP) benchmarks, emphasizing the need for more challenging ones to properly assess LLM capabilities. However, domain-specific and multilingual…

Temporal awareness is crucial in many information retrieval tasks, particularly in scenarios where the relevance of documents depends on their alignment with the query's temporal context. Traditional approaches such as BM25 and Dense…

信息检索 · 计算机科学 2025-04-09 Abdelrahman Abdallah , Bhawna Piryani , Jonas Wallat , Avishek Anand , Adam Jatowt

We introduce Strategic Doctrine Language Models (sdLM), a learning-system framework for multi-document strategic reasoning with doctrinal consistency constraints and calibrated uncertainty. The approach combines multi-document attention,…

机器学习 · 计算机科学 2026-01-22 Olaf Yunus Laitinen Imanov , Taner Yilmaz , Derya Umut Kulali

Large text corpora are increasingly important for a wide variety of Natural Language Processing (NLP) tasks, and automatic language identification (LangID) is a core technology needed to collect such datasets in a multilingual context.…

计算与语言 · 计算机科学 2020-10-30 Isaac Caswell , Theresa Breiner , Daan van Esch , Ankur Bapna

This competition investigates the performance of large-scale retrieval of historical document images based on writing style. Based on large image data sets provided by cultural heritage institutions and digital libraries, providing a total…

计算机视觉与模式识别 · 计算机科学 2019-12-10 Vincent Christlein , Anguelos Nicolaou , Mathias Seuret , Dominique Stutzmann , Andreas Maier