English
Related papers

Related papers: Temporally Extending Existing Web Archive Collecti…

200 papers

This paper presents the third edition of the LongEval Lab, part of the CLEF 2025 conference, which continues to explore the challenges of temporal persistence in Information Retrieval (IR). The lab features two tasks designed to provide…

We here introduce a substantially extended version of JeSemE, an interactive website for visually exploring computationally derived time-variant information on word meanings and lexical emotions assembled from five large diachronic text…

Computation and Language · Computer Science 2020-12-21 Johannes Hellrich , Sven Buechel , Udo Hahn

The motivation, concept, design and implementation of latent semantic search for search engines have limited semantic search, entity extraction and property attribution features, have insufficient accuracy and response time of latent…

Information Retrieval · Computer Science 2019-12-03 Anton Kolonin

Sentence-by-sentence information extraction from long documents is an exhausting and error-prone task. As the indicator of document skeleton, catalogs naturally chunk documents into segments and provide informative cascade semantics, which…

Computation and Language · Computer Science 2023-05-01 Tong Zhu , Guoliang Zhang , Zechang Li , Zijian Yu , Junfei Ren , Mengsong Wu , Zhefeng Wang , Baoxing Huai , Pingfu Chao , Wenliang Chen

Wikipedia serves as a key infrastructure for public access to scientific knowledge, but it faces challenges in maintaining the credibility of cited sources--especially when scientific papers are retracted. This paper investigates how…

Human-Computer Interaction · Computer Science 2026-01-27 Haohan Shi , Yulin Yu , Daniel M. Romero , Emőke-Ágnes Horvát

This paper introduces a large collection of time series data derived from Twitter, postprocessed using word embedding techniques, as well as specialized fine-tuned language models. This data comprises the past five years and captures…

Computation and Language · Computer Science 2023-08-07 Daniel Loureiro , Kiamehr Rezaee , Talayeh Riahi , Francesco Barbieri , Leonardo Neves , Luis Espinosa Anke , Jose Camacho-Collados

In large-scale construction projects, the continuous evolution of decisions generates extensive records, most often captured in meeting minutes. Since decisions may override previous ones, professionals often need to reconstruct the history…

Computation and Language · Computer Science 2026-04-17 Ioannis-Aris Kostis , Natalia Sanchiz , Steeve De Schryver , François Denis , Pierre Schaus

This paper proposes temporally aligned Large Language Models (LLMs) as a tool for longitudinal analysis of social media data. We fine-tune Temporal Adapters for Llama 3 8B on full timelines from a panel of British Twitter users, and extract…

Computers and Society · Computer Science 2025-10-14 Georg Ahnert , Max Pellert , David Garcia , Markus Strohmaier

In our continuously evolving world, entities change over time and new, previously non-existing or unknown, entities appear. We study how this evolutionary scenario impacts the performance on a well established entity linking (EL) task. For…

Computation and Language · Computer Science 2023-02-07 Klim Zaporojets , Lucie-Aimee Kaffee , Johannes Deleu , Thomas Demeester , Chris Develder , Isabelle Augenstein

Web archive collections are created with a particular purpose in mind. A curator selects seeds, or original resources, which are then captured by an archiving system and stored as archived web pages, or mementos. The systems that build web…

Digital Libraries · Computer Science 2021-01-26 Shawn M. Jones , Michele C. Weigle , Michael L. Nelson

The historical, cultural, and intellectual importance of archiving the web has been widely recognized. Today, all countries with high Internet penetration rate have established high-profile archiving initiatives to crawl and archive the…

Digital Libraries · Computer Science 2013-08-13 Zhiwu Xie , Herbert Van de Sompel , Jinyang Liu , Johann van Reenen , Ramiro Jordan

The work on large-scale graph analytics to date has largely focused on the study of static properties of graph snapshots. However, a static view of interactions between entities is often an oversimplification of several complex phenomena…

Databases · Computer Science 2015-10-01 Udayan Khurana , Amol Deshpande

We explore the availability and persistence of URLs cited in articles published in D-Lib Magazine. We extracted 4387 unique URLs referenced in 453 articles published from July 1995 to August 2004. The availability was checked three times a…

Digital Libraries · Computer Science 2011-11-09 Frank McCown , Sheffan Chan , Michael L. Nelson , Johan Bollen

In the last few years, open-domain question answering (ODQA) has advanced rapidly due to the development of deep learning techniques and the availability of large-scale QA datasets. However, the current datasets are essentially designed for…

Computation and Language · Computer Science 2022-02-23 Jiexin Wang , Adam Jatowt , Masatoshi Yoshikawa

In this Digest we review a recent study released by the Google Scholar team on the apparently increasing fraction of citations to old articles from studies published in the last 24 years (1990-2013). First, we describe the main findings of…

The enormous amount of discourse taking place online poses challenges to the functioning of a civil and informed public sphere. Efforts to standardize online discourse data, such as ClaimReview, are making available a wealth of new data…

Social and Information Networks · Computer Science 2026-05-19 Matthew Sumpter , Giovanni Luca Ciampaglia

Differentiable Search Index is a recently proposed paradigm for document retrieval, that encodes information about a corpus of documents within the parameters of a neural network and directly maps queries to corresponding documents. These…

Information Retrieval · Computer Science 2024-08-20 Varsha Kishore , Chao Wan , Justin Lovelace , Yoav Artzi , Kilian Q. Weinberger

Determining the sustainability impact of companies is a highly complex subject which has garnered more and more attention over the past few years. Today, investors largely rely on sustainability-ratings from established rating-providers in…

Information Retrieval · Computer Science 2024-12-20 Fabian Billert , Stefan Conrad

We analyze the online response to the preprint publication of a cohort of 4,606 scientific articles submitted to the preprint database arXiv.org between October 2010 and May 2011. We study three forms of responses to these preprints:…

Social and Information Networks · Computer Science 2015-06-04 Xin Shuai , Alberto Pepe , Johan Bollen

Working with Web archives raises a number of issues caused by their temporal characteristics. Depending on the age of the content, additional knowledge might be needed to find and understand older texts. Especially facts about entities are…

Computation and Language · Computer Science 2017-02-07 Helge Holzmann , Thomas Risse