English
Related papers

Related papers: The Ephemeral Web and the Case for Proactive Archi…

200 papers

Recent progress in large language models (LLMs) has enabled the development of autonomous web agents capable of navigating and interacting with real websites. However, evaluating such agents remains challenging due to the instability and…

Information Retrieval · Computer Science 2025-08-14 Zihao Sun , Ling Chen

Personal and private Web archives are proliferating due to the increase in the tools to create them and the realization that Internet Archive and other public Web archives are unable to capture personalized (e.g., Facebook) and private…

Digital Libraries · Computer Science 2018-06-05 Mat Kelly , Michael L. Nelson , Michele C. Weigle

The use of the internet, and in particular web browsing, offers many potential advantages for educational institutions as students have access to a wide range of information previously not available. However, there are potential negative…

Social and Information Networks · Computer Science 2011-10-31 Scott Hazelhurst , Yestin Johnson , Ian Sanders

We introduce the Tidal Tales Plugin, a Firefox extension for efficiently collecting and archiving of Instagram stories, addressing the challenges of ephemeral data in social media research. It enables an automated collection of story…

Social and Information Networks · Computer Science 2024-09-04 Michael Achmann-Denkler , Christian Wolff

Crossing multiple planetary boundaries places us in a zone of uncertainty that is characterized by considerable fluctuations in climatic events. The situation is exacerbated by the relentless use of resources and energy required to develop…

Computers and Society · Computer Science 2025-08-13 Olivier Michel , Emilie Frenkiel

Auditing differential privacy has emerged as an important area of research that supports the design of privacy-preserving mechanisms. Privacy audits help to obtain empirical estimates of the privacy parameter, to expose flawed…

Cryptography and Security · Computer Science 2025-09-25 Önder Askin , Tim Kutta , Holger Dette

Scholarly resources, just like any other resources on the web, are subject to reference rot as they frequently disappear or significantly change over time. Digital Object Identifiers (DOIs) are commonplace to persistently identify scholarly…

Digital Libraries · Computer Science 2020-04-08 Martin Klein , Lyudmila Balakireva

We consider a task of scheduling a crawler to retrieve content from several sites with ephemeral content. A user typically loses interest in ephemeral content, like news or posts at social network groups, after several days or hours. Thus,…

Information Retrieval · Computer Science 2015-03-31 Konstantin Avrachenkov , Vivek Borkar

This text advances the hypothesis that the meaning of the Web as an object of study has diluted as a clear research domain. One example of this phenomenon is the identity crisis of the Web Conference and the International Semantic Web…

Computers and Society · Computer Science 2026-04-15 Claudio Gutierrez , Daniel Hernández

The reliability and success of any organization such as academic institution rely on its ability to provide secure, accurate and timely data about its operations. Erstwhile managing student information in academic institution was done…

Computers and Society · Computer Science 2022-11-02 Oluwatosin Samuel Falebita

Archives play a crucial role in the construction and advancement of society. Humans place a great deal of trust in archives and depend on them to craft public policies and to preserve languages, cultures, self-identity, views and values.…

Computers and Society · Computer Science 2020-08-12 Abhishek Gupta , Nikitasha Kapoor

The evolution analysis on Web service ecosystems has become a critical problem as the frequency of service changes on the Internet increases rapidly. Developers need to understand these evolution patterns to assist in their decision-making…

Software Engineering · Computer Science 2021-08-30 Mingyi Liu , Zhiying Tu , Yeqi Zhu , Xiaofei Xu , Zhongjie Wang , Quan Z. Sheng

Common Crawl is a multi-petabyte longitudinal dataset containing over 100 billion web pages which is widely used as a source of language data for sequence model training and in web science research. Each of its constituent archives is on…

Networking and Internet Architecture · Computer Science 2024-04-16 Henry S. Thompson

The web constitutes a complex infrastructure and as demonstrated by numerous attacks, rigorous analysis of standards and web applications is indispensable. Inspired by successful prior work, in particular the work by Akhawe et al. as well…

Cryptography and Security · Computer Science 2019-01-31 Daniel Fett , Ralf Kuesters , Guido Schmitz

The arXiv is the most popular preprint repository in the world. Since its inception in 1991, the arXiv has allowed researchers to freely share publication-ready articles prior to formal peer review. The growth and the popularity of the…

Digital Libraries · Computer Science 2017-09-22 Alberto Pepe , Matteo Cantiello , Josh Nicholson

GitHub natively supports workflow automation through GitHub Actions. Yet, workflow maintenance is often considered a burden for software developers, who frequently face difficulties in writing, testing, debugging, and maintaining workflows.…

Software Engineering · Computer Science 2026-04-13 Hassan Onsori Delicheh , Guillaume Cardoen , Alexandre Decan , Tom Mens

Science projects are data publishers. The scale and complexity of current and future science data changes the nature of the publication process. Publication is becoming a major project component. At a minimum, a project must preserve the…

Digital Libraries · Computer Science 2015-06-25 Jim Gray , Alexander S. Szalay , Ani R. Thakar , Christopher Stoughton , Jan vandenBerg

In high energy physics, scholarly papers circulate primarily through online preprint archives based on a centralized repository, arXiv, that physicists simply refer to as "the archive". This is not just a tool for preservation and memory,…

Physics and Society · Physics 2016-06-24 Alessandro Delfanti

Analysis pipelines commonly use high-level technologies that are popular when created, but are unlikely to be readable, executable, or sustainable in the long term. A set of criteria is introduced to address this problem: Completeness (no…

As Web sites are now ordinary products, it is necessary to explicit the notion of quality of a Web site. The quality of a site may be linked to the easiness of accessibility and also to other criteria such as the fact that the site is up to…

Information Retrieval · Computer Science 2007-05-23 Thierry Despeyroux
‹ Prev 1 3 4 5 6 7 10 Next ›