English
Related papers

Related papers: Longitudinal Sampling of URLs From the Wayback Mac…

200 papers

arXiv is the largest open-access repository for scientific literature. When submitting a paper, authors upload the manuscript's source files, from which the final PDF is compiled. These source files are also publicly downloadable,…

Networking and Internet Architecture · Computer Science 2026-01-19 Giovanni Apruzzese , Aurore Fass

The collapse of social contexts has been amplified by digital infrastructures but surprisingly received insufficient attention from Web privacy scholars. Users are persistently identified within and across distinct Web contexts, in varying…

Cryptography and Security · Computer Science 2025-03-03 Ido Sivan-Sevilla , Parthav Poudel

Twitter is among the commonest sources of data employed in social media research mainly because of its convenient APIs to collect tweets. However, most researchers do not have access to the expensive Firehose and Twitter Historical Archive,…

Computers and Society · Computer Science 2016-11-28 Daniel Gayo-Avello

The Web is a tangled mass of interconnected services, where websites import a range of external resources from various third-party domains. However, the latter can further load resources hosted on other domains. For each website, this…

Cryptography and Security · Computer Science 2019-02-19 Muhammad Ikram , Rahat Masood , Gareth Tyson , Mohamed Ali Kaafar , Noha Loizon , Roya Ensafi

In 1998, the Getty Center hosted the ''Time and Bits: Managing Digital Continuity'' conference, gathering the founders and thinkers of two San Francisco non-profit organizations interested in long-term thinking and archiving: the Internet…

Digital Libraries · Computer Science 2024-12-12 Julie Momméja

The focused web-harvesting is deployed to realize an automated and comprehensive index databases as an alternative way for virtual topical data integration. The web-harvesting has been implemented and extended by not only specifying the…

Information Retrieval · Computer Science 2008-09-05 Z. Akbar , L. T. Handoko

User tracking on the Internet can come in various forms, e.g., via cookies or by fingerprinting web browsers. A technique that got less attention so far is user tracking based on TLS and specifically based on the TLS session resumption…

Cryptography and Security · Computer Science 2019-03-01 Erik Sy , Christian Burkert , Hannes Federrath , Mathias Fischer

In a Web plagued by disappearing resources, Web archive collections provide a valuable means of preserving Web resources important to the study of past events ranging from elections to disease outbreaks. These archived collections start…

Digital Libraries · Computer Science 2019-05-30 Alexander C. Nwala , Michele C. Weigle , Michael L. Nelson

Web API specifications are machine-readable descriptions of APIs. These specifications, in combination with related tooling, simplify and support the consumption of APIs. However, despite the increased distribution of web APIs,…

Software Engineering · Computer Science 2018-01-29 Jinqiu Yang , Erik Wittern , Annie T. T. Ying , Julian Dolby , Lin Tan

The Data Web refers to the vast and rapidly increasing quantity of scientific, corporate, government and crowd-sourced data published in the form of Linked Open Data, which encourages the uniform representation of heterogeneous data items…

Malicious URL, a.k.a. malicious website, is a common and serious threat to cybersecurity. Malicious URLs host unsolicited content (spam, phishing, drive-by exploits, etc.) and lure unsuspecting users to become victims of scams (monetary…

Machine Learning · Computer Science 2019-08-22 Doyen Sahoo , Chenghao Liu , Steven C. H. Hoi

Web archives are a valuable resource for researchers of various disciplines. However, to use them as a scholarly source, researchers require a tool that provides efficient access to Web archive data for extraction and derivation of smaller…

Digital Libraries · Computer Science 2017-02-06 Helge Holzmann , Vinay Goel , Avishek Anand

Purpose: To provide a critical review of Bergman's 2001 study on the Deep Web. In addition, we bring a new concept into the discussion, the Academic Invisible Web (AIW). We define the Academic Invisible Web as consisting of all databases…

Digital Libraries · Computer Science 2019-01-15 Dirk Lewandowski , Philipp Mayr

In recent years, journalists and other researchers have used web archives as an important resource for their study of disinformation. This paper provides several examples of this use and also brings together some of the work that the Old…

Digital Libraries · Computer Science 2023-06-19 Michele C. Weigle

We propose a new large-scale (nearly a million questions) ultra-long-context (more than 50,000 words average document length) reading comprehension dataset. Using GPT 3.5, we summarized each scene in 1,500 hand-curated fiction books from…

Computation and Language · Computer Science 2023-12-11 Arseny Moskvichev , Ky-Vinh Mai

Event collections are frequently built by crawling the live web on the basis of seed URIs nominated by human experts. Focused web crawling is a technique where the crawler is guided by reference content pertaining to the event. Given the…

Digital Libraries · Computer Science 2018-04-06 Martin Klein , Lyudmila Balakireva , Herbert Van de Sompel

Internet-based personal digital belongings present different vulnerabilities than locally stored materials. We use responses to a survey of people who have recovered lost websites, in combination with supplementary interviews, to paint a…

Digital Libraries · Computer Science 2007-05-23 Catherine C. Marshall , Frank McCown , Michael L. Nelson

In the mid-90's, it was shown that the statistics of aggregated time series from Internet traffic departed from those of traditional short range dependent models, and were instead characterized by asymptotic self-similarity. Following this…

Networking and Internet Architecture · Computer Science 2017-03-07 Romain Fontugne , Patrice Abry , Kensuke Fukuda , Darryl Veitch , Kenjiro Cho , Pierre Borgnat , Herwig Wendt

In this paper, we analyze the nature and distribution of structured data on the Web. Web-scale information extraction, or the problem of creating structured tables using extraction from the entire web, is gathering lots of research…

Databases · Computer Science 2012-03-30 Nilesh Dalvi , Ashwin Machanavajjhala , Bo Pang

Large language models (LLMs) rely heavily on web-scale datasets like Common Crawl, which provides over 80\% of training data for some modern models. However, the indiscriminate nature of web crawling raises challenges in data quality,…

Computation and Language · Computer Science 2025-09-01 Inés Altemir Marinas , Anastasiia Kucherenko , Andrei Kucharavy