English
Related papers

Related papers: Longitudinal Sampling of URLs From the Wayback Mac…

200 papers

World Wide Web is a huge data repository and is growing with the explosive rate of about 1 million pages a day. As the information available on World Wide Web is growing the usage of the web sites is also growing. Web log records each…

Information Retrieval · Computer Science 2009-08-03 Ratnesh Kumar Jain , Dr. R. S. Kasana , Dr. Suresh Jain

Many emerging Web services, such as email, photo sharing, and web site archives, need to preserve large amounts of quickly-accessible data indefinitely into the future. In this paper, we make the case that these applications' demands on…

Digital Libraries · Computer Science 2007-05-23 Mary Baker , Mehul Shah , David S. H. Rosenthal , Mema Roussopoulos , Petros Maniatis , TJ Giuli , Prashanth Bungale

As part of their efforts to consistently and reliably identify and preserve the archival records of their organizations, archivists are trying to address the issue of appraising and collecting World Wide Web documents. This paper discusses…

History and Philosophy of Physics · Physics 2007-05-23 Jean M. Deken

As part of its scholarly data efforts, the Internet Archive (IA) releases a first version of a citation graph dataset, named refcat, derived from scholarly publications and additional data sources. It is composed of data gathered by the…

Digital Libraries · Computer Science 2021-10-15 Martin Czygan , Helge Holzmann , Bryan Newbold

Users' detailed browsing activity - such as what sites they are spending time on and for how long, and what tabs they have open and which one is focused at any given time - is useful for a number of research and practical applications.…

Human-Computer Interaction · Computer Science 2021-02-09 Geza Kovacs

As web archives' holdings grow, archivists subdivide them into collections so they are easier to understand and manage. In this work, we review the collection structures of eight web archive platforms: : Archive-It, Conifer, the Croatian…

Nowadays, the huge amount of information distributed through the Web motivates studying techniques to be adopted in order to extract relevant data in an efficient and reliable way. Both academia and enterprises developed several approaches…

Artificial Intelligence · Computer Science 2013-06-06 Emilio Ferrara , Robert Baumgartner

Established in 2005, YouTube has become the most successful Internet site providing a new generation of short video sharing service. Today, YouTube alone comprises approximately 20% of all HTTP traffic, or nearly 10% of all traffic on the…

Networking and Internet Architecture · Computer Science 2007-07-26 Xu Cheng , Cameron Dale , Jiangchuan Liu

URLs are central to a myriad of cyber-security threats, from phishing to the distribution of malware. Their inherent ease of use and familiarity is continuously abused by attackers to evade defences and deceive end-users. Seemingly…

Cryptography and Security · Computer Science 2021-08-31 Mahathir Almashor , Ejaz Ahmed , Benjamin Pick , Sharif Abuadbba , Raj Gaire , Seyit Camtepe , Surya Nepal

As defined by the Memento Framework, TimeMaps are ma-chine-readable lists of time-specific copies -- called "mementos" -- of an archived original resource. In theory, as an archive acquires additional mementos over time, a TimeMap should be…

Digital Libraries · Computer Science 2013-07-23 Justin F. Brunelle , Michael L. Nelson

Session length is a very important aspect in determining a user's satisfaction with a media streaming service. Being able to predict how long a session will last can be of great use for various downstream tasks, such as recommendations and…

Information Retrieval · Computer Science 2017-08-02 Theodore Vasiloudis , Hossein Vahabi , Ross Kravitz , Valery Rashkov

The arXiv has collected 1.5 million pre-print articles over 28 years, hosting literature from scientific fields including Physics, Mathematics, and Computer Science. Each pre-print features text, figures, authors, citations, categories, and…

Information Retrieval · Computer Science 2019-05-02 Colin B. Clement , Matthew Bierbaum , Kevin P. O'Keeffe , Alexander A. Alemi

Domain probe lists--used to determine which URLs to probe for Web censorship--play a critical role in Internet censorship measurement studies. Indeed, the size and accuracy of the domain probe list limits the set of censored pages that can…

Cryptography and Security · Computer Science 2024-07-12 Jenny Tang , Leo Alvarez , Arjun Brar , Nguyen Phong Hoang , Nicolas Christin

With the large-scale penetration of the internet, for the first time, humanity has become linked by a single, open, communications platform. Harnessing this fact, we report insights arising from a unified internet activity and location…

Economics · Quantitative Finance 2017-01-23 Klaus Ackermann , Simon D Angus , Paul A Raschky

GoogleTrendArchive is a comprehensive archive of Google Trending Now data spanning over one year (from November 28, 2024 to January 3, 2026) across 125 countries and 1,358 locations. Unlike Google Trends, which requires specifying search…

Information Retrieval · Computer Science 2026-03-24 Aleksandra Urman , Anikó Hannák , Joachim Baumann

We perform a large-scale analysis of third-party trackers on the World Wide Web from more than 3.5 billion web pages of the CommonCrawl 2012 corpus. We extract a dataset containing more than 140 million third-party embeddings in over 41…

Social and Information Networks · Computer Science 2016-08-01 Sebastian Schelter , Jérôme Kunegis

As Digital Libraries (DL) become more aligned with the web architecture, their functional components need to be fundamentally rethought in terms of URIs and HTTP. Annotation, a core scholarly activity enabled by many DL solutions, exhibits…

Digital Libraries · Computer Science 2010-03-22 Robert Sanderson , Herbert Van de Sompel

Active Internet measurement studies rely on a list of targets to be scanned. While probing the entire IPv4 address space is feasible for scans of limited complexity, more complex scans do not scale to measuring the full Internet. Thus, a…

Networking and Internet Architecture · Computer Science 2018-02-09 Quirin Scheitle , Jonas Jelten , Oliver Hohlfeld , Luca Ciprian , Georg Carle

Log files contain information about User Name, IP Address, Time Stamp, Access Request, number of Bytes Transferred, Result Status, URL that Referred and User Agent. The log files are maintained by the web servers. By analysing these log…

Databases · Computer Science 2011-02-01 L. K. Joshila Grace , V. Maheswari , Dhinaharan Nagamalai

Cross-lingual document alignment aims to identify pairs of documents in two distinct languages that are of comparable content or translations of each other. In this paper, we exploit the signals embedded in URLs to label web documents at…

Computation and Language · Computer Science 2020-10-13 Ahmed El-Kishky , Vishrav Chaudhary , Francisco Guzman , Philipp Koehn
‹ Prev 1 3 4 5 6 7 10 Next ›