English
Related papers

Related papers: Web Archive Analytics

200 papers

We perform a large-scale analysis of third-party trackers on the World Wide Web from more than 3.5 billion web pages of the CommonCrawl 2012 corpus. We extract a dataset containing more than 140 million third-party embeddings in over 41…

Social and Information Networks · Computer Science 2016-08-01 Sebastian Schelter , Jérôme Kunegis

Wikipedia is one of the most visited websites in the world and is also a frequent subject of scientific research. However, the analytical possibilities of Wikipedia information have not yet been analyzed considering at the same time both a…

Digital Libraries · Computer Science 2022-11-18 Wenceslao Arroyo-Machado , Daniel Torres-Salinas , Rodrigo Costas

The ability to programmatically retrieve vast quantities of data from online sources has given rise to increasing usage of web-scraped datasets for various purposes across government, industry and academia. Contemporaneously, there has also…

Digital Libraries · Computer Science 2025-11-19 Cynthia A. Huang , Tina Lam

Websites are capable of learning a wide range of information about the platform on which a browser is executing. One major source of such information is the set of standardised Application Programming Interfaces (APIs) provided within the…

Cryptography and Security · Computer Science 2019-10-17 Zhaoyi Fan

This paper presents a comprehensive analysis of global web usage patterns based on data from SimilarWeb, a leading source for estimating web traffic. Leveraging a dataset comprising over 250,000 websites, we estimate the total web traffic…

Computers and Society · Computer Science 2024-11-27 Henrique S. Xavier

The underlying data source for web usage mining (WUM) is commonly thought to be server logs. However, access log files ensure quite limited data about the clients. Identifying sessions from this messy data takes a considerable effort, and…

Information Retrieval · Computer Science 2025-01-09 Ozkan Canay , Umit Kocabicak

Web Engineering is the application of systematic, disciplined and quantifiable approaches to development, operation, and maintenance of Web-based applications. It is both a pro-active approach and a growing collection of theoretical and…

Software Engineering · Computer Science 2007-05-23 Yogesh Deshpande , San Murugesan , Athula Ginige , Steve Hansen , Daniel Schwabe , Martin Gaedke , Bebo White

Although user access patterns on the live web are well-understood, there has been no corresponding study of how users, both humans and robots, access web archives. Based on samples from the Internet Archive's public Wayback Machine, we…

Digital Libraries · Computer Science 2013-09-17 Yasmin AlNoamany , Michele C. Weigle , Michael L. Nelson

The World Wide Web is not only one of the most important platforms of communication and information at present, but also an area of growing interest for scientific research. This motivates a lot of work and projects that require large…

Computer Vision and Pattern Recognition · Computer Science 2021-05-18 Christian Mejia-Escobar , Miguel Cazorla , Ester Martinez-Martin

With the huge amount of information available online, the World Wide Web is a fertile area for data mining research. The Web mining research is at the cross road of research from several research communities, such as database, information…

Machine Learning · Computer Science 2007-05-23 Raymond Kosala , Hendrik Blockeel

Purpose: Advanced usage of Web Analytics tools allows to capture the content of user queries. Despite their relevant nature, the manual analysis of large volumes of user queries is problematic. This paper demonstrates the potential of using…

Information Retrieval · Computer Science 2017-09-25 Anne Chardonnens , Ettore Rizza , Mathias Coeckelbergs , Seth van Hooland

Aggregations of Web resources are increasingly important in scholarship as it adopts new methods that are data-centric, collaborative, and networked-based. The same notion of aggregations of resources is common to the mashed-up, socially…

Digital Libraries · Computer Science 2009-06-12 Herbert Van de Sompel , Carl Lagoze , Michael L. Nelson , Simeon Warner , Robert Sanderson , Pete Johnston

The Data Web refers to the vast and rapidly increasing quantity of scientific, corporate, government and crowd-sourced data published in the form of Linked Open Data, which encourages the uniform representation of heterogeneous data items…

The Web has been around and maturing for 25 years. The popular websites of today have undergone vast changes during this period, with a few being there almost since the beginning and many new ones becoming popular over the years. This makes…

Digital Libraries · Computer Science 2017-02-07 Helge Holzmann , Wolfgang Nejdl , Avishek Anand

Wikipedia is a rich and invaluable source of information. Its central place on the Web makes it a particularly interesting object of study for scientists. Researchers from different domains used various complex datasets related to Wikipedia…

Information Retrieval · Computer Science 2019-03-21 Nicolas Aspert , Volodymyr Miz , Benjamin Ricaud , Pierre Vandergheynst

Internet is one of the main sources of information for millions of people. One can find information related to practically all matters on internet. Moreover if we want to retrieve information about some particular topic we may find…

Information Retrieval · Computer Science 2012-10-01 Deepika Sharma , Deepak Garg

Digital computational outputs are now ubiquitous in the research workflow and the way in which these data are stored and cataloged is becoming more standardized across fields of research. However, even with accessible data and code, the…

Digital Libraries · Computer Science 2025-08-19 Sabar Dasgupta , Paul Nuyujukian

Public data archives are the backbone of modern biological and biomedical research. While archives for biological molecules and structures are well-established, resources for imaging data do not yet cover the full range of spatial and…

Quantitative Methods · Quantitative Biology 2018-11-05 Jan Ellenberg , Jason R Swedlow , Mary Barlow , Charles E Cook , Ardan Patwardhan , Alvis Brazma , Ewan Birney

Log files contain information about User Name, IP Address, Time Stamp, Access Request, number of Bytes Transferred, Result Status, URL that Referred and User Agent. The log files are maintained by the web servers. By analysing these log…

Databases · Computer Science 2011-02-01 L. K. Joshila Grace , V. Maheswari , Dhinaharan Nagamalai

Web information extraction (WIE) is the task of automatically extracting data from web pages, offering high utility for various applications. The evaluation of WIE systems has traditionally relied on benchmarks built from HTML snapshots…

Computation and Language · Computer Science 2026-03-17 Seungbin Yang , Jihwan Kim , Jaemin Choi , Dongjin Kim , Soyoung Yang , ChaeHun Park , Jaegul Choo
‹ Prev 1 3 4 5 6 7 10 Next ›