English
Related papers

Related papers: Access Patterns for Robots and Humans in Web Archi…

200 papers

Search engines provide cached copies of indexed content so users will have something to "click on" if the remote resource is temporarily or permanently unavailable. Depending on their proprietary caching strategies, search engines will…

Digital Libraries · Computer Science 2007-05-23 Frank McCown , Michael L. Nelson

Our experience of web access slowing down is a consequence of the aggregated web access pattern of web users. This is just one example among several human oriented services which are strongly affected by human activity patterns. Recent…

Physics and Society · Physics 2009-11-11 Alexei Vazquez

The field of web archiving provides a unique mix of human and automated agents collaborating to achieve the preservation of the web. Centuries old theories of archival appraisal are being transplanted into the sociotechnical environment of…

Digital Libraries · Computer Science 2016-11-09 Ed Summers , Ricardo Punzalan

When a user views an archived page using the archive's user interface (UI), the user selects a datetime to view from a list. The archived web page, if available, is then displayed. From this display, the web archive UI attempts to simulate…

Digital Libraries · Computer Science 2013-09-24 Scott G. Ainsworth , Michael L. Nelson

The increasing adoption of econometric and machine-learning approaches by empirical researchers has led to a widespread use of one data collection method: web scraping. Web scraping refers to the use of automated computer programs to access…

General Economics · Economics 2023-08-07 Jens Foerderer

A promising approach to solving challenging long-horizon tasks has been to extract behavior priors (skills) by fitting generative models to large offline datasets of demonstrations. However, such generative models inherit the biases of the…

Machine Learning · Computer Science 2021-08-13 Xiaofei Wang , Kimin Lee , Kourosh Hakhamaneshi , Pieter Abbeel , Michael Laskin

As the exploration of digital behavioral data revolutionizes communication research, understanding the nuances of data collection methodologies becomes increasingly pertinent. This study focuses on one prominent data collection approach,…

Computers and Society · Computer Science 2024-12-03 Roberto Ulloa , Frank Mangold , Felix Schmidt , Judith Gilsbach , Sebastian Stier

World Wide Web is a huge repository of web pages and links. It provides abundance of information for the Internet users. The growth of web is tremendous as approximately one million pages are added daily. Users' accesses are recorded in web…

Information Retrieval · Computer Science 2010-04-09 V. Chitraa , Dr. Antony Selvdoss Davamani

Significant parts of cultural heritage are produced on the web during the last decades. While easy accessibility to the current web is a good baseline, optimal access to the past web faces several challenges. This includes dealing with…

Digital Libraries · Computer Science 2017-01-31 Nattiya Kanhabua , Philipp Kemkes , Wolfgang Nejdl , Tu Ngoc Nguyen , Felipe Reis , Nam Khanh Tran

As "Big Data" has become pervasive, an increasing amount of research has connected the dots between human behaviour in the offline and online worlds. Consequently, researchers have exploited these new findings to create models that better…

Computers and Society · Computer Science 2020-02-14 Marco De Nadai , Bruno Lepri , Nuria Oliver

Despite the importance and pervasiveness of Wikipedia as one of the largest platforms for open knowledge, surprisingly little is known about how people navigate its content when seeking information. To bridge this gap, we present the first…

Computers and Society · Computer Science 2023-01-19 Tiziano Piccardi , Martin Gerlach , Akhil Arora , Robert West

As defined by the Memento Framework, TimeMaps are ma-chine-readable lists of time-specific copies -- called "mementos" -- of an archived original resource. In theory, as an archive acquires additional mementos over time, a TimeMap should be…

Digital Libraries · Computer Science 2013-07-23 Justin F. Brunelle , Michael L. Nelson

Web browsers are a portal to the internet, where much of human activity is undertaken. Thus, there has been significant research work in AI agents that interact with the internet through web browsing. However, there is also another…

Computation and Language · Computer Science 2025-06-18 Yueqi Song , Frank Xu , Shuyan Zhou , Graham Neubig

Accurately analyzing and modeling online browsing behavior play a key role in understanding users and technology interactions. In this work, we design and conduct a user study to collect browsing data from 31 participants continuously for…

Computers and Society · Computer Science 2021-08-17 Yuliia Lut , Michael Wang , Elissa M. Redmiles , Rachel Cummings

The World Wide Web is the most wide known information source that is easily available and searchable. It consists of billions of interconnected documents Web pages are authored by millions of people. Accesses made by various users to pages…

Databases · Computer Science 2014-08-26 Priyanka Verma , Nishtha Kesswani

Understanding how people interact with the web is key for a variety of applications, e.g., from the design of effective web pages to the definition of successful online marketing campaigns. Browsing behavior has been traditionally…

Computers and Society · Computer Science 2021-05-05 Luca Vassio , Idilio Drago , Marco Mellia , Zied Ben Houidi , Mohamed Lamine Lamali

A mass of traces of human activities show diverse dynamic patterns. In this paper, we comprehensively investigate the dynamic pattern of human attention defined by the quantity of interests on subdisciplines in an online academic…

Physics and Society · Physics 2017-11-29 Zhi-Dan Zhao , Ya-Chun Gao , Shi-Min Cai

With the World Wide Web's ubiquity increase and the rapid development of various online businesses, the complexity of web sites grow. The analysis of web user's navigational pattern within a web site can provide useful information for…

Information Retrieval · Computer Science 2008-03-07 Biswajit Biswal

Web archives, query and proxy logs, and so on, can all be very large and highly repetitive; and are accessed only sporadically and partially, rather than continually and holistically. This type of data is ideal for compression-based…

Information Theory · Computer Science 2016-03-01 Matthias Petri , Alistair Moffat , P. C. Nagesh , Anthony Wirth

As the World Wide Web is growing rapidly, it is getting increasingly challenging to gather representative information about it. Instead of crawling the web exhaustively one has to resort to other techniques like sampling to determine the…

Data Structures and Algorithms · Computer Science 2009-02-11 Eda Baykan , Monika Henzinger , Stefan F. Keller , Sebastian De Castelberg , Markus Kinzler