English
Related papers

Related papers: A Brief History of Web Crawlers

200 papers

Web crawl is a main source of large language models' (LLMs) pretraining data, but the majority of crawled web pages are discarded in pretraining due to low data quality. This paper presents Craw4LLM, an efficient web crawling method that…

Computation and Language · Computer Science 2025-06-24 Shi Yu , Zhiyuan Liu , Chenyan Xiong

Automated analysis of privacy policies has proved a fruitful research direction, with developments such as automated policy summarization, question answering systems, and compliance detection. Prior research has been limited to analysis of…

Computers and Society · Computer Science 2021-07-22 Ryan Amos , Gunes Acar , Eli Lucherini , Mihir Kshirsagar , Arvind Narayanan , Jonathan Mayer

Web services are widely used in many areas via callable APIs, however, data are not always available in this way. We always need to get some data from web pages whose structure is not in order. Many developers use web data extraction…

Databases · Computer Science 2019-10-18 Naibo Wang , Zhiling Luo , Xiya Lyu , Zitong Yang , Jianwei Yin

Web search engines have marked everyone's life by transforming how one searches and accesses information. Search engines give special attention to the user interface, especially search engine result pages (SERP). The well-known ''10 blue…

Information Retrieval · Computer Science 2023-01-23 B. Oliveira , C. T. Lopes

Deep Research systems based on web agents have shown strong potential in solving complex information-seeking tasks, yet their search efficiency remains underexplored. We observe that many state-of-the-art open-source web agents rely on long…

Artificial Intelligence · Computer Science 2026-05-11 Junjie Wang , Zequn Xie , Dan Yang , Jie Feng , Yue Shen , Duolin Sun , Meixiu Long , Yihan Jiao , Zhehao Tan , Jian Wang , Peng Wei , Jinjie Gu

With the evolution of the online advertisement and tracking ecosystem, content-filtering has become the reference tool for improving the security, privacy and browsing experience when surfing the Internet. It is also commonly believed that…

Networking and Internet Architecture · Computer Science 2020-02-06 Ismael Castell-Uroz , Josep Solé-Pareta , Pere Barlet-Ros

The internet contains large amounts of low-quality content, yet users expect web search engines to deliver high-quality, relevant results. The abundant presence of low-quality pages can negatively impact retrieval and crawling processes by…

Information Retrieval · Computer Science 2025-04-16 Francesca Pezzuti , Ariane Mueller , Sean MacAvaney , Nicola Tonellotto

More and more people use the Internet to work on duties of their daily work routine. To find the right information online, Web search engines are the tools of their choice. Apart from finding facts, people use Web search engines to also…

Information Retrieval · Computer Science 2012-09-17 Georg Singer , Ulrich Norbisrath , Dirk Lewandowski

Accurately analyzing and modeling online browsing behavior play a key role in understanding users and technology interactions. In this work, we design and conduct a user study to collect browsing data from 31 participants continuously for…

Computers and Society · Computer Science 2021-08-17 Yuliia Lut , Michael Wang , Elissa M. Redmiles , Rachel Cummings

Search conducted in a work context is an everyday activity that has been around since long before the Web was invented, yet we still seem to understand little about its general characteristics. With this paper we aim to contribute to a…

Information Retrieval · Computer Science 2019-10-03 Suzan Verberne , Jiyin He , Gineke Wiggers , Tony Russell-Rose , Udo Kruschwitz , Arjen P. de Vries

Web browsers have come a long way since their inception, evolving from a simple means of displaying text documents over the network to complex software stacks with advanced graphics and network capabilities. As personal computers grew in…

Computers and Society · Computer Science 2023-05-01 Naif Mehanna , Walter Rudametkin

Web tracking is an omnipresent phenomenon in today's web, affecting users in their day-to-day lives. Filter lists and blockers were invented to detect trackers and to protect users. Due to limitations of said tools, researchers developed…

Cryptography and Security · Computer Science 2026-05-06 Wolf Rieder , Philip Raschke , Thomas Cory , Christian René Sechting , Aditya Kumar , Axel Küpper

The World Wide Web is the most wide known information source that is easily available and searchable. It consists of billions of interconnected documents Web pages are authored by millions of people. Accesses made by various users to pages…

Databases · Computer Science 2014-08-26 Priyanka Verma , Nishtha Kesswani

Web archives preserve portions of the web, but quantifying their completeness remains challenging. Prior approaches have estimated the coverage of a crawl by either comparing the outcomes of multiple crawlers, or by comparing the results of…

Physics and Society · Physics 2026-04-07 Michael Paris , Grigori Paris , Fabian Baumann

Performance in web applications is a key aspect of user experience and system scalability. Among the different techniques used to improve web application performance, caching has been widely used. While caching has been widely explored in…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-02-09 Mohammad Umar , Bharat Tripathi

Contextual retrieval is a critical technique for today's search engines in terms of facilitating queries and returning relevant information. This paper reports on the development and evaluation of a system designed to tackle some of the…

Information Retrieval · Computer Science 2014-07-24 Dilip K. Limbu , Andy M. Connor , Russel Pears , Stephen G. MacDonell

Biclustering is a two way clustering approach involving simultaneous clustering along two dimensions of the data matrix. Finding biclusters of web objects (i.e. web users and web pages) is an emerging topic in the context of web usage…

Neural and Evolutionary Computing · Computer Science 2011-06-14 R. Rathipriya , Dr. K. Thangavel , J. Bagyamani

It is widely known that people become better at an activity if they perform this activity long and often. Yet, the question is whether being active in related areas like communicating online, writing blog articles or commenting on community…

Information Retrieval · Computer Science 2015-11-19 Georg Singer , Pille Pruulmann-Vengerfeldt , Ulrich Norbisrath , Dirk Lewandowski

With the advance of technology, Criminal Justice agencies are being confronted with an increased need to investigate crimes perpetuated partially or entirely over the Internet. These types of crime are known as cybercrimes. In order to…

Cryptography and Security · Computer Science 2017-10-27 Christopher Warren , Eman El-Sheikh , Nhien-An Le-Khac

Web is a wide term which mainly consists of surface web and hidden web. One can easily access the surface web using traditional web crawlers, but they are not able to crawl the hidden portion of the web. These traditional crawlers retrieve…

Information Retrieval · Computer Science 2015-09-24 Manvi , Komal Kumar Bhatia , Ashutosh Dixit
‹ Prev 1 3 4 5 6 7 10 Next ›