中文
相关论文

相关论文: Using Exclusive Web Crawlers to Store Better Resul…

200 篇论文

A web crawler is a system designed to collect web pages, and efficient crawling of new pages requires appropriate algorithms. While website features such as XML sitemaps and the frequency of past page updates provide important clues for…

信息检索 · 计算机科学 2025-05-13 Yuichi Sasazawa , Yasuhiro Sogawa

The exponential growth of information source on the web and in turn continuing technological progress of searching the information by using tools like Search Engines gives rise to many problems for the user to know which tool is best for…

信息检索 · 计算机科学 2025-03-13 Rajender Nath , Satinder Bal

Nowadays, more and more people use the Web as their primary source of up-to-date information. In this context, fast crawling and indexing of newly created Web pages has become crucial for search engines, especially because user traffic to a…

信息检索 · 计算机科学 2013-07-25 Damien Lefortier , Liudmila Ostroumova , Egor Samosvat , Pavel Serdyukov

Given the vast scale of the Web, crawling prioritisation techniques based on link graph traversal, popularity, link analysis, and textual content are frequently applied to surface documents that are most likely to be valuable. While…

信息检索 · 计算机科学 2025-07-03 Francesca Pezzuti , Sean MacAvaney , Nicola Tonellotto

Dark web crawling is a complex process that involves specific methodologies and techniques to navigate the Tor network and extract data from hidden services. This study proposes a general dark web crawler designed to extract pages handling…

密码学与安全 · 计算机科学 2024-05-13 Daniel De Pascale , Giuseppe Cascavilla , Damian A. Tamburri , Willem-Jan Van Den Heuvel

Journalistic fact-checking, as well as social or economic research, require analyzing high-quality statistics datasets (SDs, in short). However, retrieving SD corpora at scale may be hard, inefficient, or impossible, depending on how they…

信息检索 · 计算机科学 2026-02-13 Antoine Gauquier , Ioana Manolescu , Pierre Senellart

The Hidden Web is the vast repository of informational databases available only through search form interfaces, accessible by therein typing a set of keywords in the search forms. Typically, a Hidden Web crawler is employed to autonomously…

信息检索 · 计算机科学 2013-11-05 Sonali Gupta , Komal Kumar Bhatia

The internet contains large amounts of low-quality content, yet users expect web search engines to deliver high-quality, relevant results. The abundant presence of low-quality pages can negatively impact retrieval and crawling processes by…

信息检索 · 计算机科学 2025-04-16 Francesca Pezzuti , Ariane Mueller , Sean MacAvaney , Nicola Tonellotto

Internet search engines function in a present which changes continuously. The search engines update their indices regularly, overwriting Web pages with newer ones, adding new pages to the index, and losing older ones. Some search engines…

信息检索 · 计算机科学 2009-11-19 Iina Hellsten , Loet Leydesdorff , Paul Wouters

Nowadays, many web databases "hidden" behind their restrictive search interfaces (e.g., Amazon, eBay) contain rich and valuable information that is of significant interests to various third parties. Recent studies have demonstrated the…

数据库 · 计算机科学 2016-11-22 Saad Bin Suhaim , Weimo Liu , Nan Zhang

Web crawl is a main source of large language models' (LLMs) pretraining data, but the majority of crawled web pages are discarded in pretraining due to low data quality. This paper presents Craw4LLM, an efficient web crawling method that…

计算与语言 · 计算机科学 2025-06-24 Shi Yu , Zhiyuan Liu , Chenyan Xiong

With the rapid advance of the Internet, search engines (e.g., Google, Bing, Yahoo!) are used by billions of users for each day. The main function of a search engine is to locate the most relevant webpages corresponding to what the user…

应用统计 · 统计学 2018-03-15 Xinzhi Han , Sen Lei

Searching accounts for one of the most frequently performed computations over the Internet as well as one of the most important applications of outsourced computing, producing results that critically affect users' decision-making behaviors.…

A typical web search engine consists of three principal parts: crawling engine, indexing engine, and searching engine. The present work aims to optimize the performance of the crawling engine. The crawling engine finds new web pages and…

网络与互联网体系结构 · 计算机科学 2012-01-20 Konstantin Avrachenkov , Alexander Dudin , Valentina Klimenok , Philippe Nain , Olga Semenova

The rapid growth of web has resulted in vast volume of information. Information availability at a rapid speed to the user is vital. English language (or any for that matter) has lot of ambiguity in the usage of words. So there is no…

信息检索 · 计算机科学 2011-08-30 Jeevan H E , Prashanth P P , Punith Kumar S N , Vinay Hegde

Web refresh crawling is the problem of keeping a cache of web pages fresh, that is, having the most recent copy available when a page is requested, given a limited bandwidth available to the crawler. Under the assumption that the change and…

Search engines are the preferred tools for finding information on the Web. They are advancing to be the common helpers to answer any of our search needs. We use them to carry out simple look-up tasks and also to work on rather time…

信息检索 · 计算机科学 2012-06-13 Georg Singer , Ulrich Norbisrath , Dirk Lewandowski

The ranked retrieval model has rapidly become the de-facto way for search query processing in web databases. Despite the extensive efforts on designing better ranking mechanisms, in practice, many such databases fail to address the diverse…

数据库 · 计算机科学 2018-07-17 Yeshwanth D. Gunasekaran , Abolfazl Asudeh , Sona Hasani , Nan Zhang , Ali Jaoua , Gautam Das

Getting informed of what is registered in the Web space on time, can greatly help the psychologists, marketers and political analysts to familiarize, analyse, make decision and act correctly based on the society`s different needs. The great…

信息检索 · 计算机科学 2012-02-10 Mehdi Naghavi , Mohsen Sharifi

Major search engines deploy personalized Web results to enhance users' experience, by showing them data supposed to be relevant to their interests. Even if this process may bring benefits to users while browsing, it also raises concerns on…

信息检索 · 计算机科学 2015-08-18 Van Tien Hoang , Angelo Spognardi , Francesco Tiezzi , Marinella Petrocchi , Rocco De Nicola