中文
相关论文

相关论文: Using Exclusive Web Crawlers to Store Better Resul…

200 篇论文

Blogs are undoubtedly the richest source of information available in cyberspace. Blogs can be of various natures i.e. personal blogs which contain posts on mixed issues or blogs can be domain specific which contains posts on particular…

信息检索 · 计算机科学 2013-08-01 Deepti Kapri , Rosy Madaan , A. K Sharma , Ashutosh Dixit

General purpose Search Engines (SEs) crawl all domains (e.g., Sports, News, Entertainment) of the Web, but sometimes the informational need of a query is restricted to a particular domain (e.g., Medical). We leverage the work of SEs as part…

信息检索 · 计算机科学 2016-05-04 Alexander Nwala , Michael Nelson

Query-based searching and browsing-based navigation are the two main components of exploratory search. Search lets users dig in deep by controlling their actions to focus on and find just the information they need, whereas navigation helps…

信息检索 · 计算机科学 2017-11-02 Mostafa Dehghani , Glorianna Jagfeld , Hosein Azarbonyad , Alex Olieman , Jaap Kamps , Maarten Marx

In contemporary times, people rely heavily on the internet and search engines to obtain information, either directly or indirectly. However, the information accessible to users constitutes merely 4% of the overall information present on the…

信息检索 · 计算机科学 2023-07-04 Ali Fayzi , Mohammad Fayzi , Kourosh Dadashtabar Ahmadi

We perform a large-scale analysis of third-party trackers on the World Wide Web from more than 3.5 billion web pages of the CommonCrawl 2012 corpus. We extract a dataset containing more than 140 million third-party embeddings in over 41…

社会与信息网络 · 计算机科学 2016-08-01 Sebastian Schelter , Jérôme Kunegis

The number of RDF knowledge graphs available on the Web grows constantly. Gathering these graphs at large scale for downstream applications hence requires the use of crawlers. Although Data Web crawlers exist, and general Web crawlers could…

CoWrangler is a data-wrangling recommender system designed to streamline data processing tasks. Recognizing that data processing is often time-consuming and complex for novice users, we aim to simplify the decision-making process regarding…

数据库 · 计算机科学 2024-09-18 Yuqing Wang , Anna Fariha

Query in a search engine is generally based on natural language. A query can be expressed in more than one way without changing its meaning as it depends on thinking of human being at a particular moment. Aim of the searcher is to get most…

信息检索 · 计算机科学 2013-03-08 Raj Kishor Bisht , Ila Pant Bisht

The emergence of Linked Data on the WWW has spawned research interest in an online execution of declarative queries over this data. A particularly interesting approach is traversal-based query execution which fetches data by traversing data…

数据库 · 计算机科学 2016-07-06 Olaf Hartig , M. Tamer Özsu

Tables on the Web contain a vast amount of knowledge in a structured form. To tap into this valuable resource, we address the problem of table retrieval: answering an information need with a ranked list of tables. We investigate this…

信息检索 · 计算机科学 2021-05-14 Shuo Zhang , Krisztian Balog

Missing web pages, URIs that return the 404 "Page Not Found" error or the HTTP response code 200 but dereference unexpected content, are ubiquitous in today's browsing experience. We use Internet search engines to relocate such missing…

信息检索 · 计算机科学 2010-04-19 Martin Klein , Jeffery Shipman , Michael L. Nelson

Browser-based language models often use retrieval-augmented generation (RAG) but typically rely on fixed, outdated indices that give users no control over which sources are consulted. This can lead to answers that mix trusted and untrusted…

人机交互 · 计算机科学 2026-01-27 Saber Zerhoudi , Michael Dinzinger , Michael Granitzer , Jelena Mitrovic

We describe Voyageur, which is an application of experiential search to the domain of travel. Unlike traditional search engines for online services, experiential search focuses on the experiential aspects of the service under consideration.…

Data is the central asset of today's dynamically operating organization and their business. This data is usually stored in database. A major consideration is applied on the security of that data from the unauthorized access and intruders.…

数据库 · 计算机科学 2013-08-25 Manish Sharma , Atul Chaudhary , Santosh Kumar

This work presents a systematic study of navigational tracking, the latest development in the cat-and-mouse game between browsers and online trackers. Navigational tracking allows trackers to 'aggregate users' activities and behaviors…

密码学与安全 · 计算机科学 2022-07-14 Audrey Randall , Peter Snyder , Alisha Ukani , Alex Snoeren , Geoff Voelker , Stefan Savage , Aaron Schulman

Indexes are the best apposite choice for quickly retrieving the records. This is nothing but cutting down the number of Disk IO. Instead of scanning the complete table for the results, we can decrease the number of IO's or page fetches…

数据库 · 计算机科学 2019-03-21 Sourav Mukherjee

Search engines are the most commonly used type of tool for finding relevant information on the Internet. However, today's search engines are far from perfect. Typical search queries are short, often one or two words, and can be ambiguous…

信息检索 · 计算机科学 2014-07-24 Dilip K. Limbu , Andy M. Connor , Stephen G. MacDonell

Many web databases are "hidden" behind proprietary search interfaces that enforce the top-$k$ output constraint, i.e., each query returns at most $k$ of all matching tuples, preferentially selected and returned according to a proprietary…

数据库 · 计算机科学 2017-05-10 Abolfazl Asudeh , Saravanan Thirumuruganathan , Nan Zhang , Gautam Das

Deep Research systems based on web agents have shown strong potential in solving complex information-seeking tasks, yet their search efficiency remains underexplored. We observe that many state-of-the-art open-source web agents rely on long…

人工智能 · 计算机科学 2026-05-11 Junjie Wang , Zequn Xie , Dan Yang , Jie Feng , Yue Shen , Duolin Sun , Meixiu Long , Yihan Jiao , Zhehao Tan , Jian Wang , Peng Wei , Jinjie Gu

The vastness of the web imposes a prohibitive cost on building large-scale search engines with limited resources. Crawl frontiers thus need to be optimized to improve the coverage and freshness of crawled content. In this paper, we propose…