English
Related papers

Related papers: A Comparative Study of Hidden Web Crawlers

200 papers

The rapid growth of web has resulted in vast volume of information. Information availability at a rapid speed to the user is vital. English language (or any for that matter) has lot of ambiguity in the usage of words. So there is no…

Information Retrieval · Computer Science 2011-08-30 Jeevan H E , Prashanth P P , Punith Kumar S N , Vinay Hegde

In this paper we argue that policies are an increasing concern for organizations that are operating a web site. Examples of policies that are relevant in the domain of the web address issues such as privacy of personal data, accessibility…

Computers and Society · Computer Science 2008-07-31 Holger M. Kienle , Hausi A. Müller

This paper presents a new general framework of information hiding, in which the hidden information is embedded into a collection of activities conducted by selected human and computer entities (e.g., a number of online accounts of one or…

Cryptography and Security · Computer Science 2018-10-04 Shujun Li , Anthony T. S. Ho , Zichi Wang , Xinpeng Zhang

Data collection is a key component of an information system. The widespread penetration of ICT tools in organizations and institutions has resulted in a shift in the way the data is collected. Data may be collected in printed-form, by…

Computers and Society · Computer Science 2013-03-27 Ruchika Thukral , Anita Goel

Searching accounts for one of the most frequently performed computations over the Internet as well as one of the most important applications of outsourced computing, producing results that critically affect users' decision-making behaviors.…

Beyond the information stored in pages of the World Wide Web, novel types of ``meta-information'' are created when they connect to each other. This information is a collective effect of independent users writing and linking pages, hidden…

Disordered Systems and Neural Networks · Physics 2009-11-07 Jean-Pierre Eckmann , Elisha Moses

Structure information extraction refers to the task of extracting structured text fields from web pages, such as extracting a product offer from a shopping page including product title, description, brand and price. It is an important…

Computation and Language · Computer Science 2022-02-02 Qifan Wang , Yi Fang , Anirudh Ravula , Fuli Feng , Xiaojun Quan , Dongfang Liu

The Semantic Web initiative puts emphasis not primarily on putting data on the Web, but rather on creating links in a way that both humans and machines can explore the Web of data. When such users access the Web, they leave a trail as Web…

Information Retrieval · Computer Science 2011-04-07 Markus Kirchberg , Ryan K L Ko , Bu Sung Lee

Nowadays, all major web browsers have a private browsing mode. However, the mode's benefits and limitations are not particularly understood. Through the use of survey studies, prior work has found that most users are either unaware of…

Human-Computer Interaction · Computer Science 2019-06-04 Ruba Abu-Salma , Benjamin Livshits

Search engines have vast technical capabilities to retain Internet search logs for each user and thus present major privacy vulnerabilities to both individuals and organizations in revealing user intent. Additionally, many of the web search…

Cryptography and Security · Computer Science 2025-06-11 Kato Mivule , Kenneth Hopkinson

Search Engine has become a major tool for searching any information from the World Wide Web (WWW). While searching the huge digital library available in the WWW, every effort is made to retrieve the most relevant results. But in WWW…

Information Retrieval · Computer Science 2011-09-26 Debajyoti Mukhopadhyay , Sukanta Sinha

Event collections are frequently built by crawling the live web on the basis of seed URIs nominated by human experts. Focused web crawling is a technique where the crawler is guided by reference content pertaining to the event. Given the…

Digital Libraries · Computer Science 2018-04-06 Martin Klein , Lyudmila Balakireva , Herbert Van de Sompel

Web crawl is a main source of large language models' (LLMs) pretraining data, but the majority of crawled web pages are discarded in pretraining due to low data quality. This paper presents Craw4LLM, an efficient web crawling method that…

Computation and Language · Computer Science 2025-06-24 Shi Yu , Zhiyuan Liu , Chenyan Xiong

Variety, size and complexity of data types, services and applications in Internet is continuously growing up. This increasing of complexity needs more powerful and sophisticated equipment's. One group of these devices that has essential…

Networking and Internet Architecture · Computer Science 2022-03-04 Mazdak Fatahi , Masou Soursouri , Pooya Pourmohammad , Mahmood Ahmadi

The internet is a common place for businesses to collect and store as much client data as possible and computer storage capacity has increased exponentially due to this trend. Businesses utilize this data to enhance customer satisfaction,…

Cryptography and Security · Computer Science 2024-10-08 Rabia Bajwa , Farah Tasnur Meem

Cloud computing has become a potential resource for businesses and individuals to outsource their data to remote but highly accessible servers. However, potentials of the cloud services have not been fully unleashed due to users' concerns…

Cryptography and Security · Computer Science 2018-11-27 Hoang Pham , Jason Woodworth , Mohsen Amini Salehi

After the Internet and the World Wide Web have become popular and widely-available, the electronically stored online interactions of individuals have fast emerged as a challenge for researchers and, perhaps even faster, as a source of…

Social and Information Networks · Computer Science 2012-08-23 Matus Medo

The World Wide Web is a popular and interactive medium to distribute information in this scenario. The web is huge, diverse, ever changing, widely disseminated global information service center. We are familiar with terms like e-commerce,…

Other Computer Science · Computer Science 2012-08-30 Priyanka Rahi

Web mining is the nontrivial process to discover valid, novel, potentially useful knowledge from web data using the data mining techniques or methods. It may give information that is useful for improving the services offered by web portals…

Information Retrieval · Computer Science 2011-10-03 R. Rathipriya , K. Thangavel , J. Bagyamani

The availability of sophisticated technologies and methods of perpetrating criminogenic activities in the cyberspace is a pertinent societal problem. Darknet is an encrypted network technology that uses the internet infrastructure and can…

Cryptography and Security · Computer Science 2020-03-18 Victor Adewopo , Bilal Gonen , Said Varlioglu , Murat Ozer