English
Related papers

Related papers: A Comparative Study of Hidden Web Crawlers

200 papers

Biclustering is a two way clustering approach involving simultaneous clustering along two dimensions of the data matrix. Finding biclusters of web objects (i.e. web users and web pages) is an emerging topic in the context of web usage…

Neural and Evolutionary Computing · Computer Science 2011-06-14 R. Rathipriya , Dr. K. Thangavel , J. Bagyamani

Health information websites offer instantaneous access to information, but have important privacy implications as they can associate a visitor with specific medical conditions. We interviewed 35 residents of Canada to better understand…

Human-Computer Interaction · Computer Science 2026-03-30 Martin P. Robillard , Lihn V. Nguyen , Deeksha Arya , Jin L. C. Guo

The field of Web services is an important paradigm in distributed application development. Currently, many businesses are seeking to convert their applications into web services because of its ability to promote inter-operability among…

Software Engineering · Computer Science 2013-04-08 A Anji Reddy , S Sowmya Kamath

The random surfer model is a frequently used model for simulating user navigation behavior on the Web. Various algorithms, such as PageRank, are based on the assumption that the model represents a good approximation of users browsing a…

Social and Information Networks · Computer Science 2015-08-05 Florian Geigl , Daniel Lamprecht , Rainer Hofmann-Wellenhof , Simon Walk , Markus Strohmaier , Denis Helic

Web is a primary and essential service to share information among users and organizations at present all over the world. Despite the current significance of such a kind of traffic on the Internet, the so-called Surface Web traffic has been…

Cryptography and Security · Computer Science 2021-12-03 Roberto Magán-Carrión , Alberto Abellán-Galera , Gabriel Maciá-Fernández , Pedro García-Teodoro

Most present day organisations make use of some automated information system. This usually means that a large body of vital corporate information is stored in these information systems. As a result, an essential function of information…

Information Retrieval · Computer Science 2021-05-20 H. A. Proper

Machine-learning (ML) shortcuts or spurious correlations are artifacts in datasets that lead to very good training and test performance but severely limit the model's generalization capability. Such shortcuts are insidious because they go…

Artificial Intelligence · Computer Science 2023-10-31 Nicolas M. Müller , Maximilian Burgert , Pascal Debus , Jennifer Williams , Philip Sperl , Konstantin Böttinger

Web services are playing an important role in e-business and e-commerce applications. As web service applications are interoperable and can work on any platform, large scale distributed systems can be developed easily using web services.…

Information Retrieval · Computer Science 2012-06-26 Debajyoti Mukhopadhyay , Archana Chougule

In this paper, we analyze the nature and distribution of structured data on the Web. Web-scale information extraction, or the problem of creating structured tables using extraction from the entire web, is gathering lots of research…

Databases · Computer Science 2012-03-30 Nilesh Dalvi , Ashwin Machanavajjhala , Bo Pang

The abundance of the data in the Internet facilitates the improvement of extraction and processing tools. The trend in the open data publishing encourages the adoption of structured formats like CSV and RDF. However, there is still a…

Information Retrieval · Computer Science 2016-08-08 Mikhail Galkin , Dmitry Mouromtsev , Sören Auer

Peer to peer (P2P) networks are an overlay on IP network of the internet and they can shape the future of computing by their involvement in distributed systems with the increased of use of low priced personal computers to form big clusters…

Networking and Internet Architecture · Computer Science 2020-01-07 Mohamed Elsharnouby

Web accessibility is actually the most important aspect for providing access to information and interaction for people with disabilities. However, it seems that the ability of users with disabilities to navigate over the Web is not…

Human-Computer Interaction · Computer Science 2013-07-22 Abdelhakim Herrouz , Chabane Khentout , Mahieddine Djoudi

Content distribution networks (CDNs) which serve to deliver web objects (e.g., documents, applications, music and video, etc.) have seen tremendous growth since its emergence. To minimize the retrieving delay experienced by a user with a…

Networking and Internet Architecture · Computer Science 2012-10-02 Jing Zhang

Web archives preserve portions of the web, but quantifying their completeness remains challenging. Prior approaches have estimated the coverage of a crawl by either comparing the outcomes of multiple crawlers, or by comparing the results of…

Physics and Society · Physics 2026-04-07 Michael Paris , Grigori Paris , Fabian Baumann

Can a Web crawler efficiently locate an unknown relevant page? While this question is receiving much empirical attention due to its considerable commercial value in the search engine community [Cho98,Chakrabarti99,Menczer00,Menczer01],…

Information Retrieval · Computer Science 2007-05-23 Filippo Menczer

There has been recently a growth of interest in developing the current machine-readable Web towards the next generation of machine-understandable Web - Semantic Web. The development of the Web to a global business was reasonably fast,…

Software Engineering · Computer Science 2021-05-07 Bryar A. Hassan

Nowadays, the size of the Internet is experiencing rapid growth. As of December 2014, the number of global Internet websites has more than 1 billion and all kinds of information resources are integrated together on the Internet, however,the…

Distributed, Parallel, and Cluster Computing · Computer Science 2015-06-02 Qingpei Guo , Chao Xu , Yang Song

The immense amounts of source code provide ample challenges and opportunities during software development. To handle the size of code bases, developers commonly search for code, e.g., when trying to find where a particular feature is…

Software Engineering · Computer Science 2022-10-06 Luca Di Grazia , Michael Pradel

The historical, cultural, and intellectual importance of archiving the web has been widely recognized. Today, all countries with high Internet penetration rate have established high-profile archiving initiatives to crawl and archive the…

Digital Libraries · Computer Science 2013-08-13 Zhiwu Xie , Herbert Van de Sompel , Jinyang Liu , Johann van Reenen , Ramiro Jordan

Over the last 30 years, the World Wide Web has changed significantly. In this paper, we argue that common practices to prepare web pages for delivery conflict with many efforts to present content with minimal latency, one fundamental goal…

Networking and Internet Architecture · Computer Science 2024-03-26 Lucas Vogel , Thomas Springer , Matthias Wählisch