English
Related papers

Related papers: Web Usage mining framework for Data Cleaning and I…

200 papers

As the amount of data on the World Wide Web continues to grow exponentially, access to semantically structured information remains limited. The Semantic Web has emerged as a solution to enhance the machine-readability of data, making it…

Digital Libraries · Computer Science 2023-06-21 Muhammad Zohaib

World Wide Web (WWW) is the most popular global information sharing and communication system consisting of three standards .i.e., Uniform Resource Identifier (URL), Hypertext Transfer Protocol (HTTP) and Hypertext Mark-up Language (HTML).…

Artificial Intelligence · Computer Science 2010-08-11 Zeeshan Ahmed , Detlef Gerhard

Web 3.0 is the new generation of the Internet that is reconstructed with distributed technology, which focuses on data ownership and value expression. Also, it operates under the principle that data and digital assets should be owned and…

Artificial Intelligence · Computer Science 2023-09-19 Meng Shen , Zhehui Tan , Dusit Niyato , Yuzhi Liu , Jiawen Kang , Zehui Xiong , Liehuang Zhu , Wei Wang , Xuemin , Shen

Web search is frequently used by people to acquire new knowledge and to satisfy learning-related objectives. In this context, informational search missions with an intention to obtain knowledge pertaining to a topic are prominent. The…

Human-Computer Interaction · Computer Science 2018-05-03 Ran Yu , Ujwal Gadiraju , Peter Holtz , Markus Rokicki , Philipp Kemkes , Stefan Dietze

A large amount of data on the WWW remains inaccessible to crawlers of Web search engines because it can only be exposed on demand as users fill out and submit forms. The Hidden web refers to the collection of Web data which can be accessed…

Information Retrieval · Computer Science 2014-07-23 Sonali Gupta , Komal Kumar Bhatia

The World Wide Web, a ubiquitous source of information, serves as a primary resource for countless individuals, amassing a vast amount of data from global internet users. However, this online data, when scraped, indexed, and utilized for…

Networking and Internet Architecture · Computer Science 2023-11-07 Dawen Zhang , Boming Xia , Yue Liu , Xiwei Xu , Thong Hoang , Zhenchang Xing , Mark Staples , Qinghua Lu , Liming Zhu

It is widely known that people become better at an activity if they perform this activity long and often. Yet, the question is whether being active in related areas like communicating online, writing blog articles or commenting on community…

Information Retrieval · Computer Science 2015-11-19 Georg Singer , Pille Pruulmann-Vengerfeldt , Ulrich Norbisrath , Dirk Lewandowski

The technological revolution of the Internet has digitized the social, economic, political, and cultural activities of billions of humans. While researchers have been paying due attention to concerns of misinformation and bias, these…

Computers and Society · Computer Science 2025-10-14 Saurabh Khanna

The World Wide Web no longer consists just of HTML pages. Our work sheds light on a number of trends on the Internet that go beyond simple Web pages. The hidden Web provides a wealth of data in semi-structured form, accessible through Web…

Artificial Intelligence · Computer Science 2011-05-11 Fabian Suchanek , Aparna Varde , Richi Nayak , Pierre Senellart

Process mining methods allow analysts to use logs of historical executions of business processes in order to gain knowledge about the actual behavior of these processes. One of the most widely studied process mining operations is automated…

Software Engineering · Computer Science 2018-06-11 Fabrizio Maria Maggi , Andrea Marrella , Fredrik Milani , Allar Soo , Silva Kasela

The growth of world-wide-web (WWW) spreads its wings from an intangible quantities of web-pages to a gigantic hub of web information which gradually increases the complexity of crawling process in a search engine. A search engine handles a…

Machine Learning · Computer Science 2012-08-15 Sudarshan Nandy , Partha Pratim Sarkar , Achintya Das

The amount of information available on the Web grows at an incredible high rate. Systems and procedures devised to extract these data from Web sources already exist, and different approaches and techniques have been investigated during the…

Artificial Intelligence · Computer Science 2012-02-13 Emilio Ferrara , Robert Baumgartner

The battle for a more secure Internet is waged on many fronts, including the most basic of networking protocols. Our focus is the IPv4 Identifier (IPID), an IPv4 header field as old as the Internet with an equally long history as an…

Networking and Internet Architecture · Computer Science 2025-11-10 Joshua J. Daymude , Antonio M. Espinoza , Holly Bergen , Benjamin Mixon-Baca , Jeffrey Knockel , Jedidiah R. Crandall

We perform a large-scale analysis of third-party trackers on the World Wide Web from more than 3.5 billion web pages of the CommonCrawl 2012 corpus. We extract a dataset containing more than 140 million third-party embeddings in over 41…

Social and Information Networks · Computer Science 2016-08-01 Sebastian Schelter , Jérôme Kunegis

While scans of the IPv4 space are ubiquitous, today little is known about scanning activity in the IPv6 Internet. In this work, we present a longitudinal and detailed empirical study on large-scale IPv6 scanning behavior in the Internet,…

Networking and Internet Architecture · Computer Science 2022-10-20 Philipp Richter , Oliver Gasser , Arthur Berger

This project is an exploration into analysing WiFi probe requests, a management frame described as part of the IEEE 802.11 protocol which publicly broadcasts the senders MAC address. The intention was to collect these probe requests to use…

Networking and Internet Architecture · Computer Science 2018-05-22 Oisín Kyne

Mature social networking services are one of the greatest assets of today's organizations. This valuable asset, however, can also be a threat to an organization's confidentiality. Members of social networking websites expose not only their…

Social and Information Networks · Computer Science 2013-09-03 Michael Fire , Rami Puzis , Yuval Elovici

Nowadays online searches are undeniably the most common form of information gathering, as witnessed by billions of clicks generated each day on search engines. In this work we describe online searches as foraging processes that take place…

Physics and Society · Physics 2017-04-05 Xiangwen Wang , Michel Pleimling

In this paper, we uncover the essential features of websites that allow intelligent models to distinguish between phishing and legitimate sites. Phishing websites are those that are made with a similar user interface and a near similar…

Social and Information Networks · Computer Science 2022-05-09 Arash Negahdari Kia , Finbarr Murphy , Zahra Dehghani Mohammadabadi , Parisa Shamsi

The web is the prominent way information is exchanged in the 21st century. However, ensuring web-based information is accessible is complicated, particularly with web applications that rely on JavaScript and other technologies to deliver…

Information Retrieval · Computer Science 2019-08-09 Trevor Bostic , Jeff Stanley , John Higgins , Rachael L. Bradley-Montgomery , Justin F. Brunelle , Daniel Chudnov