English
Related papers

Related papers: CRATOR: a Dark Web Crawler

200 papers

Understanding and extracting of information from large documents, such as business opportunities, academic articles, medical documents and technical reports, poses challenges not present in short documents. Such large documents may be…

Computation and Language · Computer Science 2019-10-10 Muhammad Mahbubur Rahman , Tim Finin

Exploring the darknet can be a daunting task; this paper explores the application of data mining the darknet within a Canadian cybercrime perspective. Measuring activity through marketplace analysis and vendor attribution has proven…

Cryptography and Security · Computer Science 2021-05-31 Edward Crowder , Jay Lansiquot

Existing techniques for efficiently crawling social media sites rely on URL patterns, query logs, and human supervision. This paper describes SOUrCe, a structure-oriented unsupervised crawler that uses page structures to learn how to crawl…

Information Retrieval · Computer Science 2018-04-10 Keyang Xu , Kyle Yingkai Gao , Jamie Callan

As the amount of data on the World Wide Web continues to grow exponentially, access to semantically structured information remains limited. The Semantic Web has emerged as a solution to enhance the machine-readability of data, making it…

Digital Libraries · Computer Science 2023-06-21 Muhammad Zohaib

We carry out a comparative study on the problem for a walker searching on several typical complex networks. The search efficiency is evaluated for various strategies. Having no knowledge of the global properties of the underlying networks…

Disordered Systems and Neural Networks · Physics 2009-11-10 Shi-Jie Yang

The use of the un-indexed web, commonly known as the deep web and dark web, to commit or facilitate criminal activity has drastically increased over the past decade. The dark web is an in-famously dangerous place where all kinds of criminal…

Cryptography and Security · Computer Science 2023-12-05 Mohamed Chahine Ghanem , Patrick Mulvihill , Karim Ouazzane , Ramzi Djemai , Dipo Dunsin

Recent research has suggested that there are clear differences in the language used in the Dark Web compared to that of the Surface Web. As studies on the Dark Web commonly require textual analysis of the domain, language models specific to…

Computation and Language · Computer Science 2023-05-19 Youngjin Jin , Eugene Jang , Jian Cui , Jin-Woo Chung , Yongjae Lee , Seungwon Shin

Web tracking is an omnipresent phenomenon in today's web, affecting users in their day-to-day lives. Filter lists and blockers were invented to detect trackers and to protect users. Due to limitations of said tools, researchers developed…

Cryptography and Security · Computer Science 2026-05-06 Wolf Rieder , Philip Raschke , Thomas Cory , Christian René Sechting , Aditya Kumar , Axel Küpper

Browser fingerprinting is a pervasive online tracking technique used increasingly often for profiling and targeted advertising. Prior research on the prevalence of fingerprinting heavily relied on automated web crawls, which inherently…

Cryptography and Security · Computer Science 2025-02-04 Meenatchi Sundaram Muthu Selva Annamalai , Igor Bilogrevic , Emiliano De Cristofaro

Peer-to-Peer protocols currently form the most heavily used protocol class in the Internet, with BitTorrent, the most popular protocol for content distribution, as its flagship. A high number of studies and investigations have been…

Networking and Internet Architecture · Computer Science 2015-03-17 Răzvan Deaconescu , Marius Sandu-Popa , Adriana Drăghici , Nicolae Tăpus

Internet censors seek ways to identify and block internet access to information they deem objectionable. Increasingly, censors deploy advanced networking tools such as deep-packet inspection (DPI) to identify such connections. In response,…

Networking and Internet Architecture · Computer Science 2016-11-15 Lucas Dixon , Thomas Ristenpart , Thomas Shrimpton

We describe the development, characteristics and availability of a test collection for the task of Web table retrieval, which uses a large-scale Web Table Corpora extracted from the Common Crawl. Since a Web table usually has rich context…

Information Retrieval · Computer Science 2021-05-07 Zhiyu Chen , Shuo Zhang , Brian D. Davison

In the world of big data, many people find it difficult to access the information they need quickly and accurately. In order to overcome this, research on the system that recommends information accurately to users is continuously conducted.…

Computers and Society · Computer Science 2019-09-19 Keum Gang Cha , Soo-Ryeon Lee , Jung-Woo Lee , Seung Bin Baik

Programs for extracting structured information from text, namely information extractors, often operate separately on document segments obtained from a generic splitting operation such as sentences, paragraphs, k-grams, HTTP requests, and so…

Databases · Computer Science 2021-05-21 Johannes Doleschal , Benny Kimelfeld , Wim Martens , Frank Neven , Matthias Niewerth

Protecting users from accessing malicious web sites is one of the important management tasks for network operators. There are many open-source and commercial products to control web sites users can access. The most traditional approach is…

Networking and Internet Architecture · Computer Science 2021-11-12 Keiichi Shima , Daisuke Miyamoto , Hiroshi Abe , Tomohiro Ishihara , Kazuya Okada , Yuji Sekiya , Hirochika Asai , Yusuke Doi

Internet of Things (IoT) is a whole new ecosystem comprised of heterogeneous connected devices -i.e. computers, laptops, smart-phones and tablets as well as embedded devices and sensors-that communicate to deliver capabilities making our…

Cryptography and Security · Computer Science 2021-09-21 Stavros Shiaeles , Nicholas Kolokotronis , Emanuele Bellini

CAPTCHA is a human-centred test to distinguish a human operator from bots, attacking programs, or other computerised agents that tries to imitate human intelligence. In this research, we investigate a way to crack visual CAPTCHA tests by an…

Computer Vision and Pattern Recognition · Computer Science 2020-06-26 Zahra Noury , Mahdi Rezaei

The paper analyses current versions of top three used Internet browsers and compare their security levels to a research done in 2006. The security is measured by analyzing how user data is stored. Data recorded during different browsing…

Cryptography and Security · Computer Science 2011-12-30 Catalin Boja

In this paper, we uncover the essential features of websites that allow intelligent models to distinguish between phishing and legitimate sites. Phishing websites are those that are made with a similar user interface and a near similar…

Social and Information Networks · Computer Science 2022-05-09 Arash Negahdari Kia , Finbarr Murphy , Zahra Dehghani Mohammadabadi , Parisa Shamsi

Guaranteeing the security of information transmitted through the Internet, against passive or active attacks, is a major concern. The discovery of new pseudo-random number generators with a strong level of security is a field of research in…

Cryptography and Security · Computer Science 2015-03-19 Jacques M. Bahi , Xiaole Fang , Christophe Guyeux , Qianxue Wang