中文
相关论文

相关论文: Automatic Generation of Web Censorship Probe Lists

200 篇论文

Testing web forms is an essential activity for ensuring the quality of web applications. It typically involves evaluating the interactions between users and forms. Automated test-case generation remains a challenge for web-form testing: Due…

软件工程 · 计算机科学 2025-05-20 Tao Li , Chenhui Cui , Rubing Huang , Dave Towey , Lei Ma

This paper presents a new method for automatically detecting words with lexical gender in large-scale language datasets. Currently, the evaluation of gender bias in natural language processing relies on manually compiled lexicons of…

计算与语言 · 计算机科学 2022-06-29 Marion Bartl , Susan Leavy

As the World Wide Web is growing rapidly, it is getting increasingly challenging to gather representative information about it. Instead of crawling the web exhaustively one has to resort to other techniques like sampling to determine the…

数据结构与算法 · 计算机科学 2009-02-11 Eda Baykan , Monika Henzinger , Stefan F. Keller , Sebastian De Castelberg , Markus Kinzler

Online toxic content has grown into a pervasive phenomenon, intensifying during times of crisis, elections, and social unrest. A significant amount of research has been focused on detecting or analyzing toxic content using machine-learning…

计算与语言 · 计算机科学 2025-09-19 Gautam Kishore Shahi , Tim A. Majchrzak

Internet censorship is a phenomenon of societal importance and attracts investigation from multiple disciplines. Several research groups, such as Censored Planet, have deployed large scale Internet measurement platforms to collect network…

机器学习 · 计算机科学 2023-02-28 Shawn P. Duncan , Hui Chen

Modern botnets rely on domain-generation algorithms (DGAs) to build resilient command-and-control infrastructures. Recent works focus on recognizing automatically generated domains (AGDs) from DNS traffic, which potentially allows to…

密码学与安全 · 计算机科学 2013-11-25 Stefano Schiavoni , Federico Maggi , Lorenzo Cavallaro , Stefano Zanero

Topic modelling is a popular unsupervised method for identifying the underlying themes in document collections that has many applications in information retrieval. A topic is usually represented by a list of terms ranked by their…

信息检索 · 计算机科学 2020-06-02 Areej Alokaili , Nikolaos Aletras , Mark Stevenson

General purpose Search Engines (SEs) crawl all domains (e.g., Sports, News, Entertainment) of the Web, but sometimes the informational need of a query is restricted to a particular domain (e.g., Medical). We leverage the work of SEs as part…

信息检索 · 计算机科学 2016-05-04 Alexander Nwala , Michael Nelson

This paper addresses the challenge of automatically extracting attributes from news article web pages across multiple languages. Recent neural network models have shown high efficacy in extracting information from semi-structured web pages.…

计算与语言 · 计算机科学 2025-02-05 Pavel Bedrin , Maksim Varlamov , Alexander Yatskov

Cross-domain sentiment analysis (CDSA) helps to address the problem of data scarcity in scenarios where labelled data for a domain (known as the target domain) is unavailable or insufficient. However, the decision to choose a domain (known…

计算与语言 · 计算机科学 2020-04-10 Akash Sheoran , Diptesh Kanojia , Aditya Joshi , Pushpak Bhattacharyya

Malicious crowdsourcing forums are gaining traction as sources of spreading misinformation online, but are limited by the costs of hiring and managing human workers. In this paper, we identify a new class of attacks that leverage deep…

密码学与安全 · 计算机科学 2017-09-11 Yuanshun Yao , Bimal Viswanath , Jenna Cryan , Haitao Zheng , Ben Y. Zhao

Semantic code search is the task of retrieving relevant code given a natural language query. While related to other information retrieval tasks, it requires bridging the gap between the language used in code (often abbreviated and highly…

机器学习 · 计算机科学 2020-06-09 Hamel Husain , Ho-Hsiang Wu , Tiferet Gazit , Miltiadis Allamanis , Marc Brockschmidt

Large language models' (LLMs) abilities are drawn from their pretraining data, and model development begins with data curation. However, decisions around what data is retained or removed during this initial stage are under-scrutinized. In…

计算与语言 · 计算机科学 2024-06-24 Li Lucy , Suchin Gururangan , Luca Soldaini , Emma Strubell , David Bamman , Lauren F. Klein , Jesse Dodge

Malicious advertisement URLs pose a security risk since they are the source of cyber-attacks, and the need to address this issue is growing in both industry and academia. Generally, the attacker delivers an attack vector to the user by…

机器学习 · 计算机科学 2022-04-29 Ehsan Nowroozi , Abhishek , Mohammadreza Mohammadi , Mauro Conti

Malicious web domains represent a big threat to web users' privacy and security. With so much freely available data on the Internet about web domains' popularity and performance, this study investigated the performance of well-known machine…

密码学与安全 · 计算机科学 2019-02-26 Zhongyi Hu , Raymond Chiong , Ilung Pranata , Willy Susilo , Yukun Bao

Many text databases on the web are "hidden" behind search interfaces, and their documents are only accessible through querying. Search engines typically ignore the contents of such search-only databases. Recently, Yahoo-like directories…

数据库 · 计算机科学 2007-05-23 Panagiotis Ipeirotis , Luis Gravano , Mehran Sahami

Using author provided tags to predict tags for a new document often results in the overgeneration of tags. In the case where the author doesn't provide any tags, our documents face the severe under-tagging issue. In this paper, we present a…

信息检索 · 计算机科学 2020-05-04 Maharshi R. Pandya , Jessica Reyes , Bob Vanderheyden

A major challenge in paraphrase research is the lack of parallel corpora. In this paper, we present a new method to collect large-scale sentential paraphrases from Twitter by linking tweets through shared URLs. The main advantage of our…

计算与语言 · 计算机科学 2017-08-02 Wuwei Lan , Siyu Qiu , Hua He , Wei Xu

Web page categorization is one of the challenging tasks in the world of ever increasing web technologies. There are many ways of categorization of web pages based on different approach and features. This paper proposes a new dimension in…

神经与进化计算 · 计算机科学 2010-09-28 S. M. Kamruzzaman

Web tracking has been extensively studied over the last decade. To detect tracking, previous studies and user tools rely on filter lists. However, it has been shown that filter lists miss trackers. In this paper, we propose an alternative…

密码学与安全 · 计算机科学 2020-03-02 Imane Fouad , Nataliia Bielova , Arnaud Legout , Natasa Sarafijanovic-Djukic