中文
相关论文

相关论文: Automatic Generation of Web Censorship Probe Lists

200 篇论文

Many governments impose traditional censorship methods on social media platforms. Instead of removing it completely, many social media companies, including Twitter, only withhold the content from the requesting country. This makes such…

社会与信息网络 · 计算机科学 2021-08-25 Tuğrulcan Elmas , Rebekah Overdorf , Karl Aberer

The number of web pages is growing at an exponential rate, accumulating massive amounts of data on the web. It is one of the key processes to classify webpages in web information mining. Some classical methods are based on manually building…

计算与语言 · 计算机科学 2023-05-10 Qiwei Lang , Jingbo Zhou , Haoyi Wang , Shiqi Lyu , Rui Zhang

Pull Requests (PRs) are a mechanism on modern collaborative coding platforms, such as GitHub. PRs allow developers to tell others that their code changes are available for merging into another branch in a repository. A PR needs to be…

软件工程 · 计算机科学 2022-07-01 Ting Zhang , Ivana Clairine Irsan , Ferdian Thung , DongGyun Han , David Lo , Lingxiao Jiang

Training data are critical in face recognition systems. However, labeling a large scale face data for a particular domain is very tedious. In this paper, we propose a method to automatically and incrementally construct datasets from massive…

计算机视觉与模式识别 · 计算机科学 2016-11-28 Shengyong Ding , Junyu Wu , Wei Xu , Hongyang Chao

When trained on large, unfiltered crawls from the internet, language models pick up and reproduce all kinds of undesirable biases that can be found in the data: they often generate racist, sexist, violent or otherwise toxic language. As…

计算与语言 · 计算机科学 2021-09-10 Timo Schick , Sahana Udupa , Hinrich Schütze

The escalating landscape of cyber threats, characterized by the registration of thousands of new domains daily for large-scale Internet attacks such as spam, phishing, and drive-by downloads, underscores the imperative for innovative…

密码学与安全 · 计算机科学 2024-07-11 Furkan Çolhak , Mert İlhan Ecevit , Hasan Dağ , Reiner Creutzburg

A web crawler is a system designed to collect web pages, and efficient crawling of new pages requires appropriate algorithms. While website features such as XML sitemaps and the frequency of past page updates provide important clues for…

信息检索 · 计算机科学 2025-05-13 Yuichi Sasazawa , Yasuhiro Sogawa

Automatic sarcasm detection methods have traditionally been designed for maximum performance on a specific domain. This poses challenges for those wishing to transfer those approaches to other existing or novel domains, which may be…

计算与语言 · 计算机科学 2018-06-12 Natalie Parde , Rodney D. Nielsen

The proliferation of unreliable news domains on the internet has had wide-reaching negative impacts on society. We introduce and evaluate interventions aimed at reducing traffic to unreliable news domains from search engines while…

信息检索 · 计算机科学 2024-04-16 Peter Carragher , Evan M. Williams , Kathleen M. Carley

Effective human learning depends on a wide selection of educational materials that align with the learner's current understanding of the topic. While the Internet has revolutionized human learning or education, a substantial resource…

Enabled by the pull-based development model, developers can easily contribute to a project through pull requests (PRs). When creating a PR, developers can add a free-form description to describe what changes are made in this PR and/or why.…

软件工程 · 计算机科学 2019-09-17 Zhongxin Liu , Xin Xia , Christoph Treude , David Lo , Shanping Li

With an ever-increasing number of scientific papers published each year, it becomes more difficult for researchers to explore a field that they are not closely familiar with already. This greatly inhibits the potential for…

机器学习 · 计算机科学 2021-06-08 Anna Nikiforovskaya , Nikolai Kapralov , Anna Vlasova , Oleg Shpynov , Aleksei Shpilman

With the rise of sophisticated scam websites that exploit human psychological vulnerabilities, distinguishing between legitimate and scam websites has become increasingly challenging. This paper presents ScamFerret, an innovative agent…

密码学与安全 · 计算机科学 2025-02-17 Hiroki Nakano , Takashi Koide , Daiki Chiba

Over the past decade, Internet centralization and its implications for both people and the resilience of the Internet has become a topic of active debate. While the networking community informally agrees on the definition of centralization,…

网络与互联网体系结构 · 计算机科学 2025-04-29 Gautam Akiwate , Kimberly Ruth , Rumaisa Habib , Zakir Durumeric

Advertising (ad for short) keyword suggestion is important for sponsored search to improve online advertising and increase search revenue. There are two common challenges in this task. First, the keyword bidding problem: hot ad keywords are…

计算与语言 · 计算机科学 2019-02-28 Hao Zhou , Minlie Huang , Yishun Mao , Changlei Zhu , Peng Shu , Xiaoyan Zhu

Since previous studies on open-domain targeted sentiment analysis are limited in dataset domain variety and sentence level, we propose a novel dataset consisting of 6,013 human-labeled data to extend the data domains in topics of interest…

计算与语言 · 计算机科学 2022-04-18 Yun Luo , Hongjie Cai , Linyi Yang , Yanxia Qin , Rui Xia , Yue Zhang

As news and social media exhibit an increasing amount of manipulative polarized content, detecting such propaganda has received attention as a new task for content analysis. Prior work has focused on supervised learning with training data…

计算与语言 · 计算机科学 2020-11-24 Liqiang Wang , Xiaoyu Shen , Gerard de Melo , Gerhard Weikum

Web applications are critical to modern software ecosystems, yet ensuring their reliability remains challenging due to the complexity and dynamic nature of web interfaces. Recent advances in large language models (LLMs) have shown promise…

A consumer-dependent (business-to-consumer) organization tends to present itself as possessing a set of human qualities, which is termed as the brand personality of the company. The perception is impressed upon the consumer through the…

计算与语言 · 计算机科学 2021-08-17 Soumyadeep Roy , Shamik Sural , Niyati Chhaya , Anandhavelu Natarajan , Niloy Ganguly

Deep-research agents, i.e., systems that rely on multi-agent pipelines to iteratively retrieve, synthesize, and cite Web content in order to produce structured reports, are rapidly replacing traditional search for both routine and complex…

密码学与安全 · 计算机科学 2026-05-26 Tingwei Zhang , Harold Triedman , Vitaly Shmatikov