中文
相关论文

相关论文: Efficient and Effective Spam Filtering and Re-rank…

200 篇论文

Malicious web content is a serious problem on the Internet today. In this paper we propose a deep learning approach to detecting malevolent web pages. While past work on web content detection has relied on syntactic parsing or on emulation…

密码学与安全 · 计算机科学 2018-04-16 Joshua Saxe , Richard Harang , Cody Wild , Hillary Sanders

The web contains large-scale, diverse, and abundant information to satisfy the information-seeking needs of humans. Through meticulous data collection, preprocessing, and curation, webpages can be used as a fundamental data resource for…

计算与语言 · 计算机科学 2024-06-18 Zhipeng Xu , Zhenghao Liu , Yukun Yan , Zhiyuan Liu , Ge Yu , Chenyan Xiong

In the context of depth-$k$ pooling for constructing web search test collections, we compare two approaches to ordering pooled documents for relevance assessors: the prioritisation strategy (PRI) used widely at NTCIR, and the simple…

信息检索 · 计算机科学 2022-11-03 Tetsuya Sakai , Sijie Tao , Zhaohao Zeng

Extracting structured information from HTML documents is a long-studied problem with a broad range of applications, including knowledge base construction, faceted search, and personalized recommendation. Prior works rely on a few…

信息检索 · 计算机科学 2022-08-30 Ritesh Sarkhel , Binxuan Huang , Colin Lockard , Prashant Shiralkar

The internet contains large amounts of low-quality content, yet users expect web search engines to deliver high-quality, relevant results. The abundant presence of low-quality pages can negatively impact retrieval and crawling processes by…

信息检索 · 计算机科学 2025-04-16 Francesca Pezzuti , Ariane Mueller , Sean MacAvaney , Nicola Tonellotto

As crowdsourcing emerges as an efficient and cost-effective method for obtaining labels for machine learning datasets, it is important to assess the quality of crowd-provided data, so as to improve analysis performance and reduce biases in…

人机交互 · 计算机科学 2025-06-26 Yang Ba , Michelle V. Mancenido , Erin K. Chiou , Rong Pan

Listwise reranking with large language models (LLMs) enhances top-ranked results in retrieval-based applications. Due to the limit in context size and high inference cost of long context, reranking is typically performed over a fixed size…

信息检索 · 计算机科学 2025-10-27 Soyoung Yoon , Gyuwan Kim , Gyu-Hwung Cho , Seung-won Hwang

Online reviews are a vital source of information when purchasing a service or a product. Opinion spammers manipulate these reviews, deliberately altering the overall perception of the service. Though there exists a corpus of online reviews,…

人工智能 · 计算机科学 2020-12-25 Athirai A. Irissappane , Hanfei Yu , Yankun Shen , Anubha Agrawal , Gray Stanton

The rise of large language models (LLMs) has enabled the generation of highly persuasive spam reviews that closely mimic human writing. These reviews pose significant challenges for existing detection systems and threaten the credibility of…

计算与语言 · 计算机科学 2026-04-21 Xin Liu , Rongwu Xu , Xinyi Jia , Jason Liao , Jiao Sun , Ling Huang , Wei Xu

Text-based communication is highly favoured as a communication method, especially in business environments. As a result, it is often abused by sending malicious messages, e.g., spam emails, to deceive users into relaying personal…

信息检索 · 计算机科学 2022-04-14 Annalisa Occhipinti , Louis Rogers , Claudio Angione

Pairing a lexical retriever with a neural re-ranking model has set state-of-the-art performance on large-scale information retrieval datasets. This pipeline covers scenarios like question answering or navigational queries, however, for…

信息检索 · 计算机科学 2022-10-20 Tim Baumgärtner , Leonardo F. R. Ribeiro , Nils Reimers , Iryna Gurevych

There is a tremendous increase in spam traffic these days. Spam messages muddle up users inbox, consume network resources, and build up DDoS attacks, spread worms and viruses. Our goal is to present a definite figure about the…

密码学与安全 · 计算机科学 2016-11-17 Cynthia Dhinakaran , Dhinaharan Nagamalai , Jae Kwang Lee

Decentralized unpermissioned peer-to-peer networks are inherently vulnerable to spam when they allow arbitrary participants to submit content to a common public index or registry; preventing this is difficult due to the absence of a central…

密码学与安全 · 计算机科学 2021-03-04 Alberto Inselvini

Nowadays, the size of the Internet is experiencing rapid growth. As of December 2014, the number of global Internet websites has more than 1 billion and all kinds of information resources are integrated together on the Internet, however,the…

分布式、并行与集群计算 · 计算机科学 2015-06-02 Qingpei Guo , Chao Xu , Yang Song

The proliferation of misinformation necessitates robust yet computationally efficient fact verification systems. While current state-of-the-art approaches leverage Large Language Models (LLMs) for generating explanatory rationales, these…

计算与语言 · 计算机科学 2025-11-10 Alamgir Munir Qazi , John P. McCrae , Jamal Abdul Nasir

Over the last years, online reviews became very important since they can influence the purchase decision of consumers and the reputation of businesses, therefore, the practice of writing fake reviews can have severe consequences on…

社会与信息网络 · 计算机科学 2020-12-29 Michela Fazzolari , Francesco Buccafurri , Gianluca Lax , Marinella Petrocchi

Machine unlearning for security is studied in this context. Several spam email detection methods exist, each of which employs a different algorithm to detect undesired spam emails. But these models are vulnerable to attacks. Many attackers…

机器学习 · 计算机科学 2021-12-28 Nishchal Parne , Kyathi Puppaala , Nithish Bhupathi , Ripon Patgiri

We argue that relationships between Web pages are functions of the user's intent. We identify a class of Web tasks - information-gathering - that can be facilitated by a search engine that provides links to pages which are related to the…

信息检索 · 计算机科学 2010-05-20 Amitabha Bagchi , Garima Lahoti

With malware detection techniques increasingly adopting machine learning approaches, the creation of precise training sets becomes more and more important. Large data sets of realistic web traffic, correctly classified as benign or…

密码学与安全 · 计算机科学 2018-02-19 Johann Vierthaler , Roman Kruszelnicki , Julian Schütte

The increasing reliance on online recruitment platforms coupled with the adoption of AI technologies has highlighted the critical need for efficient resume classification methods. However, challenges such as small datasets, lack of…

计算与语言 · 计算机科学 2024-07-16 Ahmed Heakl , Youssef Mohamed , Noran Mohamed , Aly Elsharkawy , Ahmed Zaky