中文
相关论文

相关论文: Efficient and Effective Spam Filtering and Re-rank…

200 篇论文

Recently, the development and implementation of phishing attacks require little technical skills and costs. This uprising has led to an ever-growing number of phishing attacks on the World Wide Web. Consequently, proactive techniques to…

密码学与安全 · 计算机科学 2020-11-09 Chidimma Opara , Bo Wei , Yingke Chen

Clickthrough data is a particularly inexpensive and plentiful resource to obtain implicit relevance feedback for improving and personalizing search engines. However, it is well known that the probability of a user clicking on a result is…

信息检索 · 计算机科学 2007-05-23 Filip Radlinski , Thorsten Joachims

The World Wide Web (WWW) is the repository of large number of web pages which can be accessed via Internet by multiple users at the same time and therefore it is Ubiquitous in nature. The search engine is a key application used to search…

数据库 · 计算机科学 2012-09-25 K. C. Srikantaiah , P. L. Srikanth , V. Tejaswi , K. Shaila , K. R. Venugopal , L. M. Patnaik

Modern phishing campaigns increasingly evade snapshot-based URL classifiers using interaction gates (e.g., checkbox/slider challenges), delayed content rendering, and logo-less credential harvesters. This shifts URL triage from static…

密码学与安全 · 计算机科学 2026-04-24 Haolin Zhang , William Reber , Yuxuan Zhang , Guofei Gu , Jeff Huang

Simulated phishing campaigns are widely deployed, yet the behavioral data they produce is endogenous: because training is triggered by clicking, the employees receiving intervention have already demonstrated susceptibility. This…

密码学与安全 · 计算机科学 2026-03-05 Muhammad Zia Hydari , Idris Adjerid , Yingda Lu , Narayan Ramasubbu

The advances in digital tools have led to the rampant spread of misinformation. While fact-checking aims to combat this, manual fact-checking is cumbersome and not scalable. It is essential for automated fact-checking to be efficient for…

信息检索 · 计算机科学 2025-02-18 Kevin Nanekhan , Venktesh V , Erik Martin , Henrik Vatndal , Vinay Setty , Avishek Anand

Large Language Models (LLMs) trained on historical web data inevitably become outdated. We investigate evaluation strategies and update methods for LLMs as new data becomes available. We introduce a web-scale dataset for time-continual…

Anomalies in emails such as phishing and spam present major security risks such as the loss of privacy, money, and brand reputation to both individuals and organizations. Previous studies on email anomaly detection relied on a single type…

密码学与安全 · 计算机科学 2022-03-22 Craig Beaman , Haruna Isah

The world is full of text data, yet text analytics has not traditionally played a large part in statistics education. We consider four different ways to provide students with opportunities to explore whether email messages are unwanted…

其他统计学 · 统计学 2022-10-11 Nicholas J. Horton , Jie Chao , William Finzer , Phebe Palmer

The reviews of customers play an essential role in online shopping. People often refer to reviews or comments of previous customers to decide whether to buy a new product. Catching up with this behavior, some people create untruths and…

计算与语言 · 计算机科学 2022-12-12 Co Van Dinh , Son T. Luu , Anh Gia-Tuan Nguyen

Recently, neural models have been leveraged to significantly improve the performance of information extraction from semi-structured websites. However, a barrier for continued progress is the small number of datasets large enough to train…

计算与语言 · 计算机科学 2023-06-16 Aidan San , Yuan Zhuang , Jan Bakus , Colin Lockard , David Ciemiewicz , Sandeep Atluri , Yangfeng Ji , Kevin Small , Heba Elfardy

Cybercriminals have leveraged the popularity of a large user base available on Online Social Networks to spread spam campaigns by propagating phishing URLs, attaching malicious contents, etc. However, another kind of spam attacks using…

社会与信息网络 · 计算机科学 2018-02-13 Srishti Gupta , Abhinav Khattar , Arpit Gogia , Ponnurangam Kumaraguru , Tanmoy Chakraborty

Due to the huge commercial interests behind online reviews, a tremendousamount of spammers manufacture spam reviews for product reputation manipulation. To further enhance the influence of spam reviews, spammers often collaboratively post…

信息检索 · 计算机科学 2020-11-17 Ziyang Wang , Wei Wei , Xian-Ling Mao , Guibing Guo , Pan Zhou , Shanshan Feng

Modern web applications rely heavily on client-side API calls to fetch data, render content, and communicate with backend services. However, the quality of these network interactions (redundant requests, missing cache headers, oversized…

软件工程 · 计算机科学 2026-02-19 Ali Hassaan Mughal , Muhammad Bilal , Noor Fatima

Improved search quality enhances users' satisfaction, which directly impacts sales growth of an E-Commerce (E-Com) platform. Traditional Learning to Rank (LTR) algorithms require relevance judgments on products. In E-Com, getting such…

信息检索 · 计算机科学 2020-07-10 Muhammad Umer Anwaar , Dmytro Rybalko , Martin Kleinsteuber

This paper presents a framework for increasing the relevancy of the web pages retrieved by the search engine. The approach introduces a Predictive Prefetching Engine (PPE) which makes use of various data mining algorithms on the log…

信息检索 · 计算机科学 2011-09-29 Jyoti , A. K. Sharma , Amit Goel

Motivated by recent commentary that has questioned today's pursuit of ever-more complex models and mathematical formalisms in applied machine learning and whether meaningful empirical progress is actually being made, this paper tries to…

信息检索 · 计算机科学 2019-04-19 Jimmy Lin

This paper presents our proposed approach that won the first prize at the ICLR competition on Hardware Aware Efficient Training. The challenge is to achieve the highest possible accuracy in an image classification task in less than 10…

机器学习 · 计算机科学 2025-05-27 Omar Mohamed Awad , Habib Hajimolahoseini , Michael Lim , Gurpreet Gosal , Walid Ahmed , Yang Liu , Gordon Deng

Systematic literature reviews (SLRs) are essential but labor-intensive due to high publication volumes and inefficient keyword-based filtering. To streamline this process, we evaluate Large Language Models (LLMs) for enhancing efficiency…

机器学习 · 计算机科学 2025-06-17 Lucas Joos , Daniel A. Keim , Maximilian T. Fischer

Large language models are increasingly used for many applications. To prevent illicit use, it is desirable to be able to detect AI-generated text. Training and evaluation of such detectors critically depend on suitable benchmark datasets.…

机器学习 · 计算机科学 2025-11-13 Philipp Dingfelder , Christian Riess