English
Related papers

Related papers: Efficient and Effective Spam Filtering and Re-rank…

200 papers

A focused crawler aims at discovering as many web pages and web sites relevant to a target topic as possible, while avoiding irrelevant ones. Reinforcement Learning (RL) has been a promising direction for optimizing focused crawling,…

Information Retrieval · Computer Science 2025-05-20 Andreas Kontogiannis , Dimitrios Kelesis , Vasilis Pollatos , George Giannakopoulos , Georgios Paliouras

We present CWRCzech, Click Web Ranking dataset for Czech, a 100M query-document Czech click dataset for relevance ranking with user behavior data collected from search engine logs of Seznam$.$cz. To the best of our knowledge, CWRCzech is…

Information Retrieval · Computer Science 2024-07-16 Josef Vonášek , Milan Straka , Rostislav Krč , Lenka Lasoňová , Ekaterina Egorova , Jana Straková , Jakub Náplava

Email continues to be a pivotal and extensively utilized communication medium within professional and commercial domains. Nonetheless, the prevalence of spam emails poses a significant challenge for users, disrupting their daily routines…

Computation and Language · Computer Science 2025-02-13 Shijing Si , Yuwei Wu , Le Tang , Yugui Zhang , Jedrek Wosik , Qinliang Su

Large language models(LLMs) have demonstrated remarkable performance on many natural language processing(NLP) tasks and have been employed in phishing email detection research. However, in current studies, well-performing LLMs typically…

Computation and Language · Computer Science 2025-05-06 Zijie Lin , Zikang Liu , Hanbo Fan

Search Engine Result Pages (SERPs) serve as the digital gateways to the vast expanse of the internet. Past decades have witnessed a surge in research primarily centered on the influence of website ranking on these pages, to determine the…

Information Retrieval · Computer Science 2023-06-06 Erik Fubel , Niclas Michael Groll , Patrick Gundlach , Qiwei Han , Maximilian Kaiser

Large Language Models (LLMs) have demonstrated exceptional performance across diverse tasks, yet their training remains highly resource-intensive and susceptible to critical challenges such as training instability. A predominant source of…

Machine Learning · Computer Science 2025-03-03 Tianjin Huang , Ziquan Zhu , Gaojie Jin , Lu Liu , Zhangyang Wang , Shiwei Liu

In information retrieval (IR), candidate set pruning has been commonly used to speed up two-stage relevance ranking. However, such an approach lacks accurate error control and often trades accuracy off against computational efficiency in an…

Information Retrieval · Computer Science 2022-05-20 Minghan Li , Xinyu Zhang , Ji Xin , Hongyang Zhang , Jimmy Lin

High-quality main content extraction from web pages is a critical prerequisite for constructing large-scale training corpora. While traditional heuristic extractors are efficient, they lack the semantic reasoning required to handle the…

Backdoor attacks pose a significant threat to the integrity of text classification models used in natural language processing. While several dirty-label attacks that achieve high attack success rates (ASR) have been proposed, clean-label…

Cryptography and Security · Computer Science 2025-08-25 Onur Alp Kirci , M. Emre Gursoy

Unsolicited bulk email (aka. spam) is a major problem on the Internet. To counter spam, several techniques, ranging from spam filters to mail protocol extensions like hashcash, have been proposed. In this paper we investigate the…

Cryptography and Security · Computer Science 2007-05-23 Flavio D. Garcia , Jaap-Henk Hoepman

In This paper we present a novel approach to spam filtering and demonstrate its applicability with respect to SMS messages. Our approach requires minimum features engineering and a small set of la- belled data samples. Features are…

Computation and Language · Computer Science 2016-06-20 Noura Al Moubayed , Toby Breckon , Peter Matthews , A. Stephen McGough

Pre-training text representations have led to significant improvements in many areas of natural language processing. The quality of these models benefits greatly from the size of the pretraining corpora as long as its quality is preserved.…

Computation and Language · Computer Science 2019-11-18 Guillaume Wenzek , Marie-Anne Lachaux , Alexis Conneau , Vishrav Chaudhary , Francisco Guzmán , Armand Joulin , Edouard Grave

Email spam detection is a critical task in modern communication systems, essential for maintaining productivity, security, and user experience. Traditional machine learning and deep learning approaches, while effective in static settings,…

Cryptography and Security · Computer Science 2025-05-06 Ghazaleh SHirvani , Saeid Ghasemshirazi

Training robust retrieval and reranker models typically relies on large-scale retrieval datasets; for example, the BGE collection contains 1.6 million query-passage pairs sourced from various data sources. However, we find that certain…

Information Retrieval · Computer Science 2025-10-21 Nandan Thakur , Crystina Zhang , Xueguang Ma , Jimmy Lin

Robust machine learning depends on clean data, yet current image data cleaning benchmarks rely on synthetic noise or narrow human studies, limiting comparison and real-world relevance. We introduce CleanPatrick, the first large-scale…

In this paper, we propose a web search retrieval approach which automatically detects recency sensitive queries and increases the freshness of the ordinary document ranking by a degree proportional to the probability of the need in recent…

Information Retrieval · Computer Science 2024-02-08 Andrey Styskin , Fedor Romanenko , Fedor Vorobyev , Pavel Serdyukov

Efficient code retrieval is critical for developer productivity, yet existing benchmarks largely focus on Python and rarely stress-test robustness beyond superficial lexical cues. To address the gap, we introduce an automated pipeline for…

Software Engineering · Computer Science 2026-03-06 Kaicheng Wang , Liyan Huang , Weike Fang , Weihang Wang

Debate portals and similar web platforms constitute one of the main text sources in computational argumentation research and its applications. While the corpora built upon these sources are rich of argumentatively relevant content and…

Computation and Language · Computer Science 2020-11-04 Jonas Dorsch , Henning Wachsmuth

A web crawler is a system designed to collect web pages, and efficient crawling of new pages requires appropriate algorithms. While website features such as XML sitemaps and the frequency of past page updates provide important clues for…

Information Retrieval · Computer Science 2025-05-13 Yuichi Sasazawa , Yasuhiro Sogawa

Data plays the most prominent role in how language models acquire skills and knowledge. The lack of massive, well-organized pre-training datasets results in costly and inaccessible data pipelines. We present Essential-Web v1.0, a…