English
Related papers

Related papers: Temporal Leakage in Search-Engine Date-Filtered We…

200 papers

In financial backtesting, large language models pretrained on internet-scale data risk introducing lookahead bias that undermines their forecasting validity, as they may have already seen the true outcome during training. To address this,…

Computation and Language · Computer Science 2026-03-13 Yutong Yan , Raphael Tang , Zhenyu Gao , Wenxi Jiang , Yao Lu

Modern large language models integrate web search to provide real-time answers, yet it remains unclear whether they are efficiently calibrated to use search when it is actually needed. We introduce a benchmark evaluating both the necessity…

Computation and Language · Computer Science 2025-11-25 Sahil Kale

Recommender models are hard to evaluate, particularly under offline setting. In this paper, we provide a comprehensive and critical analysis of the data leakage issue in recommender system offline evaluation. Data leakage is caused by not…

Information Retrieval · Computer Science 2023-08-07 Yitong Ji , Aixin Sun , Jie Zhang , Chenliang Li

Backtesting large language models on historical events requires reasoning exclusively from information available before a specified cutoff date. Yet models routinely leak post-cutoff knowledge from pre-training into their reasoning,…

Machine Learning · Computer Science 2026-05-20 Zeyu Zhang , Bradly C. Stadie

Search engine retrieval effectiveness studies are usually small-scale, using only limited query samples. Furthermore, queries are selected by the researchers. We address these issues by taking a random representative sample of 1,000…

Information Retrieval · Computer Science 2014-05-12 Dirk Lewandowski

When it comes to search engines, users generally prefer Google. Our study aims to find the differences between the results found in Google compared to other search engines. We compared the top 10 results from Google, Bing, DuckDuckGo, and…

Information Retrieval · Computer Science 2022-07-18 Nurce Yagci , Sebastian Sünkler , Helena Häußler , Dirk Lewandowski

Released Large Language Models (LLMs) are often paired with a claimed knowledge cutoff date, or the dates at which training data was gathered. Such information is crucial for applications where the LLM must provide up to date information.…

Computation and Language · Computer Science 2024-09-18 Jeffrey Cheng , Marc Marone , Orion Weller , Dawn Lawrie , Daniel Khashabi , Benjamin Van Durme

Cloud data storage solutions offer customers cost-effective and reduced data management. While attractive, data security issues remain to be a core concern. Traditional encryption protects stored documents, but hinders simple…

Cryptography and Security · Computer Science 2023-06-28 Marc Damie , Florian Hahn , Andreas Peter

Survey-based studies suggest that search engines are trusted more than social media or even traditional news, although cases of false information or defamation are known. In this study, we analyze query suggestion features of three search…

Information Retrieval · Computer Science 2019-12-03 Malte Bonart , Anastasiia Samokhina , Gernot Heisenberg , Philipp Schaer

To enhance documentation and maintenance practices, developers conventionally establish links between related software artifacts manually. Empirical research has revealed that developers frequently overlook this practice, resulting in…

Software Engineering · Computer Science 2023-04-25 Maliheh Izadi , Pooya Rostami Mazrae , Tom Mens , Arie van Deursen

Using non-linear machine learning methods and a proper backtest procedure, we critically examine the claim that Google Trends can predict future price returns. We first review the many potential biases that may influence backtests with this…

Trading and Market Microstructure · Quantitative Finance 2014-03-10 Damien Challet , Ahmed Bel Hadj Ayed

Because the World Wide Web is a dynamic collection of information, the Web search tools (or "search engines") that index the Web are dynamic. Traditional information retrieval evaluation techniques may not provide reliable results when…

Information Retrieval · Computer Science 2007-05-23 Scott Nicholson

Large pre-trained models can dramatically reduce the amount of task-specific data required to solve a problem, but they often fail to capture domain-specific nuances out of the box. The Web likely contains the information necessary to excel…

Computer Vision and Pattern Recognition · Computer Science 2023-12-01 Hamed Damirchi , Cristian Rodríguez-Opazo , Ehsan Abbasnejad , Damien Teney , Javen Qinfeng Shi , Stephen Gould , Anton van den Hengel

Generative search engines have the potential to transform how people seek information online, but generated responses from existing large language models (LLMs)-backed generative search engines may not always be accurate. Nonetheless,…

Computation and Language · Computer Science 2024-03-20 Xuming Hu , Xiaochuan Li , Junzhe Chen , Yinghui Li , Yangning Li , Xiaoguang Li , Yasheng Wang , Qun Liu , Lijie Wen , Philip S. Yu , Zhijiang Guo

Batches of pharmaceutical are sometimes recalled from the market when a safety issue or a defect is detected in specific production runs of a drug. Such problems are usually detected when patients or healthcare providers report…

Information Retrieval · Computer Science 2018-05-16 Elad Yom-Tov

How different are search engines? The search engine wars are a favorite topic of on-line analysts, as two of the biggest companies in the world, Google and Microsoft, battle for prevalence of the web search space. Differences in search…

Information Retrieval · Computer Science 2020-11-09 Konstantina Dritsa , Thodoris Sotiropoulos , Haris Skarpetis , Panos Louridas

The conspiracy theory that the US 2020 presidential election was fraudulent - the Big Lie - remained a prominent part of the media agenda months after the election. Whether and how search engines prioritized news stories that sought to…

Information Retrieval · Computer Science 2024-09-18 Ernesto de León , Mykola Makhortykh , Aleksandra Urman , Roberto Ulloa

Information retrieval systems are crucial for enabling effective access to large document collections. Recent approaches have leveraged Large Language Models (LLMs) to enhance retrieval performance through query augmentation, but often rely…

Information Retrieval · Computer Science 2025-04-15 Pengcheng Jiang , Jiacheng Lin , Lang Cao , Runchu Tian , SeongKu Kang , Zifeng Wang , Jimeng Sun , Jiawei Han

Recall, the proportion of relevant documents retrieved, is an important measure of effectiveness in information retrieval, particularly in the legal, patent, and medical domains. Where document sets are too large for exhaustive relevance…

Information Retrieval · Computer Science 2012-10-25 William Webber

The rapid dissemination of information through social media and the Internet has posed a significant challenge for fact-checking, among others in identifying check-worthy claims that fact-checkers should pay attention to, i.e. filtering…

Computation and Language · Computer Science 2024-06-27 Yufeng Li , Rrubaa Panchendrarajan , Arkaitz Zubiaga
‹ Prev 1 2 3 10 Next ›