English
Related papers

Related papers: Intriguing behavior when testing the impact of quo…

200 papers

Search-engine date filters are widely used to enforce pre-cutoff retrieval in retrospective evaluations of search-augmented forecasters. We show this approach is unreliable across two major search engines: auditing Google Search's before:…

Computation and Language · Computer Science 2026-04-22 Ali El Lahib , Ying-Jieh Xia , Zehan Li , Yuxuan Wang , Xinyu Pi

We applied a set of standard bibliometric indicators to monitor the scientific state-of-arte of 500 universities worldwide and constructed a ranking on the basis of these indicators (Leiden Ranking 2010). We find a dramatic and hitherto…

Digital Libraries · Computer Science 2010-12-24 Anthony F. J. van Raan , Thed N. van Leeuwen , Martijn S. Visser

The intriguing law of anomalous numbers, also named Benford's law, states that the significant digits of data follow a logarithmic distribution favoring the smallest values. In this work, we test the compliance with this law of the atomic…

Atomic Physics · Physics 2024-05-30 Jean-Christophe Pain , Yuri Ralchenko

Our goal is to distinguish between the following two hypotheses: (A) The Internet will remain disproportionately in English and will, over time, cause more people to learn English as second language and thus solidify the role of English as…

Computers and Society · Computer Science 2007-05-23 Neil Gandal , Carl Shapiro

Benford's law is widely used for fraud-detection nowadays. The underlying assumption for using the law is that a "regular" dataset follows the significant digit phenomenon. In this paper, we address the scenario where a shrewd fraudster…

Applications · Statistics 2021-05-21 Javad Kazemitabar

The Newcomb-Benford law, also known as the first-digit law, gives the probability distribution associated with the first digit of a dataset, so that, for example, the first significant digit has a probability of $30.1$ % of being $1$ and…

Popular Physics · Physics 2021-08-25 Andrea Burgos , Andrés Santos

In this work, we introduce a novel metric for auditing group fairness in ranked lists. Our approach offers two benefits compared to the state of the art. First, we offer a blueprint for modeling of user attention. Rather than assuming a…

Computers and Society · Computer Science 2019-05-14 Piotr Sapiezynski , Wesley Zeng , Ronald E. Robertson , Alan Mislove , Christo Wilson

The correlation between the demographics of users and the text they write has been investigated through literary texts and, more recently, social media. However, differences pertaining to language use in search engines has not been…

Computers and Society · Computer Science 2018-05-24 Elad Yom-Tov

Search engines decide what we see for a given search query. Since many people are exposed to information through search engines, it is fair to expect that search engines are neutral. However, search engine results do not necessarily cover…

Information Retrieval · Computer Science 2023-02-07 Gizem Gezici , Aldo Lipani , Yucel Saygin , Emine Yilmaz

The proliferation of surveys and review articles in academic journals has impacted citation metrics like impact factor and h-index, skewing evaluations of journal and researcher quality. This work investigates the implications of this…

Digital Libraries · Computer Science 2025-04-09 Jesus S. Aguilar-Ruiz

It is tempting to treat frequency trends from the Google Books data sets as indicators of the "true" popularity of various words and phrases. Doing so allows us to draw quantitatively strong conclusions about the evolution of cultural…

Physics and Society · Physics 2020-05-28 Eitan Adam Pechenick , Christopher M. Danforth , Peter Sheridan Dodds

Nonextensive statistics, characterized by a nonextensive parameter $q$, is a promising and practically useful generalization of the Boltzmann statistics to describe power-law behaviors from physical and social observations. We here explore…

Data Analysis, Statistics and Probability · Physics 2011-03-07 Lijing Shao , Bo-Qiang Ma

We carried out a retrieval effectiveness test on the three major web search engines (i.e., Google, Microsoft and Yahoo). In addition to relevance judgments, we classified the results according to their commercial intent and whether or not…

Information Retrieval · Computer Science 2015-11-19 Dirk Lewandowski

In this paper, we propose a measure to assess scientific impact that discounts self-citations and does not require any prior knowledge on the their distribution among publications. This index can be applied to both researchers and journals.…

Digital Libraries · Computer Science 2013-10-17 Emilio Ferrara , Alfonso E. Romero

When interacting with information retrieval (IR) systems, users, affected by confirmation biases, tend to select search results that confirm their existing beliefs on socially significant contentious issues. To understand the judgments and…

Information Retrieval · Computer Science 2024-06-18 Ben Wang , Jiqun Liu

The impact of scientific publications has traditionally been expressed in terms of citation counts. However, scientific activity has moved online over the past decade. To better capture scientific impact in the digital era, a variety of new…

Digital Libraries · Computer Science 2009-06-30 Johan Bollen , Herbert Van de Sompel , Aric Hagberg , Ryan Chute

There is inherent information captured in the order in which we write words in a list. The orderings of binomials --- lists of two words separated by `and' or `or' --- has been studied for more than a century. These binomials are common…

Social and Information Networks · Computer Science 2020-03-10 Katherine Van Koevering , Austin R. Benson , Jon Kleinberg

Purpose: To compare five major Web search engines (Google, Yahoo, MSN, Ask.com, and Seekport) for their retrieval effectiveness, taking into account not only the results but also the results descriptions. Design/Methodology/Approach: The…

Information Retrieval · Computer Science 2015-11-19 Dirk Lewandowski

This paper extends Becker (1957)'s outcome test of discrimination to settings where a (human or algorithmic) decision-maker produces a ranked list of candidates. Ranked lists are particularly relevant in the context of online platforms that…

Econometrics · Economics 2021-11-16 Jonathan Roth , Guillaume Saint-Jacques , YinYin Yu

The main objective of this paper is to empirically test whether the identification of highly-cited documents through Google Scholar is feasible and reliable. To this end, we carried out a longitudinal analysis (1950 to 2013), running a…

Digital Libraries · Computer Science 2018-04-30 Alberto Martín-Martín , Enrique Orduna-Malea , Anne-Wil Harzing , Emilio Delgado López-Cózar