English
Related papers

Related papers: On The Problem of Relevance in Statistical Inferen…

200 papers

In this comment, I discuss the use of statistical inference in citation analysis. In a recent paper, Williams and Bornmann argue in favor of the use of statistical inference in citation analysis. I present a critical analysis of their…

Digital Libraries · Computer Science 2016-12-13 Ludo Waltman

Sequence modeling faces challenges in capturing long-range dependencies across diverse tasks. Recent linear and transformer-based forecasters have shown superior performance in time series forecasting. However, they are constrained by their…

Machine Learning · Computer Science 2024-11-25 Bong Gyun Kang , Dongjun Lee , HyunGi Kim , DoHyun Chung , Sungroh Yoon

The problem of overdispersion in multivariate count data is a challenging issue. Nowadays, it covers a central role mainly due to the relevance of modern technologies data, such as Next Generation Sequencing and textual data from the web or…

Methodology · Statistics 2025-02-24 Noemi Corsini , Cinzia Viroli

Positional bias in large language models (LLMs) hinders their ability to effectively process long inputs. A prominent example is the "lost in the middle" phenomenon, where LLMs struggle to utilize relevant information situated in the middle…

Computation and Language · Computer Science 2025-05-29 Runchu Tian , Yanghao Li , Yuepeng Fu , Siyang Deng , Qinyu Luo , Cheng Qian , Shuo Wang , Xin Cong , Zhong Zhang , Yesai Wu , Yankai Lin , Huadong Wang , Xiaojiang Liu

Classifying samples in incomplete datasets is a common aim for machine learning practitioners, but is non-trivial. Missing data is found in most real-world datasets and these missing values are typically imputed using established methods,…

This paper gives a brief description of the author's database of integer sequences, now over 35 years old, together with a selection of a few of the most interesting sequences in the table. Many unsolved problems are mentioned.

Combinatorics · Mathematics 2007-05-23 N. J. A. Sloane

As a powerful representation paradigm for networked and multi-typed data, the heterogeneous information network (HIN) is ubiquitous. Meanwhile, defining proper relevance measures has always been a fundamental problem and of great pragmatic…

Social and Information Networks · Computer Science 2019-02-22 Yu Shi , Po-Wei Chan , Honglei Zhuang , Huan Gui , Jiawei Han

This paper re-examines the first normalized incomplete moment, a well-established measure of inequality with wide applications in economic and social sciences. Despite the popularity of the measure itself, existing statistical inference…

Methodology · Statistics 2025-08-26 Jiannan Lu , Peng Ding , Anqi Zhao

Credal networks are graph-based statistical models whose parameters take values in a set, instead of being sharply specified as in traditional statistical models (e.g., Bayesian networks). The computational complexity of inferences on such…

Artificial Intelligence · Computer Science 2013-09-27 Denis D. Maua , Cassio Polpo de Campos , Alessio Benavoli , Alessandro Antonucci

Being able to predict the length of a scientific paper may be helpful in numerous situations. This work defines the paper length prediction task as a regression problem and reports several experimental results using popular machine learning…

Computation and Language · Computer Science 2020-12-18 Erion Çano , Ondřej Bojar

The replication crisis has prompted many to call for statistical reform within the psychological sciences. Here we examine issues within Frequentist statistics that may have led to the replication crisis, and we examine the…

Methodology · Statistics 2018-11-09 Lincoln J Colling , Denes Szucs

Relation extraction is the task of determining the relation between two entities in a sentence. Distantly-supervised models are popular for this task. However, sentences can be long and two entities can be located far from each other in a…

Computation and Language · Computer Science 2019-12-10 Tapas Nayak , Hwee Tou Ng

Inference for impulse responses estimated with local projections presents interesting challenges and opportunities. Analysts typically want to assess the precision of individual estimates, explore the dynamic evolution of the response over…

Econometrics · Economics 2024-08-15 Atsushi Inoue , Òscar Jordà , Guido M. Kuersteiner

We consider the problem of statistical inference for ranking data, specifically rank aggregation, under the assumption that samples are incomplete in the sense of not comprising all choice alternatives. In contrast to most existing methods,…

Machine Learning · Statistics 2017-12-05 Mohsen Ahmadi Fahandar , Eyke Hüllermeier , Inés Couso

The most fundamental problem in statistics is the inference of an unknown probability distribution from a finite number of samples. For a specific observed data set, answers to the following questions would be desirable: (1) Estimation:…

Statistics Theory · Mathematics 2013-01-23 Ali Kinkhabwala

Landmarks are facts or actions that appear in all valid solutions of a planning problem. They have been used successfully to calculate heuristics that guide the search for a plan. We investigate an extension to this concept by defining a…

Artificial Intelligence · Computer Science 2024-03-13 Oliver Kim , Mohan Sridharan

This article discusses a number of incorrect statements appearing in textbooks on data analysis, machine learning, or computational methods; the common theme in all these cases is the relevance and application of statistics to the study of…

Data Analysis, Statistics and Probability · Physics 2023-01-13 Alexandros Gezerlis , Martin Williams

This paper draws attention to the potential of computational methods in reworking data generated in past qualitative studies. While qualitative inquiries often produce rich data through rigorous and resource-intensive processes, much of…

Databases · Computer Science 2025-06-06 Kaveh Mohajeri , Amir Karami

This work looks in depth at several studies that have attempted to automate the process of citation importance classification based on the publications full text. We analyse a range of features that have been previously used in this task.…

Digital Libraries · Computer Science 2017-07-14 David Pride , Petr Knoth

Linear quantile regression models aim at providing a detailed and robust picture of the (conditional) response distribution as function of a set of observed covariates. Longitudinal data represent an interesting field of application of such…

Methodology · Statistics 2015-07-30 Maria Francesca Marino , Nikos Tzavidis , Marco Alfo'