English
Related papers

Related papers: Know your population and know your model: Using mo…

200 papers

Over the past decade network theory has been applied successfully to the study of a variety of complex adaptive systems. However, the application of these techniques to non-human social networks has several shortfalls. Firstly, in most…

Populations and Evolution · Quantitative Biology 2009-03-10 David Lusseau , Hal Whitehead , Shane Gero

Over the past ten years, propensity score methods have made an important contribution to improving generalizations from studies that do not select samples randomly from a population of inference. However, these methods require assumptions…

Methodology · Statistics 2021-01-26 Wendy Chan

User profiling means exploiting the technology of machine learning to predict attributes of users, such as demographic attributes, hobby attributes, preference attributes, etc. It's a powerful data support of precision marketing. Existing…

Computation and Language · Computer Science 2018-10-17 Yunpei Zheng , Lin Li , Luo Zhong , Jianwei Zhang , Jinhang Liu

Distributed learning provides an attractive framework for scaling the learning task by sharing the computational load over multiple nodes in a network. Here, we investigate the performance of distributed learning for large-scale linear…

Machine Learning · Statistics 2021-11-03 Martin Hellkvist , Ayça Özçelikkale , Anders Ahlén

The generalization of machine learning models has a complex dependence on the data, model and learning algorithm. We study train and test performance, as well as the generalization gap given by the mean of their difference over different…

Machine Learning · Statistics 2022-06-29 Carlos A. Gomez-Uribe

Basic personality traits are typically assessed through questionnaires. Here we consider phone-based metrics as a way to asses personality traits. We use data from smartphones with custom data-collection software distributed to 730…

Social and Information Networks · Computer Science 2018-03-06 Bjarke Mønsted , Anders Mollgaard , Joachim Mathiesen

Physical contacts result in the spread of various phenomena such as viruses, gossips, ideas, packages and marketing pamphlets across a population. The spread depends on how people move and co-locate with each other, or their mobility…

Social and Information Networks · Computer Science 2021-12-20 Sepanta Zeighami , Cyrus Shahabi , John Krumm

Experimental work regularly finds that individual choices are not deterministically rationalized by well-defined preferences. Nonetheless, recent work shows that data collected from many individuals can be stochastically rationalized by a…

Theoretical Economics · Economics 2021-10-22 Changkuk Im , John Rehbeck

We consider inference from non-random samples in data-rich settings where high-dimensional auxiliary information is available both in the sample and the target population, with survey inference being a special case. We propose a regularized…

Methodology · Statistics 2021-04-13 Yutao Liu , Andrew Gelman , Qixuan Chen

Throughout the COVID-19 pandemic, government policy and healthcare implementation responses have been guided by reported positivity rates and counts of positive cases in the community. The selection bias of these data calls into question…

Applications · Statistics 2021-04-12 Len Covello , Andrew Gelman , Yajuan Si , Siquan Wang

In aligning large language models (LLMs), reward models have played an important role, but are standardly trained as discriminative models and rely only on labeled human preference data. In this paper, we explore methods that train reward…

Computation and Language · Computer Science 2026-01-27 Chenglong Wang , Yang Gan , Yifu Huo , Yongyu Mu , Qiaozhi He , Murun Yang , Bei Li , Tong Xiao , Chunliang Zhang , Tongran Liu , Jingbo Zhu

Recent advances have shown that statistical tests for the rank of cross-covariance matrices play an important role in causal discovery. These rank tests include partial correlation tests as special cases and provide further graphical…

Machine Learning · Computer Science 2025-06-13 Xinshuai Dong , Ignavier Ng , Boyang Sun , Haoyue Dai , Guang-Yuan Hao , Shunxing Fan , Peter Spirtes , Yumou Qiu , Kun Zhang

We introduce a set of resampling-based methods for quantifying uncertainty and statistical precision of evaluation metrics in multilingual and/or multitask NLP benchmarks. We show how experimental variation in performance scores arises from…

Computation and Language · Computer Science 2025-12-19 Jonne Sälevä , Duygu Ataman , Constantine Lignos

Networks, representing attitudinal survey data, expose the structure of opinion-based groups. We make use of these network projections to identify the groups reliably through community detection algorithms and to examine…

Physics and Society · Physics 2021-10-14 Alejandro Dinkelberg , David JP O'Sullivan , Michael Quayle , Pádraig MacCarron

Social media provide access to behavioural data at an unprecedented scale and granularity. However, using these data to understand phenomena in a broader population is difficult due to their non-representativeness and the bias of…

Computers and Society · Computer Science 2019-05-16 Zijian Wang , Scott A. Hale , David Adelani , Przemyslaw A. Grabowicz , Timo Hartmann , Fabian Flöck , David Jurgens

We introduce multi-population opinion dynamics models linked to the bounded confidence model, aiming to explore how interactions between individuals contribute to the emergence of consensus, polarization, or fragmentation. Existing models…

Dynamical Systems · Mathematics 2023-11-29 Tigran Bakaryan , Yuliang Gu , Naira Hovakimyan , Tarek Abdelzaher , Christian Lebiere

The well-studied problem of statistical rank aggregation has been applied to comparing sports teams, information retrieval, and most recently to data generated by human judgment. Such human-generated rankings may be substantially different…

Information Retrieval · Computer Science 2014-11-05 Andrew Mao , Hossein Azari Soufiani , Yiling Chen , David C. Parkes

Examining residuals such as Pearson and deviance residuals, is a standard tool for assessing normal regression. However, for discrete response, these residuals cluster on lines corresponding to distinct response values. Their distributions…

Methodology · Statistics 2020-07-07 Cindy Feng , Alireza Sadeghpour , Longhai Li

We study the problem of learning from aggregate observations where supervision signals are given to sets of instances instead of individual instances, while the goal is still to predict labels of unseen individuals. A well-known example is…

Machine Learning · Statistics 2021-01-08 Yivan Zhang , Nontawat Charoenphakdee , Zhenguo Wu , Masashi Sugiyama

Big data sets must be carefully partitioned into statistically similar data subsets that can be used as representative samples for big data analysis tasks. In this paper, we propose the random sample partition (RSP) data model to represent…

Distributed, Parallel, and Cluster Computing · Computer Science 2019-06-11 Salman Salloum , Yulin He , Joshua Zhexue Huang , Xiaoliang Zhang , Tamer Z. Emara , Chenghao Wei , Heping He