English
Related papers

Related papers: Mining Hidden Populations through Attributed Searc…

200 papers

In this work we analyze the problem of, given the probability distribution of a population, questioning an unknown individual that is representative of the distribution so that our uncertainty about certain characteristics is significantly…

Computational Complexity · Computer Science 2026-01-22 David Pantoja , Ismael Rodriguez , Fernando Rubio , Clara Segura

Respondent-driven sampling is a survey method for hidden or hard-to-reach populations in which sampled individuals recruit others in the study population via their social links. The most popular estimator for for the population mean assumes…

Methodology · Statistics 2015-04-15 Peter M. Aronow , Forrest W. Crawford

A plethora of problems in AI, engineering and the sciences are naturally formalized as inference in discrete probabilistic models. Exact inference is often prohibitively expensive, as it may require evaluating the (unnormalized) target…

Machine Learning · Computer Science 2019-10-16 Lars Buesing , Nicolas Heess , Theophane Weber

Millions of people express themselves on public social media, such as Twitter. Through their posts, these people may reveal themselves as potentially valuable sources of information. For example, real-time information about an event might…

Social and Information Networks · Computer Science 2014-04-09 Jalal Mahmud , Michelle Zhou , Nimrod Megiddo , Jeffrey Nichols , Clemens Drews

Diversified recommendation has attracted increasing attention from both researchers and practitioners, which can effectively address the homogeneity of recommended items. Existing approaches predominantly aim to infer the diversity of user…

Information Retrieval · Computer Science 2026-01-07 Hanyang Yuan , Ning Tang , Tongya Zheng , Jiarong Xu , Xintong Hu , Renhong Huang , Shunyu Liu , Jiacong Hu , Jiawei Chen , Mingli Song

The recent proliferation of research into transformer based natural language processing has led to a number of studies which attempt to detect the presence of human-like cognitive behavior in the models. We contend that, as is true of human…

Computation and Language · Computer Science 2024-04-01 Jesse Roberts , Kyle Moore , Drew Wilenzick , Doug Fisher

Algorithms deployed in education can shape the learning experience and success of a student. It is therefore important to understand whether and how such algorithms might create inequalities or amplify existing biases. In this paper, we…

Computers and Society · Computer Science 2022-12-21 Jade Maï Cock , Muhammad Bilal , Richard Davis , Mirko Marras , Tanja Käser

We analyze dynamic random network models where younger vertices connect to older ones with probabilities proportional to their degrees as well as a propensity kernel governed by their attribute types. Using stochastic approximation…

Probability · Mathematics 2025-10-29 Nelson Antunes , Sayan Banerjee , Shankar Bhamidi , Vladas Pipiras

Social networks are typical attributed networks with node attributes. Different from traditional attribute community detection problem aiming at obtaining the whole set of communities in the network, we study an application-oriented problem…

Social and Information Networks · Computer Science 2017-05-11 Peng Wu , Li Pan

Crowdsourcing employs human workers to solve computer-hard problems, such as data cleaning, entity resolution, and sentiment analysis. When crowdsourcing tabular data, e.g., the attribute values of an entity set, a worker's answers on the…

Databases · Computer Science 2017-08-08 Caihua Shan , Nikos Mamoulis , Guoliang Li , Reynold Cheng , Zhipeng Huang , Yudian Zheng

A network has a non-overlapping community structure if the nodes of the network can be partitioned into disjoint sets such that each node in a set is densely connected to other nodes inside the set and sparsely connected to the nodes out-…

Social and Information Networks · Computer Science 2016-07-19 Talasila Sai Deepak , Hindol Adhya , Shyamal Kejriwal , Bhanuteja Gullapalli , Saswata Shannigrahi

Thompson sampling has emerged as an effective heuristic for a broad range of online decision problems. In its basic form, the algorithm requires computing and sampling from a posterior distribution over models, which is tractable only for…

Machine Learning · Statistics 2023-04-26 Xiuyuan Lu , Benjamin Van Roy

In theory, a major advantage to the big data approach in studying online communities is that it should be possible to collect a representative random sample from a broadly defined population. However, in practice, data collection processes…

Social and Information Networks · Computer Science 2021-02-02 Muhammad Umer Gurchani

A hidden database refers to a dataset that an organization makes accessible on the web by allowing users to issue queries through a search interface. In other words, data acquisition from such a source is not by following static…

Databases · Computer Science 2012-08-02 Cheng Sheng , Nan Zhang , Yufei Tao , Xin Jin

We propose the Heterogeneous Thurstone Model (HTM) for aggregating ranked data, which can take the accuracy levels of different users into account. By allowing different noise distributions, the proposed HTM model maintains the generality…

Machine Learning · Computer Science 2019-12-04 Tao Jin , Pan Xu , Quanquan Gu , Farzad Farnoud

The goal of our research is to contribute information about how useful the crowd is at anticipating stereotypes that may be biasing a data set without a researcher's knowledge. The results of the crowd's prediction can potentially be used…

Human-Computer Interaction · Computer Science 2018-01-11 Zeyuan Hu , Julia Strout

Finding influential users in online social networks is an important problem with many possible useful applications. HITS and other link analysis methods, in particular, have been often used to identify hub and authority users in web graphs…

Social and Information Networks · Computer Science 2018-02-21 Roy Ka-Wei Lee , Tuan-Anh Hoang , Ee-Peng Lim

Even when aggregate accuracy is high, state-of-the-art NLP models often fail systematically on specific subgroups of data, resulting in unfair outcomes and eroding user trust. Additional data collection may not help in addressing these…

Computation and Language · Computer Science 2023-05-30 Zexue He , Marco Tulio Ribeiro , Fereshte Khani

Social media data provides propitious opportunities for public health research. However, studies suggest that disparities may exist in the representation of certain populations (e.g., people of lower socioeconomic status). To quantify and…

Computers and Society · Computer Science 2017-11-07 Nina Cesare , Christan Grant , Jared B. Hawkins , John S. Brownstein , Elaine O. Nsoesie

Incorporating graph side information into recommender systems has been widely used to better predict ratings, but relatively few works have focused on theoretical guarantees. Ahn et al. (2018) firstly characterized the optimal sample…

Information Theory · Computer Science 2021-09-09 Changhun Jo , Kangwook Lee