中文
相关论文

相关论文: Mining Hidden Populations through Attributed Searc…

200 篇论文

In this work we analyze the problem of, given the probability distribution of a population, questioning an unknown individual that is representative of the distribution so that our uncertainty about certain characteristics is significantly…

计算复杂性 · 计算机科学 2026-01-22 David Pantoja , Ismael Rodriguez , Fernando Rubio , Clara Segura

Respondent-driven sampling is a survey method for hidden or hard-to-reach populations in which sampled individuals recruit others in the study population via their social links. The most popular estimator for for the population mean assumes…

统计方法学 · 统计学 2015-04-15 Peter M. Aronow , Forrest W. Crawford

A plethora of problems in AI, engineering and the sciences are naturally formalized as inference in discrete probabilistic models. Exact inference is often prohibitively expensive, as it may require evaluating the (unnormalized) target…

机器学习 · 计算机科学 2019-10-16 Lars Buesing , Nicolas Heess , Theophane Weber

Millions of people express themselves on public social media, such as Twitter. Through their posts, these people may reveal themselves as potentially valuable sources of information. For example, real-time information about an event might…

社会与信息网络 · 计算机科学 2014-04-09 Jalal Mahmud , Michelle Zhou , Nimrod Megiddo , Jeffrey Nichols , Clemens Drews

Diversified recommendation has attracted increasing attention from both researchers and practitioners, which can effectively address the homogeneity of recommended items. Existing approaches predominantly aim to infer the diversity of user…

信息检索 · 计算机科学 2026-01-07 Hanyang Yuan , Ning Tang , Tongya Zheng , Jiarong Xu , Xintong Hu , Renhong Huang , Shunyu Liu , Jiacong Hu , Jiawei Chen , Mingli Song

The recent proliferation of research into transformer based natural language processing has led to a number of studies which attempt to detect the presence of human-like cognitive behavior in the models. We contend that, as is true of human…

计算与语言 · 计算机科学 2024-04-01 Jesse Roberts , Kyle Moore , Drew Wilenzick , Doug Fisher

Algorithms deployed in education can shape the learning experience and success of a student. It is therefore important to understand whether and how such algorithms might create inequalities or amplify existing biases. In this paper, we…

计算机与社会 · 计算机科学 2022-12-21 Jade Maï Cock , Muhammad Bilal , Richard Davis , Mirko Marras , Tanja Käser

We analyze dynamic random network models where younger vertices connect to older ones with probabilities proportional to their degrees as well as a propensity kernel governed by their attribute types. Using stochastic approximation…

概率论 · 数学 2025-10-29 Nelson Antunes , Sayan Banerjee , Shankar Bhamidi , Vladas Pipiras

Social networks are typical attributed networks with node attributes. Different from traditional attribute community detection problem aiming at obtaining the whole set of communities in the network, we study an application-oriented problem…

社会与信息网络 · 计算机科学 2017-05-11 Peng Wu , Li Pan

Crowdsourcing employs human workers to solve computer-hard problems, such as data cleaning, entity resolution, and sentiment analysis. When crowdsourcing tabular data, e.g., the attribute values of an entity set, a worker's answers on the…

数据库 · 计算机科学 2017-08-08 Caihua Shan , Nikos Mamoulis , Guoliang Li , Reynold Cheng , Zhipeng Huang , Yudian Zheng

A network has a non-overlapping community structure if the nodes of the network can be partitioned into disjoint sets such that each node in a set is densely connected to other nodes inside the set and sparsely connected to the nodes out-…

社会与信息网络 · 计算机科学 2016-07-19 Talasila Sai Deepak , Hindol Adhya , Shyamal Kejriwal , Bhanuteja Gullapalli , Saswata Shannigrahi

Thompson sampling has emerged as an effective heuristic for a broad range of online decision problems. In its basic form, the algorithm requires computing and sampling from a posterior distribution over models, which is tractable only for…

机器学习 · 统计学 2023-04-26 Xiuyuan Lu , Benjamin Van Roy

In theory, a major advantage to the big data approach in studying online communities is that it should be possible to collect a representative random sample from a broadly defined population. However, in practice, data collection processes…

社会与信息网络 · 计算机科学 2021-02-02 Muhammad Umer Gurchani

A hidden database refers to a dataset that an organization makes accessible on the web by allowing users to issue queries through a search interface. In other words, data acquisition from such a source is not by following static…

数据库 · 计算机科学 2012-08-02 Cheng Sheng , Nan Zhang , Yufei Tao , Xin Jin

We propose the Heterogeneous Thurstone Model (HTM) for aggregating ranked data, which can take the accuracy levels of different users into account. By allowing different noise distributions, the proposed HTM model maintains the generality…

机器学习 · 计算机科学 2019-12-04 Tao Jin , Pan Xu , Quanquan Gu , Farzad Farnoud

The goal of our research is to contribute information about how useful the crowd is at anticipating stereotypes that may be biasing a data set without a researcher's knowledge. The results of the crowd's prediction can potentially be used…

人机交互 · 计算机科学 2018-01-11 Zeyuan Hu , Julia Strout

Finding influential users in online social networks is an important problem with many possible useful applications. HITS and other link analysis methods, in particular, have been often used to identify hub and authority users in web graphs…

社会与信息网络 · 计算机科学 2018-02-21 Roy Ka-Wei Lee , Tuan-Anh Hoang , Ee-Peng Lim

Even when aggregate accuracy is high, state-of-the-art NLP models often fail systematically on specific subgroups of data, resulting in unfair outcomes and eroding user trust. Additional data collection may not help in addressing these…

计算与语言 · 计算机科学 2023-05-30 Zexue He , Marco Tulio Ribeiro , Fereshte Khani

Social media data provides propitious opportunities for public health research. However, studies suggest that disparities may exist in the representation of certain populations (e.g., people of lower socioeconomic status). To quantify and…

计算机与社会 · 计算机科学 2017-11-07 Nina Cesare , Christan Grant , Jared B. Hawkins , John S. Brownstein , Elaine O. Nsoesie

Incorporating graph side information into recommender systems has been widely used to better predict ratings, but relatively few works have focused on theoretical guarantees. Ahn et al. (2018) firstly characterized the optimal sample…

信息论 · 计算机科学 2021-09-09 Changhun Jo , Kangwook Lee