English
Related papers

Related papers: Mining Hidden Populations through Attributed Searc…

200 papers

Recent advances in natural language processing (NLP) in online social media are evidently owed to large-scale datasets. However, labeling, storing, and processing a large number of textual data points, e.g., tweets, has remained…

Computation and Language · Computer Science 2022-02-02 Toktam A. Oghaz , Ivan Garibay

Populations of interest are often hidden from data for a variety of reasons, though their magnitude remains important in determining resource allocation and appropriate policy. One popular approach to population size estimation, the…

Methodology · Statistics 2025-06-27 Mallory J Flynn , Paul Gustafson

Anomaly detection research works generally propose algorithms or end-to-end systems that are designed to automatically discover outliers in a dataset or a stream. While literature abounds concerning algorithms or the definition of metrics…

Networking and Internet Architecture · Computer Science 2022-11-21 Jose Manuel Navarro , Alexis Huet , Dario Rossi

Increasing evidence suggests that a growing amount of social media content is generated by autonomous entities known as social bots. In this work we present a framework to detect such entities on Twitter. We leverage more than a thousand…

Social and Information Networks · Computer Science 2017-03-28 Onur Varol , Emilio Ferrara , Clayton A. Davis , Filippo Menczer , Alessandro Flammini

Given a set of attributed subgraphs known to be from different classes, how can we discover their differences? There are many cases where collections of subgraphs may be contrasted against each other. For example, they may be assigned…

Social and Information Networks · Computer Science 2017-02-01 Aria Rezaei , Bryan Perozzi , Leman Akoglu

Social networks include millions of users constantly looking for new relationships for personal or professional purposes. Social network sites recommend friends based on relationship features and content information. A significant part of…

Social and Information Networks · Computer Science 2020-03-26 Ali Choumane , Zein Al Abidin Ibrahim

As known, attribute selection is a method that is used before the classification of data mining. In this study, a new data set has been created by using attributes expressing overall satisfaction in Turkey Statistical Institute (TSI) Life…

Machine Learning · Computer Science 2018-07-20 Adil Çoban , Ilhan Tarımer

Communities are an important feature of social networks. The goal of this paper is to propose a mathematical model to study the community structure in social networks. For this, we consider a particular case of a social network, namely…

Social and Information Networks · Computer Science 2020-04-14 Peter Marbach

Respondent-driven sampling (RDS) is a popular method for sampling hard-to-survey populations that leverages social network connections through peer recruitment. While RDS is most frequently applied to estimate the prevalence of infections…

Methodology · Statistics 2016-10-24 Ashton M. Verdery , Jacob C. Fisher , Nalyn Siripong , Kahina Abdesselam , Shawn Bauldry

Probabilistic models learned as density estimators can be exploited in representation learning beside being toolboxes used to answer inference queries only. However, how to extract useful representations highly depends on the particular…

Machine Learning · Computer Science 2016-08-12 Antonio Vergari , Nicola Di Mauro , Floriana Esposito

The wide use of social media sites and other digital technologies have resulted in an unprecedented availability of digital data that are being used to study human behavior across research domains. Although unsolicited opinions and…

Social and Information Networks · Computer Science 2018-06-01 Nina Cesare , Christan Grant , Quynh Nguyen , Hedwig Lee , Elaine O. Nsoesie

Unveiling individuals' preferences for connecting with similar others (choice homophily) beyond the structural factors determining the pool of opportunities, is a challenging task. Here, we introduce a robust methodology for quantifying and…

Physics and Society · Physics 2024-01-25 Sina Sajjadi , Samuel Martin-Gutierrez , Fariba Karimi

In this work, we study practical heuristics to improve the performance of prefix-tree based algorithms for differentially private heavy hitter detection. Our model assumes each user has multiple data points and the goal is to learn as many…

Machine Learning · Computer Science 2023-07-24 Karan Chadha , Junye Chen , John Duchi , Vitaly Feldman , Hanieh Hashemi , Omid Javidbakht , Audra McMillan , Kunal Talwar

Disagreement remains on what the target estimand should be for population-adjusted indirect treatment comparisons. This debate is of central importance for policy-makers and applied practitioners in health technology assessment.…

Methodology · Statistics 2022-12-06 Antonio Remiro-Azócar

Social relationships can be divided into different classes based on the regularity with which they occur and the similarity among them. Thus, rare and somewhat similar relationships are random and cause noise in a social network, thus…

Social and Information Networks · Computer Science 2018-10-08 Jeancarlo Campos Leão , Michele Amaral Brandão , Pedro O. S. Vaz de Melo , Alberto H. F. Laender

Both named entities and keywords are important in defining the content of a text in which they occur. In particular, people often use named entities in information search. However, named entities have ontological features, namely, their…

Information Retrieval · Computer Science 2018-07-17 Tru H. Cao , Vuong M. Ngo

In uses of pre-trained machine learning models, it is a known issue that the target population in which the model is being deployed may not have been reflected in the source population with which the model was trained. This can result in a…

Machine Learning · Computer Science 2023-06-27 Jose M. Alvarez , Kristen M. Scott , Salvatore Ruggieri , Bettina Berendt

Today's probabilistic language generators fall short when it comes to producing coherent and fluent text despite the fact that the underlying models perform well under standard metrics, e.g., perplexity. This discrepancy has puzzled the…

Computation and Language · Computer Science 2025-06-06 Clara Meister , Tiago Pimentel , Gian Wiher , Ryan Cotterell

This paper presents a new approach for unsupervised Spoken Term Detection with spoken queries using multiple sets of acoustic patterns automatically discovered from the target corpus. The different pattern HMM configurations(number of…

Computation and Language · Computer Science 2015-09-09 Cheng-Tao Chung , Chun-an Chan , Lin-shan Lee

We present a new design and inference method for estimating population size of a hidden population best reached through a link-tracing design. The strategy involves the Rao-Blackwell Theorem applied to a sufficient statistic markedly…

Methodology · Statistics 2014-11-26 Kyle Vincent , Steve Thompson