English
Related papers

Related papers: Building K-Anonymous User Cohorts with\\ Consecuti…

200 papers

Most epidemiologic cohorts are composed of volunteers who do not represent the general population. To enable population inference from cohorts, we and others have proposed utilizing probability survey samples as external references to…

Methodology · Statistics 2020-12-01 Lingxiao Wang , Barry I. Graubard , Hormuzd A. Katki , Yan Li

Large language models excel at many tasks but still struggle with consistent, robust reasoning. We introduce Cohort-based Consistency Learning (CC-Learn), a reinforcement learning framework that improves the reliability of LLM reasoning by…

Computation and Language · Computer Science 2025-06-19 Xiao Ye , Shaswat Shrivastava , Zhaonan Li , Jacob Dineen , Shijie Lu , Avneet Ahuja , Ming Shen , Zhikun Xu , Ben Zhou

Collecting the large datasets needed to train deep neural networks can be very difficult, particularly for the many applications for which sharing and pooling data is complicated by practical, ethical, or legal concerns. However, it may be…

Clustering is a separation of data into groups of similar objects. Every group called cluster consists of objects that are similar to one another and dissimilar to objects of other groups. In this paper, the K-Means algorithm is implemented…

Machine Learning · Computer Science 2013-04-03 P. Ashok , G. M Kadhar Nawaz , E. Elayaraja , V. Vadivel

Email service providers have employed many email classification and prioritization systems over the last decade to improve their services. In order to assist email services, we propose a personalized email community detection method to…

Social and Information Networks · Computer Science 2013-06-07 Waqas Nawaz , Yongkoo Han , Kifayat-Ullah Khan , Young-Koo Lee

Today, large amounts of valuable data are distributed among millions of user-held devices, such as personal computers, phones, or Internet-of-things devices. Many companies collect such data with the goal of using it for training machine…

Machine Learning · Computer Science 2020-08-21 Valentin Hartmann , Konark Modi , Josep M. Pujol , Robert West

Personalized health analytics increasingly rely on population benchmarks to provide contextual insights such as ''How do I compare to others like me?'' However, cohort-based aggregation of health data introduces nontrivial privacy risks,…

Cryptography and Security · Computer Science 2026-01-21 Richik Chakraborty , Lawrence Liu , Syed Hasnain

One way of investigating how genes affect human traits would be with a genome-wide association study (GWAS). Genetic markers, known as single-nucleotide polymorphism (SNP), are used in GWAS. This raises privacy and security concerns as…

Applications · Statistics 2019-08-02 Jun Jie Sim , Fook Mun Chan , Shibin Chen , Benjamin Hong Meng Tan , Khin Mi Mi Aung

In this paper, we investigate representation learning for low-resource keyword spotting (KWS). The main challenges of KWS are limited labeled data and limited available device resources. To address those challenges, we explore…

Sound · Computer Science 2023-03-21 Fan Cui , Liyong Guo , Quandong Wang , Peng Gao , Yujun Wang

Crowdsensing is an emerging paradigm of ubiquitous sensing, through which a crowd of workers are recruited to perform sensing tasks collaboratively. Although it has stimulated many applications, an open fundamental problem is how to select…

Computers and Society · Computer Science 2022-05-09 Feng Li , Jichao Zhao , Dongxiao Yu , Xiuzhen Cheng , Weifeng Lv

We describe CFW, a computationally efficient algorithm for collaborative filtering that uses posteriors over weights of evidence. In experiments on real data, we show that this method predicts as well or better than other methods in…

Information Retrieval · Computer Science 2015-05-19 Carl Kadie , Christopher Meek , David Heckerman

Motivated by undetectable risks in generative AI, we study a general robust aggregation problem: how to aggregate several probability distributions to boost safety. We present consensus sampling, a black-box algorithm that, given k…

Artificial Intelligence · Computer Science 2026-05-12 Adam Tauman Kalai , Yael Tauman Kalai , Or Zamir

A keyword spotting (KWS) system determines the existence of, usually predefined, keyword in a continuous speech stream. This paper presents a query-by-example on-device KWS system which is user-specific. The proposed system consists of two…

Machine Learning · Computer Science 2020-01-15 Byeonggeun Kim , Mingu Lee , Jinkyu Lee , Yeonseok Kim , Kyuwoong Hwang

Creating large, good quality labeled data has become one of the major bottlenecks for developing machine learning applications. Multiple techniques have been developed to either decrease the dependence of labeled data (zero/few-shot…

Computation and Language · Computer Science 2023-02-08 Abhinav Bohra , Huy Nguyen , Devashish Khatwani

The neighbourhood-based Collaborative Filtering is a widely used method in recommender systems. However, the risks of revealing customers' privacy during the process of filtering have attracted noticeable public concern recently.…

Cryptography and Security · Computer Science 2015-06-05 Zhigang Lu , Hong Shen

Unsourced random access is a novel communication paradigm designed for handling a large number of uncoordinated users that sporadically transmit very short messages. Under this model, coded compressed sensing (CCS) has emerged as a…

Information Theory · Computer Science 2021-07-22 Jamison R. Ebert , Vamsi K. Amalladinne , Stefano Rini , Jean-Francois Chamberland , Krishna R. Narayanan

In this paper, the problem of content-aware user clustering and content caching in wireless small cell networks is studied. In particular, a service delay minimization problem is formulated, aiming at optimally caching contents at the small…

Networking and Internet Architecture · Computer Science 2016-11-15 Mohammed S. ElBamby , Mehdi Bennis , Walid Saad , Matti Latva-aho

Minwise hashing (MinHash) is a standard algorithm widely used in the industry, for large-scale search and learning applications with the binary (0/1) Jaccard similarity. One common use of MinHash is for processing massive n-gram text…

Machine Learning · Statistics 2023-06-14 Xiaoyun Li , Ping Li

Genome-wide association studies are pivotal in understanding the genetic underpinnings of complex traits and diseases. Collaborative, multi-site GWAS aim to enhance statistical power but face obstacles due to the sensitive nature of genomic…

Cryptography and Security · Computer Science 2025-12-12 Arjhun Swaminathan , Anika Hannemann , Ali Burak Ünal , Nico Pfeifer , Mete Akgün

Recently, numerous community search methods for large graphs have been proposed, at the core of which is defining and measuring cohesion. This paper experimentally evaluates the effectiveness of these community search algorithms w.r.t.…

Information Retrieval · Computer Science 2025-05-02 Yining Zhao , Sourav S Bhowmick , Nastassja L. Fischer , SH Annabel Chen