English
Related papers

Related papers: Individual Data Protected Integrative Regression A…

200 papers

Data security and availability for operational use are frequently seen as conflicting goals. Research on searchable encryption and homomorphic encryption are a start, but they typically build from encryption methods that, at best, provide…

Cryptography and Security · Computer Science 2015-12-02 David Zage , Helen Xu , Thomas Kroeger , Bridger Hahn , Nolan Donoghue , Thomas Benson

In many areas, practitioners need to analyze large datasets that challenge conventional single-machine computing. To scale up data analysis, distributed and parallel computing approaches are increasingly needed. Here we study a fundamental…

Statistics Theory · Mathematics 2020-06-04 Edgar Dobriban , Yue Sheng

Regression discontinuity design (RDD) is widely adopted for causal inference under intervention determined by a continuous variable. While one is interested in treatment effect heterogeneity by subgroups in many applications, RDD typically…

Methodology · Statistics 2024-11-11 Shonosuke Sugasawa , Takuya Ishihara , Daisuke Kurisu

Past few years have witnessed a growing recognition of intelligent techniques for the construction of efficient and reliable intrusion detection systems. Due to increasing incidents of cyber attacks, building effective intrusion detection…

Artificial Intelligence · Computer Science 2007-05-23 Srinivas Mukkamala , Andrew H. Sung , Ajith Abraham , Vitorino Ramos

A private information retrieval (PIR) scheme is a protocol that allows a user to retrieve a file from a database without revealing the identity of the desired file to a curious database. Given a distributed data storage system, efficient…

Information Retrieval · Computer Science 2025-08-08 Camilla Hollanti , Neehar Verma

The dramatic growth of big datasets presents a new challenge to data storage and analysis. Data reduction, or subsampling, that extracts useful information from datasets is a crucial step in big data analysis. We propose an orthogonal…

Methodology · Statistics 2021-06-01 Lin Wang , Jake Elmstedt , Weng Kee Wong , Hongquan Xu

Covariance regression offers an effective way to model the large covariance matrix with the auxiliary similarity matrices. In this work, we propose a sparse covariance regression (SCR) approach to handle the potentially high-dimensional…

Methodology · Statistics 2024-10-17 Yuan Gao , Zhiyuan Zhang , Zhanrui Cai , Xuening Zhu , Tao Zou , Hansheng Wang

Outlier detection is critical in real applications to prevent financial fraud, defend network intrusions, or detecting imminent device failures. To reduce the human effort in evaluating outlier detection results and effectively turn the…

Machine Learning · Computer Science 2023-09-04 Yu Wang , Lei Cao , Yizhou Yan , Samuel Madden

Scaling Federated Learning (FL) to billion-parameter models forces a challenging trade-off between privacy, scalability, and model utility. Existing solutions often tackle these challenges in isolation, sacrificing accuracy, relying on…

Machine Learning · Computer Science 2026-05-12 Dario Fenoglio , Pasquale Polverino , Jacopo Quizi , Martin Gjoreski , Akash Dhasade , Marc Langheinrich

Insurance companies often operate across multiple interrelated lines of business (LOBs), and accounting for dependencies between them is essential for accurate reserve estimation and risk capital determination. In our previous work on the…

Methodology · Statistics 2025-09-09 Pengfei Cai , Anas Abdallah , Pratheepa Jeganathan

Understanding network influence and its determinants are key challenges in political science and network analysis. Traditional latent variable models position actors within a social space based on network dependencies but often do not…

Applications · Statistics 2025-08-28 Shahryar Minhas , Peter D. Hoff

Composed Image Retrieval (CIR) is an emerging yet challenging task that allows users to search for target images using a multimodal query, comprising a reference image and a modification text specifying the user's desired changes to the…

Multimedia · Computer Science 2025-03-05 Xuemeng Song , Haoqiang Lin , Haokun Wen , Bohan Hou , Mingzhu Xu , Liqiang Nie

We consider the task of meta-analysis in high-dimensional settings in which the data sources are similar but non-identical. To borrow strength across such heterogeneous datasets, we introduce a global parameter that emphasizes…

Methodology · Statistics 2022-07-01 Subha Maity , Yuekai Sun , Moulinath Banerjee

It is an important task to model realized volatilities for high-frequency data in finance and economics and, as arguably the most popular model, the heterogeneous autoregressive (HAR) model has dominated the applications in this area.…

Methodology · Statistics 2023-03-07 Huiling Yuan , Kexin Lu , Yifeng Guo , Guodong Li

We study an individual-based stochastic SIR epidemic model with infection-age dependent infectivity on a large random graph, capturing individual heterogeneity and non-homogeneous connectivity. Each individual is associated with particular…

Probability · Mathematics 2026-02-11 Guodong Pang , Étienne Pardoux , Aurélien Velleret

Ultrahigh dimensional data sets are becoming increasingly prevalent in areas such as bioinformatics, medical imaging, and social network analysis. Sure independent screening of such data is commonly used to analyze such data. Nevertheless,…

Methodology · Statistics 2020-10-15 Randall Reese , Xiaotian Dai , Guifang Fu

This paper describes the hierarchical infinite relational model (HIRM), a new probabilistic generative model for noisy, sparse, and heterogeneous relational data. Given a set of relations defined over a collection of domains, the model…

Machine Learning · Computer Science 2022-02-25 Feras A. Saad , Vikash K. Mansinghka

Extraordinary amounts of data are being produced in many branches of science. Proven statistical methods are no longer applicable with extraordinary large data sets due to computational limitations. A critical step in big data analysis is…

Methodology · Statistics 2019-06-27 HaiYing Wang , Min Yang , John Stufken

We consider information-theoretic privacy in federated submodel learning, where a global server has multiple submodels. Compared to the privacy considered in the conventional federated submodel learning where secure aggregation is adopted…

Information Theory · Computer Science 2020-08-19 Minchul Kim , Jungwoo Lee

We systematically investigate the preservation of differential privacy in functional data analysis, beginning with functional mean estimation and extending to varying coefficient model estimation. Our work introduces a distributed learning…

Statistics Theory · Mathematics 2026-02-11 Gengyu Xue , Zhenhua Lin , Yi Yu