中文
相关论文

相关论文: Using saturated count models for user-friendly syn…

200 篇论文

We propose a method to generate statistically representative synthetic data from a given dataset. The main goal of our method is for the created data set to mimic the inter--feature correlations present in the original data, while also…

机器学习 · 计算机科学 2025-06-25 Nicklas Jävergård , Rainey Lyons , Adrian Muntean , Jonas Forsman

We develop a model-based methodology for integrating gene-set information with an experimentally-derived gene list. The methodology uses a previously reported sampling model, but takes advantage of natural constraints in the…

统计方法学 · 统计学 2015-06-02 Zhishi Wang , Qiuling He , Bret Larget , Michael A. Newton

The appropriateness of the Poisson model is frequently challenged when examining spatial count data marked by unbalanced distributions, over-dispersion, or under-dispersion. Moreover, traditional parametric models may inadequately capture…

统计方法学 · 统计学 2025-03-26 Mahsa Nadifar , Andriette Bekker , Mohammad Arashi , Abel Ramoelo

Synthetic data has garnered attention as a Privacy Enhancing Technology (PET) in sectors such as healthcare and finance. When using synthetic data in practical applications, it is important to provide protection guarantees. In the…

Synthetic data generation, leveraging generative machine learning techniques, offers a promising approach to mitigating privacy concerns associated with real-world data usage. Synthetic data closely resembles real-world data while…

机器学习 · 计算机科学 2025-08-25 Weijie Niu , Alberto Huertas Celdran , Karoline Siarsky , Burkhard Stiller

Hierarchically-organized data arise naturally in many psychology and neuroscience studies. As the standard assumption of independent and identically distributed samples does not hold for such data, two important problems are to accurately…

统计理论 · 数学 2018-09-03 Irene Dowding , Stefan Haufe

Time series of counts arise in a variety of forecasting applications, for which traditional models are generally inappropriate. This paper introduces a hierarchical Bayesian formulation applicable to count time series that can easily…

机器学习 · 统计学 2014-05-16 Nicolas Chapados

The use of synthetic data to deidentify data and to improve predictive models is well-attested to. The augmentation of datasets using synthetically generated data is an alluring proposition: in the best case, it generates realistic data…

统计方法学 · 统计学 2026-03-20 Reid Dale , Jordan Rodu , Mike Baiocchi

Unit-level models for survey data offer many advantages over their area-level counterparts, such as potential for more precise estimates and a natural benchmarking property. However two main challenges occur in this context: accounting for…

统计方法学 · 统计学 2020-05-18 Paul A. Parker , Scott H. Holan , Ryan Janicki

Benchmarking is crucial for evaluating a DBMS, yet existing benchmarks often fail to reflect the varied nature of user workloads. As a result, there is increasing momentum toward creating databases that incorporate real-world user data to…

数据库 · 计算机科学 2025-04-11 Yunqing Ge , Jianbin Qin , Shuyuan Zheng , Yongrui Zhong , Bo Tang , Yu-Xuan Qiu , Rui Mao , Ye Yuan , Makoto Onizuka , Chuan Xiao

Formulating efficient SQL queries requires several cycles of tuning and execution, particularly for inexperienced users. We examine methods that can accelerate and improve this interaction by providing insights about SQL queries prior to…

数据库 · 计算机科学 2020-02-24 Zainab Zolaktaf , Mostafa Milani , Rachel Pottinger

Data privacy is a core tenet of responsible computing, and in the United States, differential privacy (DP) is the dominant technical operationalization of privacy-preserving data analysis. With this study, we qualitatively examine one class…

人机交互 · 计算机科学 2024-12-18 Lucas Rosenblatt , Bill Howe , Julia Stoyanovich

This paper introduces SynDiffix, a mechanism for generating statistically accurate, anonymous synthetic data for structured data. Recent open source and commercial systems use Generative Adversarial Networks or Transformed Auto Encoders to…

密码学与安全 · 计算机科学 2023-11-17 Paul Francis , Cristian Berneanu , Edon Gashi

Most statistical agencies release randomly selected samples of Census microdata, usually with sample fractions under 10% and with other forms of statistical disclosure control (SDC) applied. An alternative to SDC is data synthesis, which…

密码学与安全 · 计算机科学 2022-07-08 Claire Little , Mark Elliot , Richard Allmendinger

Due to their data-driven nature, Machine Learning (ML) models are susceptible to bias inherited from data, especially in classification problems where class and group imbalances are prevalent. Class imbalance (in the classification target)…

机器学习 · 计算机科学 2024-09-10 Emmanouil Panagiotou , Arjun Roy , Eirini Ntoutsi

Imbalanced classification and spurious correlation are common challenges in data science and machine learning. Both issues are linked to data imbalance, with certain groups of data samples significantly underrepresented, which in turn would…

机器学习 · 统计学 2026-02-10 Ryumei Nakada , Yichen Xu , Lexin Li , Linjun Zhang

Privacy is an important concern for our society where sharing data with partners or releasing data to the public is a frequent occurrence. Some of the techniques that are being used to achieve privacy are to remove identifiers, alter…

数据库 · 计算机科学 2018-07-04 Noseong Park , Mahmoud Mohammadi , Kshitij Gorde , Sushil Jajodia , Hongkyu Park , Youngmin Kim

A growing number of approaches exist to generate explanations for image classification. However, few of these approaches are subjected to human-subject evaluations, partly because it is challenging to design controlled experiments with…

人工智能 · 计算机科学 2021-05-07 Martin Schuessler , Philipp Weiß , Leon Sixt

Datasets in engineering domains are often small, sparsely labeled, and contain numerical as well as categorical conditions. Additionally. computational resources are typically limited in practical applications which hinders the adoption of…

机器学习 · 计算机科学 2025-05-23 Phillip Mueller , Jannik Wiese , Sebastian Mueller , Lars Mikelsons

This study addresses the reliability of automatic summarization in high-risk scenarios and proposes a large language model framework that integrates uncertainty quantification and risk-aware mechanisms. Starting from the demands of…

计算与语言 · 计算机科学 2025-10-03 Shuaidong Pan , Di Wu
‹ 上一页 1 8 9 10 下一页 ›