中文
相关论文

相关论文: Statistical Distortion: Consequences of Data Clean…

200 篇论文

We propose a practical methodology to protect a user's private data, when he wishes to publicly release data that is correlated with his private data, in the hope of getting some utility. Our approach relies on a general statistical…

Recent advancements in Large Language Models (LLMs) have demonstrated significant progress in various areas, such as text generation and code synthesis. However, the reliability of performance evaluation has come under scrutiny due to data…

计算与语言 · 计算机科学 2025-06-06 Yuxing Cheng , Yi Chang , Yuan Wu

Data science has become increasingly essential for the production of official statistics, as it enables the automated collection, processing, and analysis of large amounts of data. With such data science practices in place, it enables more…

机器学习 · 统计学 2023-06-08 Cedric De Boom , Michael Reusens

One of the main cost factors in software development is the detection and removal of defects. However, the relationships and influencing factors of the costs and revenues of defect-detection techniques are still not well understood. This…

软件工程 · 计算机科学 2016-12-13 Stefan Wagner

Substances such as chemical compounds are invisible to human eyes, they are usually captured by sensing equipments with their spectral fingerprints. Though spectra of pure chemicals can be identified by visual inspection, the spectra of…

数值分析 · 数学 2015-01-07 Yuanchang Sun , Wensong Wu , Jack Xin

We introduce statistical constraints, a declarative modelling tool that links statistics and constraint programming. We discuss two statistical constraints and some associated filtering algorithms. Finally, we illustrate applications to…

人工智能 · 计算机科学 2014-09-09 Roberto Rossi , Steven Prestwich , S. Armagan Tarim

The problem of private information "leakage" (inadvertently or by malicious design) from the myriad large centralized searchable data repositories drives the need for an analytical framework that quantifies unequivocally how safe private…

信息论 · 计算机科学 2010-02-09 Lalitha Sankar , S. Raj Rajagopalan , H. Vincent Poor

With the rapid development of large language models (LLMs), the quality of training data has become crucial. Among the various types of training data, mathematical data plays a key role in enabling LLMs to acquire strong reasoning…

计算与语言 · 计算机科学 2025-02-27 Hao Liang , Meiyi Qiang , Yuying Li , Zefeng He , Yongzhen Guo , Zhengzhou Zhu , Wentao Zhang , Bin Cui

Spectral methods have emerged as a simple yet surprisingly effective approach for extracting information from massive, noisy and incomplete data. In a nutshell, spectral methods refer to a collection of algorithms built upon the eigenvalues…

机器学习 · 统计学 2021-10-26 Yuxin Chen , Yuejie Chi , Jianqing Fan , Cong Ma

Statistical significance measures the reliability of a result obtained from a random experiment. We investigate the number of repetitions needed for a statistical result to have a certain significance. In the first step, we consider…

统计方法学 · 统计学 2024-06-19 Maike Tormählen , Galiya Klinkova , Michael Grabinski

Computing the rate-distortion function for continuous sources is commonly regarded as a standard continuous optimization problem. When numerically addressing this problem, a typical approach involves discretizing the source space and…

信息论 · 计算机科学 2024-05-02 Lingyi Chen , Shitong Wu , Wenyi Zhang , Huihui Wu , Hao Wu

In this paper, by proposing two new kinds of distributional uncertainty sets, we explore robustness of distortion risk measures against distributional uncertainty. To be precise, we first consider a distributional uncertainty set which is…

风险管理 · 定量金融 2025-08-15 Xiangyu Han , Yijun Hu , Ran Wang , Linxiao Wei

The popularity of deep learning has led to the curation of a vast number of massive and multifarious datasets. Despite having close-to-human performance on individual tasks, training parameter-hungry models on large datasets poses…

机器学习 · 计算机科学 2023-09-27 Noveen Sachdeva , Julian McAuley

Machine learning (ML) has employed various discretization methods to partition numerical attributes into intervals. However, an effective discretization technique remains elusive in many ML applications, such as association rule mining.…

机器学习 · 计算机科学 2023-11-07 Minakshi Kaushik , Rahul Sharma , Dirk Draheim

Measuring inter-dataset similarity is an important task in machine learning and data mining with various use cases and applications. Existing methods for measuring inter-dataset similarity are computationally expensive, limited, or…

机器学习 · 计算机科学 2025-05-06 Muhammad Rajabinasab , Anton D. Lautrup , Arthur Zimek

In any other circumstance, it might make sense to define the extent of the terrain (Data Science) first, and then locate and describe the landmarks (Principles). But this data revolution we are experiencing defies a cadastral survey. Areas…

其他统计学 · 统计学 2021-02-04 Noel Cressie

Statistically sound pattern discovery harnesses the rigour of statistical hypothesis testing to overcome many of the issues that have hampered standard data mining approaches to pattern discovery. Most importantly, application of…

统计方法学 · 统计学 2019-01-07 Wilhelmiina Hämäläinen , Geoffrey I. Webb

Curriculum learning is a training strategy that sorts the training examples by some measure of their difficulty and gradually exposes them to the learner to improve the network performance. Motivated by our insights from implicit curriculum…

机器学习 · 计算机科学 2021-07-28 Vinu Sankar Sadasivan , Anirban Dasgupta

In data mining applications, feature selection is an essential process since it reduces a model's complexity. The cost of obtaining the feature values must be taken into consideration in many domains. In this paper, we study the…

机器学习 · 计算机科学 2013-06-04 Hong Zhao , Fan Min , William Zhu

We propose a novel measure of statistical depth, the metric spatial depth, for data residing in an arbitrary metric space. The measure assigns high (low) values for points located near (far away from) the bulk of the data distribution,…

统计理论 · 数学 2023-06-19 Joni Virta