中文
相关论文

相关论文: Toward Compact Data from Big Data

200 篇论文

Dataset distillation is the task of synthesizing a small dataset such that a model trained on the synthetic set will match the test accuracy of the model trained on the full dataset. In this paper, we propose a new formulation that…

计算机视觉与模式识别 · 计算机科学 2022-03-23 George Cazenavette , Tongzhou Wang , Antonio Torralba , Alexei A. Efros , Jun-Yan Zhu

This is a thought piece on data-intensive science requirements for databases and science centers. It argues that peta-scale datasets will be housed by science centers that provide substantial storage and processing for scientists who access…

数据库 · 计算机科学 2007-05-23 Jim Gray , David T. Liu , Maria Nieto-Santisteban , Alexander S. Szalay , David DeWitt , Gerd Heber

One of the increasingly important technologies dealing with the growing complexity of the digitalization of almost all human activities is Artificial intelligence, more precisely machine learning Despite the fact, that we live in a Big data…

机器学习 · 计算机科学 2021-03-02 Peter Kokol , Marko Kokol , Sašo Zagoranski

Big data analytics is one of the most promising areas of new research and development in computer science, enterprises, e-commerce, and defense. For many organizations, big data is considered one of their most important strategic assets.…

信息检索 · 计算机科学 2025-07-16 Santanu Acharjee , Ripunjoy Choudhury

With ubiquitous sensors continuously monitoring and collecting large amounts of information, there is no doubt that this is an era of big data. One of the important sources for scientific big data is the datasets collected by Internet of…

其他计算机科学 · 计算机科学 2017-05-04 Yongshuai Shao , Zhe Chen

In Big data era, information integration often requires abundant data extracted from massive data sources. Due to a large number of data sources, data source selection plays a crucial role in information integration, since it is costly and…

数据库 · 计算机科学 2016-11-01 Yiming Lin , Hongzhi Wang , Jianzhong Li , Hong Gao

Data quality describes the degree to which data meet specific requirements and are fit for use by humans and/or downstream tasks (e.g., artificial intelligence). Data quality can be assessed across multiple high-level concepts called…

数据库 · 计算机科学 2025-07-24 Vasileios Papastergios , Lisa Ehrlinger , Anastasios Gounaris

Big data have the characteristics of enormous volume, high velocity, diversity, value-sparsity, and uncertainty, which lead the knowledge learning from them full of challenges. With the emergence of crowdsourcing, versatile information can…

机器学习 · 计算机科学 2022-06-22 Jing Zhang

A set of preferred records can be obtained from a large database in a multi-criteria setting using various computational methods which either depend on the concept of dominance or on the concept of utility or scoring function based on the…

数据库 · 计算机科学 2022-03-18 Anagha Radhakrishnan

Subsampling from a large data set is useful in many supervised learning contexts to provide a global view of the data based on only a fraction of the observations. Diverse (or space-filling) subsampling is an appealing subsampling approach…

统计方法学 · 统计学 2023-11-27 Boyang Shang , Daniel W. Apley , Sanjay Mehrotra

This book chapter attempts to counter anxieties in the humanities and social science about the role of big data in research by focusing on approaches which, by being firmly grounded in the traditional values of disciplines, enhance existing…

计算机与社会 · 计算机科学 2016-05-23 Tobias Blanke , Andrew Prescott

Dataset distillation aims at synthesizing a dataset by a small number of artificially generated data items, which, when used as training data, reproduce or approximate a machine learning (ML) model as if it were trained on the entire…

机器学习 · 计算机科学 2024-03-27 Radu-Andrei Rosu , Mihaela-Elena Breaban , Henri Luchian

The Big Data management is a problem right now. The Big Data growth is very high. It is very difficult to manage due to various characteristics. This manuscript focuses on Big Data analytics in cloud environment using Hadoop. We have…

分布式、并行与集群计算 · 计算机科学 2016-10-17 Mansaf Alam , Kashish Ara Shakil

'Big' high-dimensional data are commonly analyzed in low-dimensions, after performing a dimensionality-reduction step that inherently distorts the data structure. For the same purpose, clustering methods are also often used. These methods…

机器学习 · 统计学 2019-02-20 Tom Lorimer , Karlis Kanders , Ruedi Stoop

Over the past decade, the data lake concept has emerged as an alternative to data warehouses for storing and analyzing big data. A data lake allows storing data without any predefined schema. Therefore, data querying and analysis depend on…

Unstructured data, such as text, images, audio, and video, comprises the vast majority of the world's information, yet it remains poorly supported by traditional data systems that rely on structured formats for computation. We argue for a…

数据库 · 计算机科学 2025-09-19 Mushtari Sadia , Amrita Roy Chowdhury , Ang Chen

Advances in science are being sought in newly available opportunities to collect massive quantities of data about complex systems. While key advances are being made in detailed mapping of systems, how to relate this data to solving many of…

物理与社会 · 物理学 2016-04-05 Yaneer Bar-Yam

The quality of the data in a dataset can have a substantial impact on the performance of a machine learning model that is trained and/or evaluated using the dataset. Effective dataset management, including tasks such as data cleanup,…

数据库 · 计算机科学 2023-03-16 Ze Mao , Yang Xu , Erick Suarez

This explainer document aims to provide an overview of the current state of the rapidly expanding work on synthetic data technologies, with a particular focus on privacy. The article is intended for a non-technical audience, though some…

Conformal Prediction is a machine learning methodology that produces valid prediction regions under mild conditions. In this paper, we explore the application of making predictions over multiple data sources of different sizes without…

机器学习 · 统计学 2018-06-15 Ola Spjuth , Lars Carlsson , Niharika Gauraha