English
Related papers

Related papers: Towards High-Performance Exploratory Data Analysis…

200 papers

Structural equation modeling (SEM) is a popular tool in the social and behavioural sciences, where it is being applied to ever more complex data types. The high-dimensional data produced by modern sensors, brain images, or (epi)genetic…

Methodology · Statistics 2019-10-11 Erik-Jan van Kesteren , Daniel L. Oberski

Scientific research increasingly relies on distributed computational resources, storage systems, networks, and instruments, ranging from HPC and cloud systems to edge devices. Event-driven architecture (EDA) benefits applications targeting…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-10-01 Haochen Pan , Ryan Chard , Sicheng Zhou , Alok Kamatar , Rafael Vescovi , Valérie Hayot-Sasson , André Bauer , Maxime Gonthier , Kyle Chard , Ian Foster

The problem of organizing data that evolves over time into clusters is encountered in a number of practical settings. We introduce evolutionary subspace clustering, a method whose objective is to cluster a collection of evolving data points…

Computer Vision and Pattern Recognition · Computer Science 2019-01-30 Abolfazl Hashemi , Haris Vikalo

Token sampling strategies critically influence text generation quality in large language models (LLMs). However, existing methods introduce additional hyperparameters, requiring extensive tuning and complicating deployment. We present…

Computation and Language · Computer Science 2025-12-02 Xiaodong Cai , Hai Lin , Shaoxiong Zhan , Weiqi Luo , Hong-Gee Kim , Hongyan Hao , Yu Yang , Hai-Tao Zheng

Spectral Embedding (SE) has often been used to map data points from non-linear manifolds to linear subspaces for the purpose of classification and clustering. Despite significant advantages, the subspace structure of data in the original…

Computer Vision and Pattern Recognition · Computer Science 2023-05-16 Hira Yaseen , Arif Mahmood

Discovering meaningful insights from a large dataset, known as Exploratory Data Analysis (EDA), is a challenging task that requires thorough exploration and analysis of the data. Automated Data Exploration (ADE) systems use goal-oriented…

Artificial Intelligence · Computer Science 2024-10-22 Abhijit Manatkar , Ashlesha Akella , Parthivi Gupta , Krishnasuri Narayanam

We study sparse principal component analysis in the high-dimensional, sample-limited regime, aiming to recover a leading component supported on a few coordinates. Despite extensive progress, most methods and analyses are tailored to the…

Information Theory · Computer Science 2025-12-18 Mengchu Xu , Jian Wang , Yonina C. Eldar

This paper proposes AEDA (An Easier Data Augmentation) technique to help improve the performance on text classification tasks. AEDA includes only random insertion of punctuation marks into the original text. This is an easier technique to…

Computation and Language · Computer Science 2021-08-31 Akbar Karimi , Leonardo Rossi , Andrea Prati

As the size of circuit designs continues to grow rapidly, artificial intelligence technologies are being extensively used in Electronic Design Automation (EDA) to assist with circuit design. Placement and routing are the most time-consuming…

Computer Vision and Pattern Recognition · Computer Science 2023-10-27 Jialv Zou , Xinggang Wang , Jiahao Guo , Wenyu Liu , Qian Zhang , Chang Huang

We construct a theory to introduce the concept of topologically robust exceptional points (EP). Starting from an ordered system with $N$ elements, we find the necessary condition to have the highest order exceptional point, namely…

Optics · Physics 2018-12-07 Cem Yuce , Hamidreza Ramezani

Even with the rise in popularity of over-parameterized models, simple dimensionality reduction and clustering methods, such as PCA and k-means, are still routinely used in an amazing variety of settings. A primary reason is the combination…

Methodology · Statistics 2020-09-08 Debolina Paul , Saptarshi Chakraborty , Didong Li , David Dunson

Business Intelligence (BI) analysis is evolving towards Exploratory BI, an iterative, multi-round exploration paradigm where analysts progressively refine their understanding. However, traditional BI systems impose critical limits for…

Databases · Computer Science 2026-03-27 Yunkai Lou , Shunyang Li , Longbin Lai , Jianke Yu , Wenyuan Yu , Ying Zhang

Clustering techniques are very attractive for extracting and identifying patterns in datasets. However, their application to very large spatial datasets presents numerous challenges such as high-dimensionality data, heterogeneity, and high…

Databases · Computer Science 2018-02-27 Malika Bendechache , Nhien-An Le-Khac , M-Tahar Kechadi

Data selection is designed to accelerate learning with preserved performance. To achieve this, a fundamental thought is to identify informative data samples with significant contributions to the training. In this work, we propose…

Machine Learning · Computer Science 2025-09-30 Ziheng Cheng , Zhong Li , Jiang Bian

Ensemble Kalman Sampler (EKS) is a method to find approximately $i.i.d.$ samples from a target distribution. As of today, why the algorithm works and how it converges is mostly unknown. The continuous version of the algorithm is a set of…

Numerical Analysis · Mathematics 2025-03-07 Zhiyan Ding , Qin Li

Organizations struggle to share data across departments that have adopted different data analytics platforms. If n datasets must serve m environments, up to n*m replicas can emerge, increasing inconsistency and cost. Traditional warehouses…

Databases · Computer Science 2025-12-04 Ryoto Miyamoto , Akira Kasuga

Expert estimation of objects takes place when there are no benchmark values of object weights, but these weights still have to be defined. That is why it is problematic to define the efficiency of expert estimation methods. We propose to…

Artificial Intelligence · Computer Science 2019-11-13 Sergii Kadenko , Vitaliy Tsyganok

When the data are stored in a distributed manner, direct application of traditional statistical inference procedures is often prohibitive due to communication cost and privacy concerns. This paper develops and investigates two…

Machine Learning · Statistics 2021-08-04 Jianqing Fan , Yongyi Guo , Kaizheng Wang

In graph-based active learning, algorithms based on expected error minimization (EEM) have been popular and yield good empirical performance. The exact computation of EEM optimally balances exploration and exploitation. In practice,…

Machine Learning · Statistics 2016-09-06 Kwang-Sung Jun , Robert Nowak

It is becoming common to archive research datasets that are not only large but also numerous. In addition, their corresponding metadata and the software required to analyse or display them need to be archived. Yet the manual curation of…

Digital Libraries · Computer Science 2011-08-24 Daniel Lemire , Andre Vellino