中文
相关论文

相关论文: Towards Ground Truth Explainability on Tabular Dat…

200 篇论文

Causal machine learning has the potential to revolutionize decision-making by combining the predictive power of machine learning algorithms with the theory of causal inference. However, these methods remain underutilized by the broader…

In this paper, we observe that semi-structured tabulated text is ubiquitous; understanding them requires not only comprehending the meaning of text fragments, but also implicit relationships between them. We argue that such data can prove…

计算与语言 · 计算机科学 2020-05-14 Vivek Gupta , Maitrey Mehta , Pegah Nokhiz , Vivek Srikumar

Dimensionality reduction and clustering techniques are frequently used to analyze complex data sets, but their results are often not easy to interpret. We consider how to support users in interpreting apparent cluster structure on scatter…

机器学习 · 计算机科学 2021-11-08 Xander Vankwikelberge , Bo Kang , Edith Heiter , Jefrey Lijffijt

The emergence of generative AI models has dramatically expanded the availability and use of synthetic data across scientific, industrial, and policy domains. While these developments open new possibilities for data analysis, they also raise…

机器学习 · 统计学 2026-03-06 Ahmad Abdel-Azim , Ruoyu Wang , Xihong Lin

Recent advancements in tabular deep learning have demonstrated exceptional practical performance, yet the field often lacks a clear understanding of why these techniques actually succeed. To address this gap, our paper highlights the…

机器学习 · 计算机科学 2025-09-05 Nikolay Kartashev , Ivan Rubachev , Artem Babenko

Knowledge Graphs are pivotal for semantic data integration. The real-world data they model is often inherently uncertain. Within knowledge graphs, uncertainty manifests in three distinct levels: imprecise attribute values, probabilistic…

人工智能 · 计算机科学 2026-05-19 Jingcheng Wu

The assessment of process mining techniques using real-life data is often compromised by the lack of ground truth knowledge, the presence of non-essential outliers in system behavior and recording errors in event logs. Using synthetically…

数据库 · 计算机科学 2025-01-27 Dominique Sommers , Natalia Sidorova , Boudewijn van Dongen

Recent developments in machine learning have introduced models that approach human performance at the cost of increased architectural complexity. Efforts to make the rationales behind the models' predictions transparent have inspired an…

计算与语言 · 计算机科学 2020-09-29 Pepa Atanasova , Jakob Grue Simonsen , Christina Lioma , Isabelle Augenstein

Public image diffusion models are now powerful enough that an attacker without the resources to train a tabular-specific generator may repurpose one off the shelf. This study tests that possibility directly. An unmodified Stable Diffusion…

密码学与安全 · 计算机科学 2026-05-04 Adam Arthur , Christopher Schwartz

The past decade has seen a substantial rise in the amount of mis- and disinformation online, from targeted disinformation campaigns to influence politics, to the unintentional spreading of misinformation about public health. This…

计算与语言 · 计算机科学 2021-12-09 Isabelle Augenstein

The machine learning community has mainly relied on real data to benchmark algorithms as it provides compelling evidence of model applicability. Evaluation on synthetic datasets can be a powerful tool to provide a better understanding of a…

机器学习 · 计算机科学 2022-11-01 Florence Regol , Anja Kroon , Mark Coates

Synthetic data is becoming increasingly integral in data-scarce fields such as medical imaging, serving as a substitute for real data. However, its inherent statistical characteristics can significantly impact downstream tasks, potentially…

计算机视觉与模式识别 · 计算机科学 2024-12-23 Krishan Agyakari Raja Babu , Rachana Sathish , Mrunal Pattanaik , Rahul Venkataramani

Synthetic tabular data generation has emerged as a promising method to address limited data availability and privacy concerns. With the sharp increase in the performance of large language models in recent years, researchers have been…

机器学习 · 计算机科学 2025-03-28 Reilly Cannon , Nicolette M. Laird , Caesar Vazquez , Andy Lin , Amy Wagler , Tony Chiang

Agricultural research has been profited by technical advances such as automation, data mining. Today, data mining is used in a vast areas and many off-the-shelf data mining system products and domain specific data mining application soft…

人工智能 · 计算机科学 2012-06-08 Jay Gholap , Anurag Ingole , Jayesh Gohil , Shailesh Gargade , Vahida Attar

Tabular data provide answers to a significant portion of search queries. However, reciting an entire result table is impractical in conversational search systems. We propose to generate natural language summaries as answers to describe the…

信息检索 · 计算机科学 2020-07-13 Shuo Zhang , Zhuyun Dai , Krisztian Balog , Jamie Callan

Foundation models for tabular data are rapidly evolving, with increasing interest in extending them to support additional modalities such as free-text features. However, existing benchmarks for tabular data rarely include textual columns,…

机器学习 · 计算机科学 2025-07-11 Martin Mráz , Breenda Das , Anshul Gupta , Lennart Purucker , Frank Hutter

In the evolving domains of Machine Learning and Data Analytics, existing dataset characterization methods such as statistical, structural, and model-based analyses often fail to deliver the deep understanding and insights essential for…

机器学习 · 计算机科学 2025-10-17 Matthew D. Merris , Tim Andersen

Privacy, data quality, and data sharing concerns pose a key limitation for tabular data applications. While generating synthetic data resembling the original distribution addresses some of these issues, most applications would benefit from…

机器学习 · 计算机科学 2024-06-04 Mark Vero , Mislav Balunović , Martin Vechev

We introduce a method to learn a hierarchy of successively more abstract representations of complex data based on optimizing an information-theoretic objective. Intuitively, the optimization searches for a set of latent factors that best…

机器学习 · 计算机科学 2014-11-03 Greg Ver Steeg , Aram Galstyan

Tabular data is more challenging to generate than text and images, due to its heterogeneous features and much lower sample sizes. On this task, diffusion-based models are the current state-of-the-art (SotA) model class, achieving almost…