English
Related papers

Related papers: Data Quality Taxonomy for Data Monetization

200 papers

Analyzing Big Data can help corporations to im-prove their efficiency. In this work we present a new vision to derive Value from Big Data using a Semantic Hierarchical Multi-label Classification called Semantic HMC based in a non-supervised…

Artificial Intelligence · Computer Science 2014-12-03 Thomas Hassan , Rafael Peixoto , Christophe Cruz , Aurlie Bertaux , Nuno Silva

The intersection of blockchain (distributed ledger) and identity management lacks a comprehensive framework for classifying distributed-ledger-based identity solutions. This paper introduces a methodologically developed taxonomy derived…

Cryptography and Security · Computer Science 2025-07-29 Awid Vaziry , Sandro Rodriguez Garzon , Patrick Herbke , Carlo Segat , Axel Kupper

The Internet of Things (IoT) is a paradigm that connects everyday items to the Internet. In the recent decade, the IoT's spreading popularity is a promising opportunity for people and industries. IoT utilizes in a wide range of respects…

Networking and Internet Architecture · Computer Science 2021-03-25 Taha Mansouri , Mohammad Reza Sadeghi Moghadam , Fatemeh Monshizadeh , Ahad Zareravasan

Spatial networks are a powerful framework for studying a large variety of systems belonging to a broad diversity of contexts: from transportation to biology, from epidemiology to communications, and migrations, to cite a few. Spatial…

Physics and Society · Physics 2020-04-08 Ignacio Morer , Alessio Cardillo , Albert Diaz-Guilera , Luce Prignano , Sergi Lozano

For fair academic recruitment at universities and research institutions, determination of the right measure based on globally accepted academic quality features is a highly delicate, challenging, but quite important problem to be addressed.…

Artificial Intelligence · Computer Science 2023-05-10 Ercan atam

Automated data quality assessment is crucial for managing big data, but existing solutions face challenges in achieving accurate context-aware assessment. This paper presents a novel knowledge-based approach to enhance automated data…

Machine Learning · Computer Science 2026-05-21 Hadi Fadlallah , Rima Kilany , Mitri Haber , Ali Jaber

In the implementation and use of research information systems (RIS) in scientific institutions, text data mining and semantic technologies are a key technology for the meaningful use of large amounts of data. It is not the collection of…

Digital Libraries · Computer Science 2018-12-12 Otmane Azeroual , Gunter Saake , Mohammad Abuosba , Joachim Schöpfel

The increasing availability of semantic data has substantially enhanced Web applications. Semantic data such as RDF data is commonly represented as entity-property-value triples. The magnitude of semantic data, in particular the large…

Information Retrieval · Computer Science 2021-05-12 Qingxia Liu , Gong Cheng , Kalpa Gunaratna , Yuzhong Qu

When it comes to cluster massive data, response time, disk access and quality of formed classes becoming major issues for companies. It is in this context that we have come to define a clustering framework for large scale heterogeneous data…

Databases · Computer Science 2017-07-04 Mohamed Ali Zoghlami , Olfa Arfaoui , Minyar Sassi Hidri , Rahma Ben Ayed

The heterogeneity of Point of Interest (POI) taxonomies is a persistent challenge for the integration of urban datasets and the development of location-based services. OpenStreetMap (OSM) adopts a flexible, community-driven tagging system,…

Post-training Quantization (PTQ) technique has been extensively adopted for large language models (LLMs) compression owing to its efficiency and low resource requirement. However, current research lacks a in-depth analysis of the superior…

Machine Learning · Computer Science 2025-05-22 Jiaqi Zhao , Ming Wang , Miao Zhang , Yuzhang Shang , Xuebo Liu , Yaowei Wang , Min Zhang , Liqiang Nie

Data-driven technologies have improved the efficiency, reliability and effectiveness of healthcare services, but come with an increasing demand for data, which is challenging due to privacy-related constraints on sharing data in healthcare…

Topological Data Analysis (TDA) is a recent approach to analyze data sets from the perspective of their topological structure. Its use for time series data has been limited to the field of financial time series primarily and as a method for…

Machine Learning · Computer Science 2019-06-18 Rodrigo Rivera-Castro , Polina Pilyugina , Alexander Pletnev , Ivan Maksimov , Wanyi Wyz , Evgeny Burnaev

Data augmentation is a series of techniques that generate high-quality artificial data by manipulating existing data samples. By leveraging data augmentation techniques, AI models can achieve significantly improved applicability in tasks…

Machine Learning · Computer Science 2025-10-16 Zaitian Wang , Pengfei Wang , Kunpeng Liu , Pengyang Wang , Yanjie Fu , Chang-Tien Lu , Charu C. Aggarwal , Jian Pei , Yuanchun Zhou

Previous work on review summarization focused on measuring the sentiment toward the main aspects of the reviewed product or business, or on creating a textual summary. These approaches provide only a partial view of the data: aspect-based…

Computation and Language · Computer Science 2021-06-15 Roy Bar-Haim , Lilach Eden , Yoav Kantor , Roni Friedman , Noam Slonim

Currently, data-driven discovery in biological sciences resides in finding segmentation strategies in multivariate data that produce sensible descriptions of the data. Clustering is but one of several approaches and sometimes falls short…

Quantitative Methods · Quantitative Biology 2022-08-12 Richard Tjörnhammar

As Grids are loosely-coupled congregations of geographically distributed heterogeneous resources, the efficient utilization of the resources requires the support of a sound Performance Prediction System (PPS). The performance prediction of…

Distributed, Parallel, and Cluster Computing · Computer Science 2013-07-15 Sena Seneviratne , David C. Levy , Rajkumar Buyya

Benchmark datasets in computer vision often contain off-topic images, near duplicates, and label errors, leading to inaccurate estimates of model performance. In this paper, we revisit the task of data cleaning and formalize it as either a…

Recent advances in generating synthetic data that allow to add principled ways of protecting privacy -- such as Differential Privacy -- are a crucial step in sharing statistical information in a privacy preserving way. But while the focus…

Machine Learning · Statistics 2021-10-04 Christian Arnold , Marcel Neunhoeffer

Data representation remains a fundamental challenge in machine learning, particularly when adapting sequence-based architectures like Transformers and Large Language Models (LLMs) for structured tabular data. Existing methods often fail to…

Machine Learning · Computer Science 2025-08-05 Kayvan Karim , Hani Ragab Hassen. Hadj Batatia