中文
相关论文

相关论文: Tackling dataset curation challenges towards relia…

200 篇论文

The rapid advancement of large language models (LLMs) has sparked interest in data synthesis techniques, aiming to generate diverse and high-quality synthetic datasets. However, these synthetic datasets often suffer from a lack of diversity…

The driving factors behind the development of large language models (LLMs) with impressive learning capabilities are their colossal model sizes and extensive training datasets. Along with the progress in natural language processing, LLMs…

Nowadays, machine learning (ML) plays a vital role in many aspects of our daily life. In essence, building well-performing ML applications requires the provision of high-quality data throughout the entire life-cycle of such applications.…

数据库 · 计算机科学 2023-02-10 Mohamed Abdelaal , Christian Hammacher , Harald Schoening

One of the main goals and challenges of materials discovery is to find the best candidates for each interest property or application. Machine learning rises in this context to efficiently optimize this search, exploring the immense…

材料科学 · 物理学 2021-08-04 Gabriel R. Schleder , Bruno Focassio , Adalberto Fazzio

The quality of foundation models depends heavily on their training data. Consequently, great efforts have been put into dataset curation. Yet most approaches rely on manual tuning of coarse-grained mixtures of large buckets of data, or…

High-throughput data generation methods and machine learning (ML) algorithms have given rise to a new era of computational materials science by learning relationships among composition, structure, and properties and by exploiting such…

Thermoelectric (TE) materials are among very few sustainable yet feasible energy solutions of present time. This huge promise of energy harvesting is contingent on identifying/designing materials having higher efficiency than presently…

Patient-trial matching requires reasoning over long, heterogeneous electronic health records (EHRs) and complex eligibility criteria, posing significant challenges for scalability, generalization, and computational efficiency. Existing…

The effective utilization at scale of complex machine learning (ML) techniques for HEP use cases poses several technological challenges, most importantly on the actual implementation of dedicated end-to-end data pipelines. A solution to…

分布式、并行与集群计算 · 计算机科学 2020-06-17 Matteo Migliorini , Riccardo Castellotti , Luca Canali , Marco Zanetti

Large language models (LLMs) are increasingly exposed to data contamination, i.e., performance gains driven by prior exposure of test datasets rather than generalization. However, in the context of tabular data, this problem is largely…

计算与语言 · 计算机科学 2026-03-31 Matteo Silvestri , Fabiano Veglianti , Flavio Giorgi , Fabrizio Silvestri , Gabriele Tolomei

Microplastics (MPs) are ubiquitous pollutants with demonstrated potential to impact ecosystems and human health. Their microscopic size complicates detection, classification, and removal, especially in biological and environmental samples.…

计算机视觉与模式识别 · 计算机科学 2025-10-22 Paul-Tiberiu Miclea , Martin Sboron , Hardik Vaghasiya , Hoang Thinh Nguyen , Meet Gadara , Thomas Schmid

Modern materials science generates vast and diverse datasets from both experiments and computations, yet these multi-source, heterogeneous data often remain disconnected in isolated "silos". Here, we introduce MaterialsGalaxy, a…

材料科学 · 物理学 2025-11-03 Tiannian Zhu , Zhong Fang , Quansheng Wu , Hongming Weng

Synthesis of advanced inorganic materials with minimum number of trials is of paramount importance towards the acceleration of inorganic materials development. The enormous complexity involved in existing multi-variable synthesis methods…

材料科学 · 物理学 2020-11-02 Bijun Tang , Yuhao Lu , Jiadong Zhou , Han Wang , Prafful Golani , Manzhang Xu , Quan Xu , Cuntai Guan , Zheng Liu

Process optimization in chemical engineering may be hindered by the limited availability of reliable thermodynamic data for fluid mixtures. Remarkable progress is being made in predicting thermodynamic mixture properties by machine learning…

计算工程、金融与科学 · 计算机科学 2025-10-14 Martin Bubel , Tobias Seidel , Michael Bortz

The quality of datasets is one of the key factors that affect the accuracy of aerodynamic data models. For example, in the uniformly sampled Burgers' dataset, the insufficient high-speed data is overwhelmed by massive low-speed data.…

机器学习 · 计算机科学 2020-10-20 Liwei Hu , Yu Xiang , Jun Zhan , Zifang Shi , Wenzheng Wang

Due to their abundant use in all-solid-state lasers, nonlinear optical (NLO) crystals are needed for many applications across diverse fields such as medicine and communication. However, because of conflicting requirements, the design of…

材料科学 · 物理学 2025-08-13 Victor Trinquet , Matthew L. Evans , Gian-Marco Rignanese

Luminescence thermometry has been extensively exploited in the last decades both from the fundamental and applied point of views. The application of photoluminescent nanoparticles on the microscopic level based on rare-earth doped (RED)…

As the number of applications that use machine learning algorithms increases, the need for labeled data useful for training such algorithms intensifies. Getting labels typically involves employing humans to do the annotation, which directly…

机器学习 · 计算机科学 2013-07-16 Alexandros Ntoulas , Omar Alonso , Vasilis Kandylas

The application of machine learning in materials presents a unique challenge of dealing with scarce and varied materials data - both experimental and theoretical. Nevertheless, several state-of-the-art machine learning models for materials…

How to generate a large, realistic set of tables along with joinability relationships, to stress-test dataset discovery methods? Dataset discovery methods aim to automatically identify related data assets in a data lake. The development and…

数据库 · 计算机科学 2025-07-09 Zhenwei Dai , Chuan Lei , Asterios Katsifodimos , Xiao Qin , Christos Faloutsos , Huzefa Rangwala