English
Related papers

Related papers: LeMat-Bulk: aggregating, and de-duplicating quantu…

200 papers

The advancement of polymer informatics has been significantly propelled by the integration of machine learning (ML) techniques, enabling the rapid prediction of polymer properties and expediting the discovery of high-performance polymeric…

Materials Science · Physics 2025-04-01 Jiaxin Xu , Gang Liu , Ruilan Guo , Meng Jiang , Tengfei Luo

Molecular property prediction integrates quantum chemistry, cheminformatics, and deep learning to connect molecular structure with physicochemical and biological behavior. This survey traces four complementary paradigms, including Quantum,…

Unraveling the hierarchical structure-property relationships is the central challenge of materials science, necessitating the interpretation of data across vast physical scales from micro to macro. Despite the rapid integration of Large…

Digital Libraries · Computer Science 2026-03-23 Yuting Zheng , Zijian Chen , Qi Jia

Softmax-based losses have achieved state-of-the-art performances on various tasks such as face recognition and re-identification. However, these methods highly relied on clean datasets with global labels, which limits their usage in many…

Computer Vision and Pattern Recognition · Computer Science 2022-01-25 Qiang Meng , Xinqian Gu , Xiaqing Xu , Feng Zhou

Reproducibility and comparability of empirical results are at the core tenet of the scientific method in any scientific field. To ease reproducibility of empirical studies, several benchmarks in software engineering research, such as…

Software Engineering · Computer Science 2021-04-01 José Campos , André Souto

While the open-source software development model has led to successful large-scale collaborations in building software systems, data science projects are frequently developed by individuals or small teams. We describe challenges to scaling…

Machine Learning · Computer Science 2021-10-26 Micah J. Smith , Jürgen Cito , Kelvin Lu , Kalyan Veeramachaneni

We present MaterialFigBench, a benchmark dataset designed to evaluate the ability of multimodal large language models (LLMs) to solve university-level materials science problems that require accurate interpretation of figures. Unlike…

Computation and Language · Computer Science 2026-03-13 Michiko Yoshitake , Yuta Suzuki , Ryo Igarashi , Yoshitaka Ushiku , Keisuke Nagato

One of the key requirements for incorporating machine learning into the drug discovery process is complete reproducibility and traceability of the model building and evaluation process. With this in mind, we have developed an end-to-end…

Peptide therapeutics are widely regarded as the "third generation" of drugs, yet progress in peptide Machine Learning (ML) are hindered by the absence of standardized benchmarks. Here we present PepBenchmark, which unifies datasets,…

Machine Learning · Computer Science 2026-04-14 Jiahui Zhang , Rouyi Wang , Kuangqi Zhou , Tianshu Xiao , Lingyan Zhu , Yaosen Min , Yang Wang

We introduce a new molecular dataset, named Alchemy, for developing machine learning models useful in chemistry and material science. As of June 20th 2019, the dataset comprises of 12 quantum mechanical properties of 119,487 organic…

As large language models (LLMs) grow in size and deployment scale, quantization has become an essential technique for reducing memory footprint and improving inference efficiency. However, existing quantization toolkits often lack…

Machine Learning · Computer Science 2025-12-01 Dong Liu , Yanxuan Yu

The Variational Quantum Eigensolver (VQE) is a promising algorithm for quantum computing applications in chemistry and materials science, particularly in addressing the limitations of classical methods for complex systems. This study…

Quantum Physics · Physics 2025-02-25 Nia Pollard , Kamal Choudhary

In the era of advanced artificial intelligence, highlighted by large-scale generative models like GPT-4, ensuring the traceability, verifiability, and reproducibility of datasets throughout their lifecycle is paramount for research…

Software Engineering · Computer Science 2024-08-19 Yue Liu , Dawen Zhang , Boming Xia , Julia Anticev , Tunde Adebayo , Zhenchang Xing , Moses Machao

Contemporary large language model (LLM) training pipelines require the assembly of internet-scale databases full of text data from a variety of sources (e.g., web, academic, and publishers). Preprocessing these datasets via deduplication --…

The identification and localization of errors is a core task in peer review, yet the exponential growth of scientific output has made it increasingly difficult for human reviewers to reliably detect errors given the limited pool of experts.…

Computation and Language · Computer Science 2025-12-01 Sarina Xi , Vishisht Rao , Justin Payan , Nihar B. Shah

A fine-grained data recipe is crucial for pre-training large language models, as it can significantly enhance training efficiency and model performance. One important ingredient in the recipe is to select samples based on scores produced by…

Computation and Language · Computer Science 2026-01-01 Ziqing Fan , Yuqiao Xian , Yan Sun , Li Shen

The development of materials science is undergoing a shift from empirical approaches to data-driven and algorithm-oriented research paradigm. The state-of-the-art platforms are confined to inorganic crystals, with limited chemical space,…

Materials Science · Physics 2025-07-08 Jifeng Wang , Jiazhe Ju , Ying Wang

Question Answering (QA) effectively evaluates language models' reasoning and knowledge depth. While QA datasets are plentiful in areas like general domain and biomedicine, academic chemistry is less explored. Chemical QA plays a crucial…

Computation and Language · Computer Science 2024-07-25 Xiuying Chen , Tairan Wang , Taicheng Guo , Kehan Guo , Juexiao Zhou , Haoyang Li , Mingchen Zhuge , Jürgen Schmidhuber , Xin Gao , Xiangliang Zhang

Neural networks are the backbone of modern artificial intelligence, but designing, evaluating, and comparing them remains labor-intensive. While numerous datasets exist for training, there are few standardized collections of the models…

In analytical applications, database systems often need to sustain workloads with multiple concurrent scans hitting the same table. The Cooperative Scans (CScans) framework, which introduces an Active Buffer Manager (ABM) component into the…

Databases · Computer Science 2012-08-22 Michał Świtakowski , Peter Boncz , Marcin Żukowski