English
Related papers

Related papers: PREFER: An Ontology for the PREcision FERmentation…

200 papers

We introduce OTTER, a unified open-set multi-label tagging framework that harmonizes the stability of a curated, predefined category set with the adaptability of user-driven open tags. OTTER is built upon a large-scale, hierarchically…

Computer Vision and Pattern Recognition · Computer Science 2025-10-02 Jieer Ouyang , Xiaoneng Xiang , Zheng Wang , Yangkai Ding

We propose a novel framework to facilitate the on-demand design of data-centric systems by exploiting domain knowledge from an existing ontology. Its key ingredient is a process that we call focusing, which allows to obtain a schema for a…

Logic in Computer Science · Computer Science 2019-04-02 Tomasz Gogacz , Víctor Gutiérrez-Basulto , Yazmín A. Ibáñez-García , Filip Murlak , Magdalena Ortiz , Mantas Šimkus

While machine learning models have achieved unprecedented success in real-world applications, they might make biased/unfair decisions for specific demographic groups and hence result in discriminative outcomes. Although research efforts…

Machine Learning · Computer Science 2022-12-08 Yuying Zhao , Yu Wang , Tyler Derr

As human machine teaming becomes central to paradigms like Industry 5.0, a critical need arises for machines to safely and effectively interpret complex human behaviors. A research gap currently exists between techno centric robotic…

Artificial Intelligence · Computer Science 2025-10-28 Alexis Ellis , Stacie Severyn , Fjollë Novakazi , Hadi Banaee , Cogan Shimizu

Dataset distillation (DD) has witnessed significant progress in creating small datasets that encapsulate rich information from large original ones. Particularly, methods based on generative priors show promising performance, while…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Jianyang Gu , Haonan Wang , Ruoxi Jia , Saeed Vahidian , Vyacheslav Kungurtsev , Wei Jiang , Yiran Chen

Data provenance, or data lineage, describes the life cycle of data. In scientific workflows on HPC systems, scientists often seek diverse provenance (e.g., origins of data products, usage patterns of datasets). Unfortunately, existing…

Distributed, Parallel, and Cluster Computing · Computer Science 2023-08-03 Runzhou Han , Mai Zheng , Suren Byna , Houjun Tang , Bin Dong , Dong Dai , Yong Chen , Dongkyun Kim , Joseph Hassoun , David Thorsley , Matthew Wolf

The development of materials science is undergoing a shift from empirical approaches to data-driven and algorithm-oriented research paradigm. The state-of-the-art platforms are confined to inorganic crystals, with limited chemical space,…

Materials Science · Physics 2025-07-08 Jifeng Wang , Jiazhe Ju , Ying Wang

The GFF3 format is a common, flexible tab-delimited format representing the structure and function of genes or other mapped features (https://github.com/The-Sequence-Ontology/Specifications/blob/master/gff3.md). However, with increasing…

Open rule refer to the implication from premise atoms to hypothesis atoms, which captures various relations between instances in the real world. Injecting open rule knowledge into the machine helps to improve the performance of downstream…

Computation and Language · Computer Science 2024-11-05 Jianyu Liu , Sheng Bi , Guilin Qi

We introduce IntFold, a controllable foundation model for general and specialized biomolecular structure prediction. Utilizing a high-performance custom attention kernel, IntFold achieves accuracy comparable to the state-of-the-art…

Biomolecules · Quantitative Biology 2025-07-08 The IntFold Team , Leon Qiao , Wayne Bai , He Yan , Gary Liu , Nova Xi , Xiang Zhang , Siqi Sun

Machine learning (ML) is increasingly being used to make decisions in our society. ML models, however, can be unfair to certain demographic groups (e.g., African Americans or females) according to various fairness metrics. Existing…

Computers and Society · Computer Science 2021-03-17 Hantian Zhang , Xu Chu , Abolfazl Asudeh , Shamkant B. Navathe

The increasing availability of high throughput data arising from gene expression studies leads to the necessity of methods for summarizing the available information. As annotation quality improves it is becoming common to rely on the Gene…

Genomics · Quantitative Biology 2007-05-23 Alex Sanchez-Pla , Miquel Salicru , Jordi Ocanya

Motivation: Ontologies are widely used in biology for data annotation, integration, and analysis. In addition to formally structured axioms, ontologies contain meta-data in the form of annotation axioms which provide valuable pieces of…

Computation and Language · Computer Science 2018-05-01 Fatima Zohra Smaili , Xin Gao , Robert Hoehndorf

Our goal is procedural text comprehension, namely tracking how the properties of entities (e.g., their location) change with time given a procedural text (e.g., a paragraph about photosynthesis, a recipe). This task is challenging as the…

Computation and Language · Computer Science 2019-06-24 Xinya Du , Bhavana Dalvi Mishra , Niket Tandon , Antoine Bosselut , Wen-tau Yih , Peter Clark , Claire Cardie

With the increased interest in computational sciences, machine learning (ML), pattern recognition (PR) and big data, governmental agencies, academia and manufacturers are overwhelmed by the constant influx of new algorithms and techniques…

Software Engineering · Computer Science 2017-07-28 André Anjos , Laurent El-Shafey , Sébastien Marcel

Motivation: The proliferation of Biomedical research articles has made the task of information retrieval more important than ever. Scientists and Researchers are having difficulty in finding articles that contain information relevant to…

Computation and Language · Computer Science 2020-11-04 Harsh Patel

While Large Language Models (LLMs) have demonstrated remarkable progress in generating functionally correct Solidity code, they continue to face critical challenges in producing gas-efficient and secure code, which are critical requirements…

Software Engineering · Computer Science 2025-10-01 Zhiyuan Peng , Xin Yin , Chenhao Ying , Chao Ni , Yuan Luo

A data-driven quantitative approach was used to develop a novel classification system for beer categories and styles. Sixty-two thousand one hundred twenty-one beer recipes were mined and analyzed, considering ingredient profiles,…

Computation and Language · Computer Science 2025-05-26 Diego Bonatto

In theory, ontology-based process modelling (OBPM) bares great potential to extend business process management. Many works have studied OBPM and are clear on the potential amenities, such as eliminating ambiguities or enabling advanced…

Artificial Intelligence · Computer Science 2021-07-14 Carl Corea , Michael Fellmann , Patrick Delfmann

Scientific data management is at a critical juncture, driven by exponential data growth, increasing cross-domain dependencies, and a severe reproducibility crisis in modern research. Traditional centralized data management approaches are…

Databases · Computer Science 2025-04-30 Sebastian Beyvers , Jannis Hochmuth , Lukas Brehm , Maria Hansen , Alexander Goesmann , Frank Förster