English
Related papers

Related papers: PLIT: An alignment-free computational tool for ide…

200 papers

Industrial multi-label document understanding pipelines score candidate labels and threshold or rank them to form a label set per document. This early selection step directly affects the accuracy of downstream information extraction from…

Information Retrieval · Computer Science 2026-05-19 Lasal Jayawardena , Nirmalie Wiratunga , Ikechukwu Nkisi-Orji , Darren Nicol

RNAs are essential molecules that carry genetic information vital for life, with profound implications for drug development and biotechnology. Despite this importance, RNA research is often hindered by the vast literature available on the…

Genomics · Quantitative Biology 2024-11-15 Yijia Xiao , Edward Sun , Yiqiao Jin , Wei Wang

Automated classification of clinical transcriptions into medical specialties is essential for routing, coding, and clinical decision support, yet prior work on the widely used MTSamples benchmark suffers from severe data leakage caused by…

Artificial Intelligence · Computer Science 2026-03-25 Pronob Kumar Barman , Pronoy Kumar Barman

Patent classification into CPC codes underpins large scale analyses of technological change but remains challenging due to its hierarchical, multi label, and highly imbalanced structure. While pre Generative AI supervised encoder based…

Computational Engineering, Finance, and Science · Computer Science 2026-02-02 Lorenzo Emer , Marco Lippi , Andrea Mina , Andrea Vandin

Collapse Lineage Tree (CLTree) is a software tool that annotates, roots, and evaluates phylogenetic trees by using lineages. A recursive algorithm was designed to annotate the branches by the common taxonomic lineage of its descendants in a…

Populations and Evolution · Quantitative Biology 2025-11-25 Guanghong Zuo

The scarcity of high-quality multimodal biomedical data limits the ability to effectively fine-tune pretrained Large Language Models (LLMs) for specialized biomedical tasks. To address this challenge, we introduce MINT (Multimodal…

Quantitative Methods · Quantitative Biology 2026-02-18 Zhanliang Wang , Da Wu , Quan Nguyen , Zhuoran Xu , Kai Wang

Single-cell RNA sequencing (scRNA-seq) technology enables systematic delineation of cellular states and interactions, providing crucial insights into cellular heterogeneity. Building on this potential, numerous computational methods have…

Genomics · Quantitative Biology 2025-11-11 Ping Xu , Zaitian Wang , Zhirui Wang , Pengjiang Li , Ran Zhang , Gaoyang Li , Hanyu Xie , Jiajia Wang , Yuanchun Zhou , Pengfei Wang

Fine-tuning pre-trained language models (PLMs), e.g., SciBERT, generally requires large numbers of annotated data to achieve state-of-the-art performance on a range of NLP tasks in the scientific domain. However, obtaining the fine-tune…

Computation and Language · Computer Science 2025-12-30 Jiawei Liu , Zi Xiong , Yi Jiang , Yongqiang Ma , Wei Lu , Yong Huang , Qikai Cheng

We propose an interactive visual analytics tool, Vis-SPLIT, for partitioning a population of individuals into groups with similar gene signatures. Vis-SPLIT allows users to interactively explore a dataset and exploit visual separations to…

Human-Computer Interaction · Computer Science 2024-02-19 Braden Roper , James C. Mathews , Saad Nadeem , Ji Hwan Park

The reuse of historical clinical trial data has significant potential to accelerate medical research and drug development. However, interoperability challenges, particularly with missing medical codes, hinders effective data integration…

Machine Learning · Computer Science 2025-03-14 Nabeel Seedat , Caterina Tozzi , Andrea Hita Ardiaca , Mihaela van der Schaar , James Weatherall , Adam Taylor

In the global challenge of understanding and characterizing biodiversity, short species-specific genomic sequences known as DNA barcodes play a critical role, enabling fine-grained comparisons among organisms within the same kingdom of…

High-throughput phenotyping refers to the non-destructive and efficient evaluation of plant phenotypes. In recent years, it has been coupled with machine learning in order to improve the process of phenotyping plants by increasing…

Computer Vision and Pattern Recognition · Computer Science 2024-07-11 Vivaan Singhvi , Langalibalele Lunga , Pragya Nidhi , Chris Keum , Varrun Prakash

We propose a modification of linear discriminant analysis, referred to as compressive regularized discriminant analysis (CRDA), for analysis of high-dimensional datasets. CRDA is specially designed for feature elimination purpose and can be…

Methodology · Statistics 2018-04-12 Muhammad Naveed Tabassum , Esa Ollila

Deep learning has markedly advanced image based plant disease diagnosis as improved hardware and dataset quality have enabled increasingly accurate neural network models. This paper presents PD36 C, a compact convolutional neural network…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Shkelqim Sherifi

Navigating the complex landscape of single-cell transcriptomic data presents significant challenges. Central to this challenge is the identification of a meaningful representation of high-dimensional gene expression patterns that sheds…

Quantitative Methods · Quantitative Biology 2023-12-13 Mu Qiao

Functional annotation of microbial genomes is often biased toward protein-coding genes, leaving a vast, unexplored landscape of non-coding RNAs (ncRNAs) that are critical for regulating bacterial and archaeal physiology, stress response and…

Genomics · Quantitative Biology 2025-07-17 Lauren Lui , Torben Nielsen

Lindenmayer systems (L-systems) are a formal grammar system that iteratively rewrites all symbols of a string, in parallel. When visualized with a graphical interpretation, the images have self-similar shapes that appear frequently in…

Artificial Intelligence · Computer Science 2017-12-05 Jason Bernard , Ian McQuillan

This paper develops statistical methods for determining the number of components in panel data finite mixture regression models with regression errors independently distributed as normal or more flexible normal mixtures. We analyze the…

Econometrics · Economics 2025-06-12 Yu Hao , Hiroyuki Kasahara

Clustering is a difficult and widely-studied data mining task, with many varieties of clustering algorithms proposed in the literature. Nearly all algorithms use a similarity measure such as a distance metric (e.g. Euclidean distance) to…

Neural and Evolutionary Computing · Computer Science 2019-10-24 Andrew Lensen , Bing Xue , Mengjie Zhang

DNA sequence classification is a fundamental task in computational biology with vast implications for applications such as disease prevention and drug design. Therefore, fast high-quality sequence classifiers are significantly important.…

Machine Learning · Computer Science 2023-11-07 Marcel Khalifa , Barak Hoffer , Orian Leitersdorf , Robert Hanhan , Ben Perach , Leonid Yavits , Shahar Kvatinsky