English
Related papers

Related papers: A quick guide for student-driven community genome …

200 papers

Recently, deep models have been successfully applied in several applications, especially with low-level representations. However, sparse, noisy samples and structured domains (with multiple objects and interactions) are some of the open…

Machine Learning · Computer Science 2019-04-16 Mayukh Das , Yang Yu , Devendra Singh Dhami , Gautam Kunapuli , Sriraam Natarajan

Clinical notes are an efficient way to record patient information but are notoriously hard to decipher for non-experts. Automatically simplifying medical text can empower patients with valuable information about their health, while saving…

Biomarker discovery is vital in advancing personalized medicine, offering insights into disease diagnosis, prognosis, and therapeutic efficacy. Traditionally, the identification and validation of biomarkers heavily depend on extensive…

Machine Learning · Computer Science 2024-09-25 Wangyang Ying , Dongjie Wang , Xuanming Hu , Ji Qiu , Jin Park , Yanjie Fu

As natural language corpora expand at an unprecedented rate, manual annotation remains a significant methodological bottleneck in corpus linguistic work. We address this challenge by presenting a scalable pipeline for automating grammatical…

Computation and Language · Computer Science 2026-02-11 Cameron Morin , Matti Marttinen Larsson

Next-generation sequencing technologies generate millions of short sequence reads, which are usually aligned to a reference genome. In many applications, the key information required for downstream analysis is the number of reads mapping to…

Genomics · Quantitative Biology 2016-07-26 Yang Liao , Gordon K Smyth , Wei Shi

Post-training (via supervised fine-tuning) improves instruction-following, but often induces semantic mode collapse by biasing models toward low-entropy fine-tuning data at the expense of the high-entropy pretraining distribution.…

Adapting production-level computer vision tools to bespoke scientific datasets is a critical "last mile" bottleneck. Current solutions are impractical: fine-tuning requires large annotated datasets scientists often lack, while manual code…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Xuefei , Wang , Kai A. Horstmann , Ethan Lin , Jonathan Chen , Alexander R. Farhang , Sophia Stiles , Atharva Sehgal , Jonathan Light , David Van Valen , Yisong Yue , Jennifer J. Sun

Gene studies are crucial for fields such as protein structure prediction, drug discovery, and cancer genomics, yet they face challenges in fully utilizing the vast and diverse information available. Gene studies require clean, factual…

Databases · Computer Science 2024-12-18 Yuwei Miao , Yuzhi Guo , Hehuan Ma , Jingquan Yan , Feng Jiang , Weizhi An , Jean Gao , Junzhou Huang

We present components of an AI-assisted academic writing system including citation recommendation and introduction writing. The system recommends citations by considering the user's current document context to provide relevant suggestions.…

Artificial Intelligence · Computer Science 2025-03-19 Daniel J. Liebling , Malcolm Kane , Madeleine Grunde-Mclaughlin , Ian J. Lang , Subhashini Venugopalan , Michael P. Brenner

Annotation cost is a bottleneck for collecting massive data in mammography, especially for training deep neural networks. In this paper, we study the use of heterogeneous levels of annotation granularity to improve predictive performances.…

Image and Video Processing · Electrical Eng. & Systems 2019-09-13 Thi-Lam-Thuy Le , Nicolas Thome , Sylvain Bernard , Vincent Bismuth , Fanny Patoureaux

This paper presents our computational methodology using Genetic Algorithms (GA) for exploring the nature of RNA editing. These models are constructed using several genetic editing characteristics that are gleaned from the RNA editing system…

Neural and Evolutionary Computing · Computer Science 2007-05-23 C. Huang , L. M. Rocha

In the context of text classification, the financial burden of annotation exercises for creating training data is a critical issue. Active learning techniques, particularly those rooted in uncertainty sampling, offer a cost-effective…

Computation and Language · Computer Science 2024-06-19 Hamidreza Rouzegar , Masoud Makrehchi

Intelligent systems for the annotation of media content are increasingly being used for the automation of parts of social science research. In this domain the problem of integrating various Artificial Intelligence (AI) algorithms into a…

Multiagent Systems · Computer Science 2018-06-05 Ilias Flaounas , Thomas Lansdall-Welfare , Panagiota Antonakaki , Nello Cristianini

This handbook is a hands-on guide on how to approach text annotation tasks. It provides a gentle introduction to the topic, an overview of theoretical concepts as well as practical advice. The topics covered are mostly technical, but…

Advancements in clinical treatment are increasingly constrained by the limitations of supervised learning techniques, which depend heavily on large volumes of annotated data. The annotation process is not only costly but also demands…

Computer Vision and Pattern Recognition · Computer Science 2024-05-20 Pranav Singh , Raviteja Chukkapalli , Shravan Chaudhari , Luoyao Chen , Mei Chen , Jinqian Pan , Craig Smuda , Jacopo Cirrone

While artificial intelligence (AI) has become widespread, many commercial AI systems are not yet accessible to individual researchers nor the general public due to the deep knowledge of the systems required to use them. We believe that AI…

As genetic sequencing costs decrease, the lack of clinical interpretation of variants has become the bottleneck in using genetics data. A major rate limiting step in clinical interpretation is the manual curation of evidence in the genetic…

Computation and Language · Computer Science 2019-09-25 Allen Nie , Arturo L. Pineda , Matt W. Wright Hannah Wand , Bryan Wulf , Helio A. Costa , Ronak Y. Patel , Carlos D. Bustamante , James Zou

Despite the growing availability of tools designed to support scholarly knowledge extraction and organization, many researchers still rely on manual methods, sometimes due to unfamiliarity with existing technologies or limited access to…

Genomic data and biomedical imaging data are undergoing exponential growth. However, our understanding of the phenotype-genotype connection linking the two types of data is lagging behind. While there are many types of software that enable…

Genomics · Quantitative Biology 2013-01-09 Andrew T. Oberlin , Dominika A. Jurkovic , Mitchell F. Balish , Iddo Friedberg

Identifying disease genes from human genome is an important and fundamental problem in biomedical research. Despite many publications of machine learning methods applied to discover new disease genes, it still remains a challenge because of…

Quantitative Methods · Quantitative Biology 2017-05-23 Peng Yang