中文
相关论文

相关论文: Using Sequence Ensembles for Seeding Alignments of…

200 篇论文

Learning to read words aloud is a major step towards becoming a reader. Many children struggle with the task because of the inconsistencies of English spelling-sound correspondences. Curricula vary enormously in how these patterns are…

机器学习 · 计算机科学 2020-07-03 Ayon Sen , Christopher R. Cox , Matthew Cooper Borkenhagen , Mark S. Seidenberg , Xiaojin Zhu

DNA is a leading candidate as the next archival storage media due to its density, durability and sustainability. To read (and write) data DNA storage exploits technology that has been developed over decades to sequence naturally occurring…

新兴技术 · 计算机科学 2022-05-12 Jasmine Quah , Omer Sella , Thomas Heinis

Annotation projection is an important area in NLP that can greatly contribute to creating language resources for low-resource languages. Word alignment plays a key role in this setting. However, most of the existing word alignment methods…

计算与语言 · 计算机科学 2021-06-17 Ehsaneddin Asgari , Masoud Jalili Sabet , Philipp Dufter , Christopher Ringlstetter , Hinrich Schütze

Identifying cell clusters is a critical step for single-cell transcriptomics study. Despite the numerous clustering tools developed recently, the rapid growth of scRNA-seq volumes prompts for a more (computationally) efficient clustering…

定量方法 · 定量生物学 2023-01-11 Nana Wei , Yating Nie , Lin Liu , Xiaoqi Zheng , Hua-Jun Wu4

Transcript enumeration methods such as SAGE, MPSS, and sequencing-by-synthesis EST ``digital northern'', are important high-throughput techniques for digital gene expression measurement. As other counting or voting processes, these…

定量方法 · 定量生物学 2013-10-29 Ricardo ZN Vêncio , Leonardo Varuzza , Carlos AB Pereira , Helena Brentani , Ilya Shmulevich

Nanopore sequencing accuracy is inherently limited by the quality of data from individual molecular translocation events, requiring advances beyond traditional sequencing-by-synthesis methods. We introduce an oxidized pyramidal sub-nm pore…

应用物理 · 物理学 2026-05-22 Jianxin Yang , Dehua Hu , Wu Yuan , Tianle Pan , Ho-Pui Ho

We report SInC (SNV, Indel and CNV) simulator and read generator, an open-source tool capable of simulating biological variants taking into account a platform-specific error model. SInC is capable of simulating and generating single- and…

定量方法 · 定量生物学 2013-08-19 Swetansu Pattnaik , Saurabh Gupta , Arjun A Rao , Binay Panda

We propose a novel neural sequence prediction method based on \textit{error-correcting output codes} that avoids exact softmax normalization and allows for a tradeoff between speed and performance. Instead of minimizing measures between the…

机器学习 · 计算机科学 2019-09-06 James O' Neill , Danushka Bollegala

Motivation: Identifying genomic variants is an essential step for connecting genotype and phenotype. The usual approach consists of statistical inference of variants from alignments of sequencing reads. State-of-the-art variant callers can…

基因组学 · 定量生物学 2018-11-07 Karel Břinda , Valentina Boeva , Gregory Kucherov

Millimeter wave (mmWave) communication with large antenna arrays is a promising technique to enable extremely high data rates due to the large available bandwidth in mmWave frequency bands. In addition, given the knowledge of an optimal…

信息论 · 计算机科学 2019-09-05 Sung-En Chiu , Nancy Ronquillo , Tara Javidi

Advances in high-throughput sequencing technology have led to significant progress in measuring gene expressions at the single-cell level. The amount of publicly available single-cell RNA-seq (scRNA-seq) data is already surpassing 50M…

机器学习 · 计算机科学 2024-02-27 Jing Gong , Minsheng Hao , Xingyi Cheng , Xin Zeng , Chiming Liu , Jianzhu Ma , Xuegong Zhang , Taifeng Wang , Le Song

Near-duplicate text alignment is the task of identifying, among the texts in a corpus, all the subsequences (substrings) that are similar to a given query. Traditional approaches rely on seeding-extension-filtering heuristics, which lack…

数据库 · 计算机科学 2025-09-03 Yuheng Zhang , Miao Qiao , Zhencan Peng , Dong Deng

The prevalent technique for DNA sequencing consists of two main steps: shotgun sequencing, where many randomly located fragments, called reads, are extracted from the overall sequence, followed by an assembly algorithm that aims to…

基因组学 · 定量生物学 2016-01-28 Shirshendu Ganguly , Elchanan Mossel , Miklos Z. Racz

Medical image segmentation annotations exhibit variations among experts due to the ambiguous boundaries of segmented objects and backgrounds in medical images. Although using multiple annotations for each image in the fully-supervised has…

计算机视觉与模式识别 · 计算机科学 2023-11-20 Shuai Wang , Tengjin Weng , Jingyi Wang , Yang Shen , Zhidong Zhao , Yixiu Liu , Pengfei Jiao , Zhiming Cheng , Yaqi Wang

A sensor network is considered where at each sensor a sequence of random variables is observed. At each time step, a processed version of the observations is transmitted from the sensors to a common node called the fusion center. At some…

统计理论 · 数学 2023-07-19 Taposh Banerjee , Venugopal V. Veeravalli

We propose novel algorithms for sequence prediction based on ideas from stringology. These algorithms are time and space efficient and satisfy mistake bounds related to particular stringological complexity measures of the sequence. In this…

形式语言与自动机理论 · 计算机科学 2026-03-31 Vanessa Kosoy

Modern preference alignment techniques, such as Best-of-N (BoN) sampling, rely on reward models trained with pairwise comparison data. While effective at learning relative preferences, this paradigm fails to capture a signal of response…

统计方法学 · 统计学 2025-10-14 Hyung Gyu Rho , Sian Lee

A new algorithm is proposed which accelerates the mini-batch k-means algorithm of Sculley (2010) by using the distance bounding approach of Elkan (2003). We argue that, when incorporating distance bounds into a mini-batch algorithm, already…

机器学习 · 统计学 2016-09-14 James Newling , François Fleuret

Graded posets frequently arise throughout combinatorics, where it is natural to try to count the number of elements of a fixed rank. These counting problems are often $\#\textbf{P}$-complete, so we consider approximation algorithms for…

数据结构与算法 · 计算机科学 2023-04-11 Prateek Bhakta , Ben Cousins , Matthew Fahrbach , Dana Randall

Error Span Detection (ESD) extends automatic machine translation (MT) evaluation by localizing translation errors and labeling their severity. Current generative ESD methods typically use Maximum a Posteriori (MAP) decoding, assuming that…