English
Related papers

Related papers: A High-Accuracy Optical Music Recognition Method B…

200 papers

Deep learning approaches to optical flow estimation have seen rapid progress over the recent years. One common trait of many networks is that they refine an initial flow estimate either through multiple stages or across the levels of a…

Computer Vision and Pattern Recognition · Computer Science 2019-04-11 Junhwa Hur , Stefan Roth

Traditional methods to tackle many music information retrieval tasks typically follow a two-step architecture: feature engineering followed by a simple learning algorithm. In these "shallow" architectures, feature engineering and learning…

Sound · Computer Science 2015-11-18 Peter Li , Jiyuan Qian , Tian Wang

The Otsu thresholding algorithm represents a fundamental technique in image segmentation, yet its computational efficiency is severely limited by exhaustive search requirements across all possible threshold values. This work presents an…

Computer Vision and Pattern Recognition · Computer Science 2025-09-22 Sai Varun Kodathala

In this study, we present a novel end-to-end approach based on the encoder-decoder framework with the attention mechanism for online handwritten mathematical expression recognition (OHMER). First, the input two-dimensional ink trajectory…

Computer Vision and Pattern Recognition · Computer Science 2017-12-13 Jianshu Zhang , Jun Du , Lirong Dai

Recently, trimap-free methods have drawn increasing attention in human video matting due to their promising performance. Nevertheless, these methods still suffer from the lack of deterministic foreground-background cues, which impairs their…

Computer Vision and Pattern Recognition · Computer Science 2025-03-04 Huayu Zhang , Dongyue Wu , Yuanjie Shao , Nong Sang , Changxin Gao

Optical Music Recognition (OMR) is an important and challenging area within music information retrieval, the accurate detection of music symbols in digital images is a core functionality of any OMR pipeline. In this paper, we introduce a…

Computer Vision and Pattern Recognition · Computer Science 2018-05-29 Lukas Tuggener , Ismail Elezi , Jurgen Schmidhuber , Thilo Stadelmann

In this paper we present an on-manifold sequence-to-sequence learning approach to motion estimation using visual and inertial sensors. It is to the best of our knowledge the first end-to-end trainable method for visual-inertial odometry…

Computer Vision and Pattern Recognition · Computer Science 2017-04-04 Ronald Clark , Sen Wang , Hongkai Wen , Andrew Markham , Niki Trigoni

Existing learning-based video compression methods still face challenges related to inaccurate motion estimates and inadequate motion compensation structures. These issues result in compression errors and a suboptimal rate-distortion…

Image and Video Processing · Electrical Eng. & Systems 2025-03-13 Md baharul Islam , Afsana Ahsan Jeny

In this work, we propose a training algorithm for an audio-visual automatic speech recognition (AV-ASR) system using deep recurrent neural network (RNN).First, we train a deep RNN acoustic model with a Connectionist Temporal Classification…

Computer Vision and Pattern Recognition · Computer Science 2016-11-10 Abhinav Thanda , Shankar M Venkatesan

This paper presents a novel training-free framework for open-vocabulary image segmentation and object recognition (OVSR), which leverages EfficientNetB0, a convolutional neural network, for unsupervised segmentation and CLIP, a…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Ying Dai , Wei Yu Chen

We present a novel end-to-end visual odometry architecture with guided feature selection based on deep convolutional recurrent neural networks. Different from current monocular visual odometry methods, our approach is established on the…

Computer Vision and Pattern Recognition · Computer Science 2018-11-27 Fei Xue , Qiuyuan Wang , Xin Wang , Wei Dong , Junqiu Wang , Hongbin Zha

This paper introduces a modeling approach that employs multi-level global processing, encompassing both short-term frame-level and long-term sample-level feature scales. In the initial stage of shallow feature extraction, various scales are…

Sound · Computer Science 2024-11-07 Chunyan Zeng , Yuhao Zhao , Zhifeng Wang

We present a framework based on neural networks to extract music scores directly from polyphonic audio in an end-to-end fashion. Most previous Automatic Music Transcription (AMT) methods seek a piano-roll representation of the pitches, that…

Sound · Computer Science 2019-10-29 Miguel A. Román , Antonio Pertusa , Jorge Calvo-Zaragoza

A new musical instrument classification method using convolutional neural networks (CNNs) is presented in this paper. Unlike the traditional methods, we investigated a scheme for classifying musical instruments using the learned features…

Sound · Computer Science 2015-12-24 Taejin Park , Taejin Lee

Machine learning approaches to auditory object recognition are traditionally based on engineered features such as those derived from the spectrum or cepstrum. More recently, end-to-end classification systems in image and auditory…

Residual networks, as discrete approximations of Ordinary Differential Equations (ODEs), have inspired significant advancements in neural network design, including multistep methods, high-order methods, and multi-particle dynamical systems.…

Computation and Language · Computer Science 2024-11-06 Bei Li , Tong Zheng , Rui Wang , Jiahao Liu , Qingyan Guo , Junliang Guo , Xu Tan , Tong Xiao , Jingbo Zhu , Jingang Wang , Xunliang Cai

Residual moveout (RMO) provides critical information for travel time tomography. The current industry-standard method for fitting RMO involves scanning high-order polynomial equations. However, this analytical approach does not accurately…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Hongtao Wang , Jiandong Liang , Lei Wang , Shuaizhe Liang , Jinping Zhu , Chunxia Zhang , Jiangshe Zhang

The goal of image ordinal estimation is to estimate the ordinal label of a given image with a convolutional neural network. Existing methods are mainly based on ordinal regression and particularly focus on modeling the ordinal mapping from…

Computer Vision and Pattern Recognition · Computer Science 2023-01-18 Yiming Lei , Zilong Li , Yangyang Li , Junping Zhang , Hongming Shan

The rise of multi-modal search requests from users has highlighted the importance of multi-modal retrieval (i.e. image-to-text or text-to-image retrieval), yet the more complex task of image-to-multi-modal retrieval, crucial for many…

Information Retrieval · Computer Science 2024-06-11 Zida Cheng , Chen Ju , Shuai Xiao , Xu Chen , Zhonghua Zhai , Xiaoyi Zeng , Weilin Huang , Junchi Yan

Long-context modeling is essential for symbolic music generation, since motif repetition and developmental variation can span thousands of musical events, yet practical workflows frequently rely on resource-limited hardware. We introduce…

Sound · Computer Science 2026-03-03 Yungang Yi , Weihua Li , Matthew Kuo , Catherine Shi , Quan Bai