中文
相关论文

相关论文: Visual Attention for Musical Instrument Recognitio…

200 篇论文

Symbolic Music Emotion Recognition(SMER) is to predict music emotion from symbolic data, such as MIDI and MusicXML. Previous work mainly focused on learning better representation via (mask) language model pre-training but ignored the…

声音 · 计算机科学 2022-01-19 Jibao Qiu , C. L. Philip Chen , Tong Zhang

Detecting singing-voice in polyphonic instrumental music is critical to music information retrieval. To train a robust vocal detector, a large dataset marked with vocal or non-vocal label at frame-level is essential. However, frame-level…

音频与语音处理 · 电气工程与系统科学 2020-08-12 Yuanbo Hou , Frank K. Soong , Jian Luan , Shengchen Li

Humans usually perceive the world in a multimodal way that vision, touch, sound are utilised to understand surroundings from various dimensions. These senses are combined together to achieve a synergistic effect where the learning is more…

计算机视觉与模式识别 · 计算机科学 2021-12-30 Guanqun Cao , Shan Luo

Surgical instrument segmentation is extremely important for computer-assisted surgery. Different from common object segmentation, it is more challenging due to the large illumination and scale variation caused by the special surgical…

计算机视觉与模式识别 · 计算机科学 2020-05-25 Zhen-Liang Ni , Gui-Bin Bian , Guan-An Wang , Xiao-Hu Zhou , Zeng-Guang Hou , Xiao-Liang Xie , Zhen Li , Yu-Han Wang

Music information retrieval is currently an active research area that addresses the extraction of musically important information from audio signals, and the applications of such information. The extracted information can be used for search…

音频与语音处理 · 电气工程与系统科学 2022-04-08 Preeti Rao

Art has long played a profound role in shaping human emotion, cognition, and behavior. While visual arts such as painting and architecture have been studied through eye tracking, revealing distinct gaze patterns between experts and novices,…

神经元与认知 · 定量生物学 2025-12-08 Taketo Akama , Zhuohao Zhang , Tsukasa Nagashima , Takagi Yutaka , Shun Minamikawa , Natalia Polouliakh

Blind music source separation has been a popular and active subject of research in both the music information retrieval and signal processing communities. To counter the lack of available multi-track data for supervised model training, a…

音频与语音处理 · 电气工程与系统科学 2020-08-07 Ching-Yu Chiu , Wen-Yi Hsiao , Yin-Cheng Yeh , Yi-Hsuan Yang , Alvin Wen-Yu Su

Currently successful methods for video description are based on encoder-decoder sentence generation using recur-rent neural networks (RNNs). Recent work has shown the advantage of integrating temporal and/or spatial attention mechanisms…

计算机视觉与模式识别 · 计算机科学 2017-03-13 Chiori Hori , Takaaki Hori , Teng-Yok Lee , Kazuhiro Sumi , John R. Hershey , Tim K. Marks

Polyphonic Sound Event Detection (SED) in real-world recordings is a challenging task because of the dynamic polyphony level, intensity, and duration of sound events. Current polyphonic SED systems fail to model the temporal structure of…

音频与语音处理 · 电气工程与系统科学 2019-08-02 Arjun Pankajakshan , Helen L. Bear , Emmanouil Benetos

In our daily life, the scenes around us are always with multiple labels especially in a smart city, i.e., recognizing the information of city operation to response and control. Great efforts have been made by using Deep Neural Networks to…

计算机视觉与模式识别 · 计算机科学 2020-12-29 Fan Lyu , Fuyuan Hu , Victor S. Sheng , Zhengtian Wu , Qiming Fu , Baochuan Fu

Current object-centric learning models such as the popular SlotAttention architecture allow for unsupervised visual scene decomposition. Our novel MusicSlots method adapts SlotAttention to the audio domain, to achieve unsupervised music…

Automatic instrument segmentation in video is an essentially fundamental yet challenging problem for robot-assisted minimally invasive surgery. In this paper, we propose a novel framework to leverage instrument motion information, by…

计算机视觉与模式识别 · 计算机科学 2019-07-19 Yueming Jin , Keyun Cheng , Qi Dou , Pheng-Ann Heng

Attention mechanisms in biological perception are thought to select subsets of perceptual information for more sophisticated processing which would be prohibitive to perform on all sensory inputs. In computer vision, however, there has been…

计算机视觉与模式识别 · 计算机科学 2018-08-02 Mateusz Malinowski , Carl Doersch , Adam Santoro , Peter Battaglia

Improving the performance of on-device audio classification models remains a challenge given the computational limits of the mobile environment. Many studies leverage knowledge distillation to boost predictive performance by transferring…

声音 · 计算机科学 2022-02-08 Kwanghee Choi , Martin Kersner , Jacob Morton , Buru Chang

Optical Music Recognition (OMR) is concerned with transcribing sheet music into a machine-readable format. The transcribed copy should allow musicians to compose, play and edit music by taking a picture of a music sheet. Complete…

计算机视觉与模式识别 · 计算机科学 2020-06-23 Elona Shatri , György Fazekas

In the realm of music information retrieval, similarity-based retrieval and auto-tagging serve as essential components. Given the limitations and non-scalability of human supervision signals, it becomes crucial for models to learn from…

This paper proposes an attentional network for the task of Continuous Sign Language Recognition. The proposed approach exploits co-independent streams of data to model the sign language modalities. These different channels of information…

计算机视觉与模式识别 · 计算机科学 2021-01-13 Fares Ben Slimane , Mohamed Bouguessa

The potential of multimodal generative artificial intelligence (mAI) to replicate human grounded language understanding, including the pragmatic, context-rich aspects of communication, remains to be clarified. Humans are known to use…

In the age of music streaming platforms, the task of automatically tagging music audio has garnered significant attention, driving researchers to devise methods aimed at enhancing performance metrics on standard datasets. Most recent…

声音 · 计算机科学 2024-02-26 Vassilis Lyberatos , Spyridon Kantarelis , Edmund Dervakos , Giorgos Stamou

This work aims to examine one of the cornerstone problems of Musical Instrument Retrieval (MIR), in particular, instrument classification. IRMAS (Instrument recognition in Musical Audio Signals) data set is chosen for this purpose. The data…

音频与语音处理 · 电气工程与系统科学 2020-04-23 Karthikeya Racharla , Vineet Kumar , Chaudhari Bhushan Jayant , Ankit Khairkar , Paturu Harish
‹ 上一页 1 8 9 10 下一页 ›