中文
相关论文

相关论文: Frequency-Temporal Attention Network for Singing M…

200 篇论文

Music generative artificial intelligence (AI) is rapidly expanding music content, necessitating automated song aesthetics evaluation. However, existing studies largely focus on speech, audio or singing quality, leaving song aesthetics…

声音 · 计算机科学 2026-01-21 Yishan Lv , Jing Luo , Boyuan Ju , Yang Zhang , Xinda Wu , Bo Yuan , Xinyu Yang

This study introduces a biologically-inspired model designed to examine the role of coincidence detection cells in speech segregation tasks. The model consists of three stages: a time-domain cochlear model that generates instantaneous rates…

神经元与认知 · 定量生物学 2024-05-13 Asaf Zorea , Miriam Furst

Patterns of brain activity are associated with different brain processes and can be used to identify different brain states and make behavioral predictions. However, the relevant features are not readily apparent and accessible. To mine…

Visual speech recognition is the task to decode the speech content from a video based on visual information, especially the movements of lips. It is also referenced as lipreading. Motivated by two problems existing in lipreading, words with…

计算机视觉与模式识别 · 计算机科学 2019-01-14 Jingyun Xiao

Chord recognition systems depend on robust feature extraction pipelines. While these pipelines are traditionally hand-crafted, recent advances in end-to-end machine learning have begun to inspire researchers to explore data-driven methods…

机器学习 · 计算机科学 2016-12-16 Filip Korzeniowski , Gerhard Widmer

Many applications of cross-modal music retrieval are related to connecting sheet music images to audio recordings. A typical and recent approach to this is to learn, via deep neural networks, a joint embedding space that correlates short…

声音 · 计算机科学 2023-09-22 Luis Carvalho , Gerhard Widmer

Auditory attention detection (AAD) aims to identify the direction of the attended speaker in multi-speaker environments from brain signals, such as Electroencephalography (EEG) signals. However, existing EEG-based AAD methods overlook the…

人机交互 · 计算机科学 2025-05-16 Cunhang Fan , Xiaoke Yang , Hongyu Zhang , Ying Chen , Lu Li , Jian Zhou , Zhao Lv

User interests are usually dynamic in the real world, which poses both theoretical and practical challenges for learning accurate preferences from rich behavior data. Among existing user behavior modeling solutions, attention networks are…

信息检索 · 计算机科学 2022-04-14 Chao Chen , Haoyu Geng , Nianzu Yang , Junchi Yan , Daiyue Xue , Jianping Yu , Xiaokang Yang

Music listening involves many simultaneous neural operations, including auditory processing, working memory, temporal sequencing, pitch tracking, anticipation, reward, and emotion, and thus, a full investigation of music cognition would…

神经元与认知 · 定量生物学 2022-01-19 Melia E. Bonomo , Anthony K. Brandt , J. Todd Frazier , Christof Karmonik

Audio-visual multi-modal modeling has been demonstrated to be effective in many speech related tasks, such as speech recognition and speech enhancement. This paper introduces a new time-domain audio-visual architecture for target speaker…

音频与语音处理 · 电气工程与系统科学 2019-09-24 Jian Wu , Yong Xu , Shi-Xiong Zhang , Lian-Wu Chen , Meng Yu , Lei Xie , Dong Yu

We present a supervised neural network model for polyphonic piano music transcription. The architecture of the proposed model is analogous to speech recognition systems and comprises an acoustic model and a music language model. The…

机器学习 · 统计学 2016-02-12 Siddharth Sigtia , Emmanouil Benetos , Simon Dixon

When songs are composed or performed, there is often an intent by the singer/songwriter of expressing feelings or emotions through it. For humans, matching the emotiveness in a musical composition or performance with the subjective…

In multichannel speech enhancement, both spectral and spatial information are vital for discriminating between speech and noise. How to fully exploit these two types of information and their temporal dynamics remains an interesting research…

音频与语音处理 · 电气工程与系统科学 2022-11-17 Yujie Yang , Changsheng Quan , Xiaofei Li

Motivated by the attention mechanism of the human visual system and recent developments in the field of machine translation, we introduce our attention-based and recurrent sequence to sequence autoencoders for fully unsupervised…

音频与语音处理 · 电气工程与系统科学 2020-05-20 Shahin Amiriparian , Pawel Winokurow , Vincent Karas , Sandra Ottl , Maurice Gerczuk , Björn W. Schuller

Assessment of voice signals has long been performed with the assumption of periodicity as this facilitates analysis. Near periodicity of normal voice signals makes short-time harmonic modeling an appealing choice to extract vocal feature…

音频与语音处理 · 电气工程与系统科学 2022-02-10 Takeshi Ikuma , Andrew J. McWhorter , Lacey Adkins , Melda Kunduk

In the field of music information retrieval, the task of simultaneously identifying the presence or absence of multiple musical instruments in a polyphonic recording remains a hard problem. Previous works have seen some success in improving…

音频与语音处理 · 电气工程与系统科学 2020-06-23 Karn Watcharasupat , Siddharth Gururani , Alexander Lerch

Attentive listening in a multispeaker environment such as a cocktail party requires suppression of the interfering speakers and the noise around. People with normal hearing perform remarkably well in such situations. Analysis of the…

音频与语音处理 · 电气工程与系统科学 2021-02-04 Ivine Kuruvila , Kubilay Can Demir , Eghart Fischer , Ulrich Hoppe

Music information retrieval is currently an active research area that addresses the extraction of musically important information from audio signals, and the applications of such information. The extracted information can be used for search…

音频与语音处理 · 电气工程与系统科学 2022-04-08 Preeti Rao

Weakly supervised video anomaly detection (WS-VAD) is a crucial area in computer vision for developing intelligent surveillance systems. This system uses three feature streams: RGB video, optical flow, and audio signals, where each stream…

计算机视觉与模式识别 · 计算机科学 2025-04-07 Yuta Kaneko , Abu Saleh Musa Miah , Najmul Hassan , Hyoun-Sup Lee , Si-Woong Jang , Jungpil Shin

Large-scale sound recognition data sets typically consist of acoustic recordings obtained from multimedia libraries. As a consequence, modalities other than audio can often be exploited to improve the outputs of models designed for…

音频与语音处理 · 电气工程与系统科学 2022-10-11 Wim Boes , Hugo Van hamme