中文
相关论文

相关论文: DCAR: A Discriminative and Compact Audio Represent…

200 篇论文

In recent years, anomaly events detection in crowd scenes attracts many researchers' attention, because of its importance to public safety. Existing methods usually exploit visual information to analyze whether any abnormal events have…

计算机视觉与模式识别 · 计算机科学 2021-10-29 Junyu Gao , Maoguo Gong , Xuelong Li

The seen birds twitter, the running cars accompany with noise, etc. These naturally audiovisual correspondences provide the possibilities to explore and understand the outside world. However, the mixed multiple objects and sounds make it…

计算机视觉与模式识别 · 计算机科学 2019-04-22 Di Hu , Feiping Nie , Xuelong Li

Speaker Identification process is to identify a particular vocal cord from a set of existing speakers. In the speaker identification processes, unknown speaker voice sample targets each of the existing speakers present in the system and…

声音 · 计算机科学 2017-04-14 Soumen Kanrar

We introduce DECAR, a self-supervised pre-training approach for learning general-purpose audio representations. Our system is based on clustering: it utilizes an offline clustering step to provide target labels that act as pseudo-labels for…

声音 · 计算机科学 2023-03-15 Sreyan Ghosh , Sandesh V Katta , Ashish Seth , S. Umesh

Accurate interpretation of electrocardiogram (ECG) signals is crucial for diagnosing cardiovascular diseases. Recent multimodal approaches that integrate ECGs with accompanying clinical reports show strong potential, but they still face two…

人工智能 · 计算机科学 2026-02-25 Ziwei Niu , Hao Sun , Shujun Bian , Xihong Yang , Lanfen Lin , Yuxin Liu , Yueming Jin

Although acoustic scenes and events include many related tasks, their combined detection and classification have been scarcely investigated. We propose three architectures of deep neural networks that are integrated to simultaneously…

音频与语音处理 · 电气工程与系统科学 2021-02-09 Jee-weon Jung , Hye-jin Shim , Ju-ho Kim , Ha-Jin Yu

Discriminatively localizing sounding objects in cocktail-party, i.e., mixed sound scenes, is commonplace for humans, but still challenging for machines. In this paper, we propose a two-stage learning framework to perform self-supervised…

计算机视觉与模式识别 · 计算机科学 2020-10-13 Di Hu , Rui Qian , Minyue Jiang , Xiao Tan , Shilei Wen , Errui Ding , Weiyao Lin , Dejing Dou

Weakly supervised video anomaly detection (WS-VAD) is a crucial area in computer vision for developing intelligent surveillance systems. This system uses three feature streams: RGB video, optical flow, and audio signals, where each stream…

计算机视觉与模式识别 · 计算机科学 2025-04-07 Yuta Kaneko , Abu Saleh Musa Miah , Najmul Hassan , Hyoun-Sup Lee , Si-Woong Jang , Jungpil Shin

Most existing audio-text retrieval (ATR) approaches typically rely on a single-level interaction to associate audio and text, limiting their ability to align different modalities and leading to suboptimal matches. In this work, we present a…

声音 · 计算机科学 2025-05-06 Yifei Xin , Zhihong Zhu , Xuxin Cheng , Xusheng Yang , Yuexian Zou

Automatic audio event recognition plays a pivotal role in making human robot interaction more closer and has a wide applicability in industrial automation, control and surveillance systems. Audio event is composed of intricate phonic…

音频与语音处理 · 电气工程与系统科学 2023-04-12 Tushar Sandhan , Sukanya Sonowal , Jin Young Choi

Deep learning (DL)-based methods have recently shown great promise in bitemporal change detection (CD). Existing discriminative methods based on Convolutional Neural Networks (CNNs) and Transformers rely on discriminative representation…

计算机视觉与模式识别 · 计算机科学 2024-03-05 Yihan Wen , Xianping Ma , Xiaokang Zhang , Man-On Pun

In the field of music information retrieval (MIR), cover song identification (CSI) is a challenging task that aims to identify cover versions of a query song from a massive collection. Existing works still suffer from high intra-song…

信息检索 · 计算机科学 2023-07-20 Jiahao Xun , Shengyu Zhang , Yanting Yang , Jieming Zhu , Liqun Deng , Zhou Zhao , Zhenhua Dong , Ruiqi Li , Lichao Zhang , Fei Wu

Machine hearing of the environmental sound is one of the important issues in the audio recognition domain. It gives the machine the ability to discriminate between the different input sounds that guides its decision making. In this work we…

声音 · 计算机科学 2022-07-20 Peter Ochieng , Dennis Kaburu

Event-based pedestrian attribute recognition (PAR) leverages motion cues to enhance RGB cameras in low-light and motion-blur scenarios, enabling more accurate inference of attributes like age and emotion. However, existing two-stream…

计算机视觉与模式识别 · 计算机科学 2026-03-23 Minghe Xu , Rouying Wu , ChiaWei Chu , Xiao Wang , Yu Li

Acoustic event detection for content analysis in most cases relies on lots of labeled data. However, manually annotating data is a time-consuming task, which thus makes few annotated resources available so far. Unlike audio event detection,…

计算机视觉与模式识别 · 计算机科学 2016-08-16 Yong Xu , Qiang Huang , Wenwu Wang , Philip J. B. Jackson , Mark D. Plumbley

Singing voice separation attempts to separate the vocal and instrumental parts of a music recording, which is a fundamental problem in music information retrieval. Recent work on singing voice separation has shown that the low-rank…

音频与语音处理 · 电气工程与系统科学 2018-01-12 Tak-Shing T. Chan , Yi-Hsuan Yang

Anomalous audio in speech recordings is often caused by speaker voice distortion, external noise, or even electric interferences. These obstacles have become a serious problem in some fields, such as high-quality music mixing and speech…

音频与语音处理 · 电气工程与系统科学 2021-02-11 Qiang Huang , Thomas Hain

Recently, the application of diffusion probabilistic models has advanced speech enhancement through generative approaches. However, existing diffusion-based methods have focused on the generation process in high-dimensional waveform or…

声音 · 计算机科学 2025-01-20 Shengkui Zhao , Zexu Pan , Kun Zhou , Yukun Ma , Chong Zhang , Bin Ma

Significant improvement has been achieved in automated audio captioning (AAC) with recent models. However, these models have become increasingly large as their performance is enhanced. In this work, we propose a knowledge distillation (KD)…

声音 · 计算机科学 2024-07-22 Xuenan Xu , Haohe Liu , Mengyue Wu , Wenwu Wang , Mark D. Plumbley

Audio captioning is an important research area that aims to generate meaningful descriptions for audio clips. Most of the existing research extracts acoustic features of audio clips as input to encoder-decoder and transformer architectures…

声音 · 计算机科学 2022-04-20 Ayşegül Özkaya Eren , Mustafa Sert