中文
相关论文

相关论文: Multi-Scale Time-Frequency Attention for Acoustic …

200 篇论文

Auditory Attention Decoding (AAD) can help to determine the identity of the attended speaker during an auditory selective attention task, by analyzing and processing measurements of electroencephalography (EEG) data. Most studies on AAD are…

信号处理 · 电气工程与系统科学 2024-09-16 Haolin Zhu , Yujie Yan , Xiran Xu , Zhongshu Ge , Pei Tian , Xihong Wu , Jing Chen

Acoustic Scene Classification (ASC) is a challenging task, as a single scene may involve multiple events that contain complex sound patterns. For example, a cooking scene may contain several sound sources including silverware clinking,…

音频与语音处理 · 电气工程与系统科学 2019-09-20 Weimin Wang , Weiran Wang , Ming Sun , Chao Wang

We study few-shot acoustic event detection (AED) in this paper. Few-shot learning enables detection of new events with very limited labeled data. Compared to other research areas like computer vision, few-shot learning for audio recognition…

机器学习 · 计算机科学 2020-02-24 Bowen Shi , Ming Sun , Krishna C. Puvvada , Chieh-Chi Kao , Spyros Matsoukas , Chao Wang

Acoustic event detection and scene classification are major research tasks in environmental sound analysis, and many methods based on neural networks have been proposed. Conventional methods have addressed these tasks separately; however,…

Environmental sound analysis is currently getting more and more attentions. In the domain, acoustic scene classification and acoustic event classification are two closely related tasks. In this letter, a two-stage method is proposed for the…

声音 · 计算机科学 2021-03-31 Weiping Zheng , Dacan Jiang , Gansen Zhao

Unsupervised multivariate time series anomaly detection (UMTSAD) plays a critical role in various domains, including finance, networks, and sensor systems. In recent years, due to the outstanding performance of deep learning in general…

机器学习 · 计算机科学 2025-04-28 Tiange Huang , Yongjun Li

Attention has become one of the most commonly used mechanisms in deep learning approaches. The attention mechanism can help the system focus more on the feature space's critical regions. For example, high amplitude regions can play an…

声音 · 计算机科学 2022-08-24 Junghun Kim , Yoojin An , Jihie Kim

Acoustic scene classification (ASC) and sound event detection (SED) are major topics in environmental sound analysis. Considering that acoustic scenes and sound events are closely related to each other, the joint analysis of acoustic scenes…

声音 · 计算机科学 2022-06-22 Kayo Nada , Keisuke Imoto , Takao Tsuchiya

Various attention mechanisms are being widely applied to acoustic scene classification. However, we empirically found that the attention mechanism can excessively discard potentially valuable information, despite improving performance. We…

机器学习 · 计算机科学 2021-12-24 Hye-jin Shim , Jee-weon Jung , Ju-ho Kim , Ha-Jin Yu

Cross-view geo-localization aims to determine the geographical location of a query image by matching it against a gallery of images. This task is challenging due to the significant appearance variations of objects observed from variable…

计算机视觉与模式识别 · 计算机科学 2025-09-17 YiTong Liu , TianZhu Liu , YanFeng GU

In this paper, we propose a novel bidirectional multiscale feature aggregation (BMFA) network with attentional fusion modules for text-independent speaker verification. The feature maps from different stages of the backbone network are…

音频与语音处理 · 电气工程与系统科学 2021-04-02 Jiajun Qi , Wu Guo , Bin Gu

Sound event detection (SED) aims to detect when and recognize what sound events happen in an audio clip. Many supervised SED algorithms rely on strongly labelled data which contains the onset and offset annotations of sound events. However,…

声音 · 计算机科学 2019-12-11 Qiuqiang Kong , Yong Xu , Iwona Sobieraj , Wenwu Wang , Mark D. Plumbley

A major advantage of a deep convolutional neural network (CNN) is that the focused receptive field size is increased by stacking multiple convolutional layers. Accordingly, the model can explore the long-range dependency of features from…

声音 · 计算机科学 2020-06-17 Xugang Lu , Peng Shen , Sheng Li , Yu Tsao , Hisashi Kawai

In this study, we introduce a convolutional time-frequency-channel "Squeeze and Excitation" (tfc-SE) module to explicitly model inter-dependencies between the time-frequency domain and multiple channels. The tfc-SE module consists of two…

音频与语音处理 · 电气工程与系统科学 2019-08-06 Wei Xia , Kazuhito Koishida

Different machines can exhibit diverse frequency patterns in their emitted sound. This feature has been recently explored in anomaly sound detection and reached state-of-the-art performance. However, existing methods rely on the manual or…

声音 · 计算机科学 2023-09-07 Hejing Zhang , Jian Guan , Qiaoxi Zhu , Feiyang Xiao , Youde Liu

Existing methods for deepfake audio detection have demonstrated some effectiveness. However, they still face challenges in generalizing to new forgery techniques and evolving attack patterns. This limitation mainly arises because the models…

The rise of highly convincing synthetic speech poses a growing threat to audio communications. Although existing Audio Deepfake Detection (ADD) methods have demonstrated good performance under clean conditions, their effectiveness drops…

音频与语音处理 · 电气工程与系统科学 2025-08-05 Haohan Shi , Xiyu Shi , Safak Dogan , Tianjin Huang , Yunxiao Zhang

Speaker verification based on ad-hoc microphone arrays has the potential of reducing the error significantly in adverse acoustic environments. However, existing approaches extract utterance-level speaker embeddings from each channel of an…

声音 · 计算机科学 2022-03-29 Chengdong Liang , Yijiang Chen , Jiadi Yao , Xiao-Lei Zhang

Detection of anomaly events is relevant for public safety and requires a combination of fine-grained motion information and contextual events at variable time-scales. To this end, we propose a Multi-Timescale Feature Learning (MTFL) method…

计算机视觉与模式识别 · 计算机科学 2025-10-29 Yiling Zhang , Erkut Akdag , Egor Bondarev , Peter H. N. De With

Recognizing and localizing events in videos is a fundamental task for video understanding. Since events may occur in auditory and visual modalities, multimodal detailed perception is essential for complete scene comprehension. Most previous…

计算机视觉与模式识别 · 计算机科学 2022-07-13 Jiashuo Yu , Ying Cheng , Rui-Wei Zhao , Rui Feng , Yuejie Zhang