中文
相关论文

相关论文: Simple Pooling Front-ends For Efficient Audio Clas…

200 篇论文

Training-free anomalous sound detection (ASD) based on pre-trained audio embedding models has recently garnered significant attention, as it enables the detection of anomalous sounds using only normal reference data while offering improved…

音频与语音处理 · 电气工程与系统科学 2026-03-06 Kevin Wilkinghoff , Sarthak Yadav , Zheng-Hua Tan

While end-to-end systems are becoming popular in auditory signal processing including automatic music tagging, models using raw audio as input needs a large amount of data and computational resources without domain knowledge. Inspired by…

音频与语音处理 · 电气工程与系统科学 2022-11-29 Yinghao Ma , Richard M. Stern

Existing studies tend tofocus onmodel modifications and integration with higher accuracy, which improve performance but also carry huge computational costs, resulting in longer detection times. Inmedical imaging, the use of time is…

图像与视频处理 · 电气工程与系统科学 2023-02-22 Weihu Song , Heng Yu

Convolutional Neural Networks (CNNs) use pooling to decrease the size of activation maps. This process is crucial to increase the receptive fields and to reduce computational requirements of subsequent convolutions. An important feature of…

计算机视觉与模式识别 · 计算机科学 2021-03-19 Alexandros Stergiou , Ronald Poppe , Grigorios Kalliatakis

Mel-scale spectrum features are used in various recognition and classification tasks on speech signals. There is no reason to expect that these features are optimal for all different tasks, including speaker verification (SV). This paper…

音频与语音处理 · 电气工程与系统科学 2022-06-16 Jingyu Li , Yusheng Tian , Tan Lee

Large Audio Language Models (LALMs) demonstrate impressive performance across diverse tasks, ranging from speech recognition to general audio understanding. However, their scalability is limited by the quadratic complexity of attention and…

音频与语音处理 · 电气工程与系统科学 2025-11-27 Saurabhchand Bhati , Samuel Thomas , Hilde Kuehne , Rogerio Feris , James Glass

There exists a plethora of techniques for inducing structured sparsity in parametric models during the optimization process, with the final goal of resource-efficient inference. However, few methods target a specific number of…

机器学习 · 计算机科学 2018-11-26 Raphael Tang , Ashutosh Adhikari , Jimmy Lin

Most audio processing pipelines involve transformations that act on fixed-dimensional input representations of audio. For example, when using the Short Time Fourier Transform (STFT) the DFT size specifies a fixed dimension for the input…

音频与语音处理 · 电气工程与系统科学 2022-03-28 Krishna Subramani , Paris Smaragdis

Anti-spoofing is the task of speech authentication. That is, identifying genuine human speech compared to spoofed speech. The main focus of this paper is to suggest new representations for genuine and spoofed speech, based on the…

音频与语音处理 · 电气工程与系统科学 2022-10-28 Matan Karo , Arie Yeredor , Itshak Lapidot

We study the problem of stereo singing voice cancellation, a subtask of music source separation, whose goal is to estimate an instrumental background from a stereo mix. We explore how to achieve performance similar to large state-of-the-art…

声音 · 计算机科学 2024-01-23 Clara Borrelli , James Rae , Dogac Basaran , Matt McVicar , Mehrez Souden , Matthias Mauch

In recent years, exploring effective sound separation (SSep) techniques to improve overlapping sound event detection (SED) attracts more and more attention. Creating accurate separation signals to avoid the catastrophic error accumulation…

音频与语音处理 · 电气工程与系统科学 2022-03-07 Yunhao Liang , Yanhua Long , Yijie Li , Jiaen Liang

Sounds carry an abundance of information about activities and events in our everyday environment, such as traffic noise, road works, music, or people talking. Recent machine learning methods, such as convolutional neural networks (CNNs),…

声音 · 计算机科学 2023-05-31 Arshdeep Singh , Haohe Liu , Mark D. Plumbley

Wildlife conservation using continuous monitoring of environmental factors and biomedical classification, which generate a vast amount of sensor data, is a challenge due to limited bandwidth in the case of remote monitoring. It becomes…

机器学习 · 计算机科学 2023-04-25 Abhishek Ramdas Nair , Pallab Kumar Nath , Shantanu Chakrabartty , Chetan Singh Thakur

Leading- and trailing-edge serrations have been widely used to reduce the leading- and trailing-edge noise in applications such as contra-rotating fans and large wind turbines. Recent studies show that these two noise problems can be…

流体动力学 · 物理学 2020-01-29 Benshuai Lyu , Lorna J. Ayton

Deep learning has celebrated resounding successes in many application areas of relevance to the Internet of Things (IoT), such as computer vision and machine listening. These technologies must ultimately be brought directly to the edge to…

声音 · 计算机科学 2022-01-19 Md Mohaimenuzzaman , Christoph Bergmeir , Bernd Meyer

Split-Federated (SplitFed) learning is an extension of federated learning that places minimal requirements on the clients computing infrastructure, since only a small portion of the overall model is deployed on the clients hardware. In…

计算机视觉与模式识别 · 计算机科学 2026-01-28 Zahra Hafezi Kafshgari , Ivan V. Bajic , Parvaneh Saeedi

Audio-Language Models (ALMs) have recently achieved remarkable success in zero-shot audio recognition tasks, which match features of audio waveforms with class-specific text prompt features, inspired by advancements in Vision-Language…

声音 · 计算机科学 2024-10-01 Asif Hanif , Maha Tufail Agro , Mohammad Areeb Qazi , Hanan Aldarmaki

Environmental sound classification (ESC) is a challenging problem due to the unstructured spatial-temporal relations that exist in the sound signals. Recently, many studies have focused on abstracting features from convolutional neural…

声音 · 计算机科学 2022-05-31 Liguang Zhou , Yuhongze Zhou , Xiaonan Qi , Junjie Hu , Tin Lun Lam , Yangsheng Xu

Environmental Sound Classification is an important problem of sound recognition and is more complicated than speech recognition problems as environmental sounds are not well structured with respect to time and frequency. Researchers have…

声音 · 计算机科学 2024-08-27 Aditya Dawn , Wazib Ansar

Although probing frozen models has become a standard evaluation paradigm, self-supervised learning in audio defaults to fine-tuning when pursuing state-of-the-art on AudioSet. A key reason is that global pooling creates an information…