English
Related papers

Related papers: Multi-dimensional frequency dynamic convolution wi…

200 papers

As speech-interfaces are getting richer and widespread, speech emotion recognition promises more attractive applications. In the continuous emotion recognition (CER) problem, tracking changes across affective states is an important and…

Sound · Computer Science 2021-10-11 Berkay Kopru , Engin Erzin

Recent advances in speech deepfake detection (SDD) have significantly improved artifacts-based detection in spoofed speech. However, most models overlook speech naturalness, a crucial cue for distinguishing bona fide speech from spoofed…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-21 Taewoo Kim , Guisik Kim , Choongsang Cho , Young Han Lee

Due to the fast inference and good performance, discriminative learning methods have been widely studied in image denoising. However, these methods mostly learn a specific model for each noise level, and require multiple models for…

Computer Vision and Pattern Recognition · Computer Science 2018-08-01 Kai Zhang , Wangmeng Zuo , Lei Zhang

Event analysis in untrimmed videos has attracted increasing attention due to the application of cutting-edge techniques such as CNN. As a well studied property for CNN-based models, the receptive field is a measurement for measuring the…

Computer Vision and Pattern Recognition · Computer Science 2021-11-24 Zhaobo Qi , Shuhui Wang , Chi Su , Li Su , Weigang Zhang , Qingming Huang

In this paper, we present a method called HODGEPODGE\footnotemark[1] for large-scale detection of sound events using weakly labeled, synthetic, and unlabeled data proposed in the Detection and Classification of Acoustic Scenes and Events…

Sound · Computer Science 2019-07-18 Ziqiang Shi , Liu Liu , Huibin Lin , Rujie Liu , Anyan Shi

Event classification is inherently sequential and multimodal. Therefore, deep neural models need to dynamically focus on the most relevant time window and/or modality of a video. In this study, we propose the Multi-level Attention Fusion…

Computer Vision and Pattern Recognition · Computer Science 2021-06-15 Mathilde Brousmiche , Jean Rouat , Stéphane Dupont

Convolutional frontends are a typical choice for Transformer-based automatic speech recognition to preprocess the spectrogram, reduce its sequence length, and combine local information in time and frequency similarly. However, the width and…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-13 Belen Alastruey , Lukas Drude , Jahn Heymann , Simon Wiesler

Building extraction $-$ needed for inventory management and planning of urban environment $-$ is affected by the misalignment between labels and off-nadir source imagery in training data. Teacher-Student learning of noise-tolerant…

Computer Vision and Pattern Recognition · Computer Science 2024-10-28 Bipul Neupane , Jagannath Aryal , Abbas Rajabifard

Deep neural networks face many problems in the field of hyperspectral image classification, lack of effective utilization of spatial spectral information, gradient disappearance and overfitting as the model depth increases. In order to…

Computer Vision and Pattern Recognition · Computer Science 2023-07-14 Guandong Li

In this study, we focus on automated approaches to detect depression from clinical interviews using multi-modal machine learning (ML). Our approach differentiates from other successful ML methods such as context-aware analysis through…

Machine Learning · Computer Science 2024-12-30 Genevieve Lam , Huang Dongyan , Weisi Lin

Sound event localization and detection (SELD) is a task for the classification of sound events and the identification of direction of arrival (DoA) utilizing multichannel acoustic signals. For effective classification and localization, a…

Audio and Speech Processing · Electrical Eng. & Systems 2025-04-18 Yusun Shul , Dayun Choi , Jung-Woo Choi

Speech enhancement in the time domain is becoming increasingly popular in recent years, due to its capability to jointly enhance both the magnitude and the phase of speech. In this work, we propose a dense convolutional network (DCN) with…

Audio and Speech Processing · Electrical Eng. & Systems 2021-03-09 Ashutosh Pandey , DeLiang Wang

Robust face detection is one of the most important pre-processing steps to support facial expression analysis, facial landmarking, face recognition, pose estimation, building of 3D facial models, etc. Although this topic has been intensely…

Computer Vision and Pattern Recognition · Computer Science 2017-01-03 Yutong Zheng , Chenchen Zhu , Khoa Luu , Chandrasekhar Bhagavatula , T. Hoang Ngan Le , Marios Savvides

Speech Emotion Recognition (SER) systems often degrade in performance when exposed to the unpredictable acoustic interference found in real-world environments. Additionally, the opacity of deep learning models hinders their adoption in…

Sound · Computer Science 2025-12-23 Sudip Chakrabarty , Pappu Bishwas , Rajdeep Chatterjee

In this study we present a kernel based convolution model to characterize neural responses to natural sounds by decoding their time-varying acoustic features. The model allows to decode natural sounds from high-dimensional neural…

Machine Learning · Statistics 2016-11-15 Ali Faisal , Anni Nora , Jaeho Seol , Hanna Renvall , Riitta Salmelin

Complex spatial dependencies in transportation networks make traffic prediction extremely challenging. Much existing work is devoted to learning dynamic graph structures among sensors, and the strategy of mining spatial dependencies from…

Machine Learning · Computer Science 2023-12-20 Yujie Li , Zezhi Shao , Yongjun Xu , Qiang Qiu , Zhaogang Cao , Fei Wang

We present a new framework SoundDet, which is an end-to-end trainable and light-weight framework, for polyphonic moving sound event detection and localization. Prior methods typically approach this problem by preprocessing raw waveform into…

Sound · Computer Science 2021-08-24 Yuhang He , Niki Trigoni , Andrew Markham

Multiple description coding (MDC) is able to stably transmit the signal in the un-reliable and non-prioritized networks, which has been broadly studied for several decades. However, the traditional MDC doesn't well leverage image's context…

Multimedia · Computer Science 2019-03-01 Lijun Zhao , Huihui Bai , Anhong Wang , Yao Zhao

These recent years have witnessed that convolutional neural network (CNN)-based methods for detecting infrared small targets have achieved outstanding performance. However, these methods typically employ standard convolutions, neglecting to…

Computer Vision and Pattern Recognition · Computer Science 2024-12-24 Jiangnan Yang , Shuangli Liu , Jingjun Wu , Xinyu Su , Nan Hai , Xueli Huang

Deep convolutional neural networks are being actively investigated in a wide range of speech and audio processing applications including speech recognition, audio event detection and computational paralinguistics, owing to their ability to…

Machine Learning · Computer Science 2018-01-16 Che-Wei Huang , Shrikanth. S. Narayanan
‹ Prev 1 8 9 10 Next ›