English
Related papers

Related papers: Full-Frequency Temporal Patching and Structured Ma…

200 papers

Environmental Sound Classification (ESC) is a rapidly evolving field that recently demonstrated the advantages of application of visual domain techniques to the audio-related tasks. Previous studies indicate that the domain-specific…

Sound · Computer Science 2021-04-26 Andrey Guzhov , Federico Raue , Jörn Hees , Andreas Dengel

Automatic pronunciation assessment (APA) analyzes second-language (L2) learners' speech by providing fine-grained pronunciation feedback at various linguistic levels. Most existing efforts on APA typically adopt segmental-level features as…

Computation and Language · Computer Science 2025-09-23 Jiun-Ting Li , Bi-Cheng Yan , Yi-Cheng Wang , Berlin Chen

Most of the current speech data augmentation methods operate on either the raw waveform or the amplitude spectrum of speech. In this paper, we propose a novel speech data augmentation method called PhasePerturbation that operates…

Sound · Computer Science 2023-12-15 Chengxi Lei , Satwinder Singh , Feng Hou , Xiaoyun Jia , Ruili Wang

Multimodal time series forecasting is crucial in real-world applications, where decisions depend on both numerical data and contextual signals. The core challenge is to effectively combine temporal numerical patterns with the context…

Machine Learning · Computer Science 2026-02-04 Huu Hiep Nguyen , Minh Hoang Nguyen , Dung Nguyen , Hung Le

The audio spectrogram is a time-frequency representation that has been widely used for audio classification. One of the key attributes of the audio spectrogram is the temporal resolution, which depends on the hop size used in the Short-Time…

Sound · Computer Science 2024-01-15 Haohe Liu , Xubo Liu , Qiuqiang Kong , Wenwu Wang , Mark D. Plumbley

Meta-learning facilitates few-shot hyperspectral target detection (HTD), but adapting deep backbones remains challenging. Full-parameter fine-tuning is inefficient and prone to overfitting, and existing methods largely ignore the…

Computer Vision and Pattern Recognition · Computer Science 2026-04-08 Luqi Gong , Qixin Xie , Yue Chen , Ziqiang Chen , Fanda Fan , Shuai Zhao , Chao Li

The core challenge in industrial equipment anoma lous sound detection (ASD) lies in modeling the time-frequency coupling characteristics of acoustic features. Existing modeling methods are limited by local receptive fields, making it…

Sound · Computer Science 2025-09-03 Chengyuan Ma , Peng Jia , Hongyue Guo , Wenming Yang

Voice spoofing attacks pose a significant threat to automated speaker verification systems. Existing anti-spoofing methods often simulate specific attack types, such as synthetic or replay attacks. However, in real-world scenarios, the…

Sound · Computer Science 2023-09-19 Awais Khan , Khalid Mahmood Malik , Shah Nawaz

Recently, masked image modeling (MIM), which learns visual representations by reconstructing the masked patches of an image, has dominated self-supervised learning in computer vision. However, the pre-training of MIM always takes massive…

Computer Vision and Pattern Recognition · Computer Science 2024-10-01 Jie Gui , Tuo Chen , Minjing Dong , Zhengqi Liu , Hao Luo , James Tin-Yau Kwok , Yuan Yan Tang

Marine Animal Segmentation (MAS) aims at identifying and segmenting marine animals from complex marine environments. Most of previous deep learning-based MAS methods struggle with the long-distance modeling issue. Recently, Segment Anything…

Computer Vision and Pattern Recognition · Computer Science 2026-03-16 Pingping Zhang , Tianyu Yan , Yuhao Wang , Yang Liu , Tongdan Tang , Yili Ma , Long Lv , Feng Tian , Weibing Sun , and Huchuan Lu

We propose FSB-LSTM, a novel long short-term memory (LSTM) based architecture that integrates full- and sub-band (FSB) modeling, for single- and multi-channel speech enhancement in the short-time Fourier transform (STFT) domain. The model…

Audio and Speech Processing · Electrical Eng. & Systems 2023-04-19 Zhong-Qiu Wang , Samuele Cornell , Shukjae Choi , Younglo Lee , Byeong-Yeol Kim , Shinji Watanabe

Acoustic environments affect acoustic characteristics of sound to be recognized by physically interacting with sound wave propagation. Thus, training acoustic models for audio and speech tasks requires regularization on various acoustic…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-08 Hyeonuk Nam , Seong-Hu Kim , Yong-Hwa Park

We propose a method using a long short-term memory (LSTM) network to estimate the noise power spectral density (PSD) of single-channel audio signals represented in the short time Fourier transform (STFT) domain. An LSTM network common to…

Signal Processing · Electrical Eng. & Systems 2020-11-11 Xiaofei Li , Simon Leglaive , Laurent Girin , Radu Horaud

This technical report describes Chung-Ang University and Korea University (CAU_KU) team's model participating in the Audio Deep Synthesis Detection (ADD) 2022 Challenge, track 1: Low-quality fake audio detection. For track 1, we propose a…

Sound · Computer Science 2022-02-10 Il-Youp Kwak , Sunmook Choi , Jonghoon Yang , Yerin Lee , Seungsang Oh

Photomultiplier tubes (PMTs) are widely deployed at neutrino and dark matter experiments for photon counting. When multiple photons hit a PMT consecutively, their photo-electron (PE) pulses pile up to hinder the precise measurements of the…

High Energy Physics - Experiment · Physics 2025-09-12 Yuyi Wang , Aiqiang Zhang , Yiyang Wu , Benda Xu , Xuewei Liu , Jiajie Chen , Zhe Wang , Shaomin Chen

Affine frequency division multiplexing (AFDM) has emerged as a promising modulation scheme for doubly selective channels, but its canonical continuous-time realization, referred to herein as piecewise continuous AFDM (PC-AFDM), has been…

Signal Processing · Electrical Eng. & Systems 2026-05-13 Yewen Cao , Yulin Shao

Time masking has become a de facto augmentation technique for speech and audio tasks, including automatic speech recognition (ASR) and audio classification, most notably as a part of SpecAugment. In this work, we propose SpliceOut, a simple…

Sound · Computer Science 2021-10-14 Arjit Jain , Pranay Reddy Samala , Deepak Mittal , Preethi Jyoti , Maneesh Singh

Previously proposed FullSubNet has achieved outstanding performance in Deep Noise Suppression (DNS) Challenge and attracted much attention. However, it still encounters issues such as input-output mismatch and coarse processing for…

Sound · Computer Science 2022-03-29 Jun Chen , Zilin Wang , Deyi Tuo , Zhiyong Wu , Shiyin Kang , Helen Meng

Recent advances in audio large language models (ALLMs) have made high-quality synthetic audio widely accessible, increasing the risk of malicious audio deepfakes across speech, environmental sounds, singing voice, and music. Real-world…

Sound · Computer Science 2026-01-07 Yuankun Xie , Xiaoxuan Guo , Jiayi Zhou , Tao Wang , Jian Liu , Ruibo Fu , Xiaopeng Wang , Haonan Cheng , Long Ye

Audio classification models, particularly the Audio Spectrogram Transformer (AST), play a crucial role in efficient audio analysis. However, optimizing their efficiency without compromising accuracy remains a challenge. In this paper, we…

Sound · Computer Science 2024-06-13 Swarup Ranjan Behera , Abhishek Dhiman , Karthik Gowda , Aalekhya Satya Narayani