English
Related papers

Related papers: Hybrid-Sep: Language-queried audio source separati…

200 papers

Sound event detection is an important facet of audio tagging that aims to identify sounds of interest and define both the sound category and time boundaries for each sound event in a continuous recording. With advances in deep neural…

Sound · Computer Science 2024-12-31 Sangwook Park , David K. Han , Mounya Elhilali

Most audio-visual speaker extraction methods rely on synchronized lip recording to isolate the speech of a target speaker from a multi-talker mixture. However, in natural human communication, co-speech gestures are also temporally aligned…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-28 Zexu Pan , Xinyuan Qian , Shengkui Zhao , Kun Zhou , Bin Ma

Large-scale pre-trained image-text models demonstrate remarkable versatility across diverse tasks, benefiting from their robust representational capabilities and effective multimodal alignment. We extend the application of these models,…

Computer Vision and Pattern Recognition · Computer Science 2023-11-08 Sooyoung Park , Arda Senocak , Joon Son Chung

The objective of this work is to extract target speaker's voice from a mixture of voices using visual cues. Existing works on audio-visual speech separation have demonstrated their performance with promising intelligibility, but maintaining…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-31 Suyeon Lee , Chaeyoung Jung , Youngjoon Jang , Jaehun Kim , Joon Son Chung

Contrastive language-audio pretraining (CLAP) is widely used for audio generation and recognition tasks. For example, CLAPScore, which utilizes the similarity of CLAP embeddings, has been a major metric for the evaluation of the relevance…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-11 Taisei Takano , Yuki Okamoto , Yusuke Kanamori , Yuki Saito , Ryotaro Nagase , Hiroshi Saruwatari

Speech Emotion Recognition (SER) is fundamental to affective computing and human-computer interaction, yet existing models struggle to generalize across diverse acoustic conditions. While Contrastive Language-Audio Pretraining (CLAP)…

Sound · Computer Science 2025-07-08 Jiacheng Shi , Yanfu Zhang , Ye Gao

Recent semi-supervised learning (SSL) methods typically include a filtering strategy to improve the quality of pseudo labels. However, these filtering strategies are usually hand-crafted and do not change as the model is updated, resulting…

Machine Learning · Computer Science 2023-09-19 Lei Zhu , Zhanghan Ke , Rynson Lau

Large Language Models (LLMs) can be adapted to extend their text capabilities to speech inputs. However, these speech-adapted LLMs consistently underperform their text-based counterparts--and even cascaded pipelines--on language…

Computation and Language · Computer Science 2026-02-24 Santiago Cuervo , Skyler Seto , Maureen de Seyssel , Richard He Bai , Zijin Gu , Tatiana Likhomanenko , Navdeep Jaitly , Zakaria Aldeneh

Recently, automatic speaker verification (ASV) based on deep learning is easily contaminated by adversarial attacks, which is a new type of attack that injects imperceptible perturbations to audio signals so as to make ASV produce wrong…

Audio and Speech Processing · Electrical Eng. & Systems 2024-07-10 Yibo Bai , Xiao-Lei Zhang , Xuelong Li

Fake speech detection systems have become a necessity to combat against speech deepfakes. Current systems exhibit poor generalizability on out-of-domain speech samples due to lack to diverse training data. In this paper, we attempt to…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-27 Rishith Sadashiv T N , Abhishek Bedge , Saisha Suresh Bore , Jagabandhu Mishra , Mrinmoy Bhattacharjee , S R Mahadeva Prasanna

Large self-supervised speech (SSL) models achieve strong downstream performance, but their size limits deployment in resource-constrained settings. We present HArnESS, an Arabic-centric self-supervised speech model family trained from…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-17 Vrunda N. Sukhadia , Shammur Absar Chowdhury

Overlapping Speech Detection (OSD) aims to identify regions where multiple speakers overlap in a conversation, a critical challenge in multi-party speech processing. This work proposes a speaker-aware progressive OSD model that leverages a…

Sound · Computer Science 2025-05-30 Zhaokai Sun , Li Zhang , Qing Wang , Pan Zhou , Lei Xie

Vision Language Models (VLMs) have shown remarkable capabilities in multimodal understanding, yet their susceptibility to perturbations poses a significant threat to their reliability in real-world applications. Despite often being…

Computer Vision and Pattern Recognition · Computer Science 2025-06-05 Jia Fu , Yongtao Wu , Yihang Chen , Kunyu Peng , Xiao Zhang , Volkan Cevher , Sepideh Pashami , Anders Holst

Active Speaker Detection (ASD) aims to identify who is currently speaking in each frame of a video. Most state-of-the-art approaches rely on late fusion to combine visual and audio features, but late fusion often fails to capture…

Computer Vision and Pattern Recognition · Computer Science 2025-12-18 Yu Wang , Juhyung Ha , Frangil M. Ramirez , Yuchen Wang , David J. Crandall

The rapid advancement of artificial intelligence (AI) has enabled sophisticated audio generation and voice cloning technologies, posing significant security risks for applications reliant on voice authentication. While existing datasets and…

Sound · Computer Science 2025-05-22 Kunyang Huang , Bin Hu

We introduce ParaSpeechCLAP, a dual-encoder contrastive model that maps speech and text style captions into a common embedding space, supporting a wide range of intrinsic (speaker-level) and situational (utterance-level) descriptors (such…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-31 Anuj Diwan , Eunsol Choi , David Harwath

Recent works show that speech separation guided diarization (SSGD) is an increasingly promising direction, mainly thanks to the recent progress in speech separation. It performs diarization by first separating the speakers and then applying…

Audio and Speech Processing · Electrical Eng. & Systems 2024-05-24 Giovanni Morrone , Samuele Cornell , Luca Serafini , Enrico Zovato , Alessio Brutti , Stefano Squartini

Advances in automatic speaker verification (ASV) promote research into the formulation of spoofing detection systems for real-world applications. The performance of ASV systems can be degraded severely by multiple types of spoofing attacks,…

Sound · Computer Science 2024-08-27 Zhenyu Wang , John H. L. Hansen

Event-based semantic segmentation has gained popularity due to its capability to deal with scenarios under high-speed motion and extreme lighting conditions, which cannot be addressed by conventional RGB cameras. Since it is hard to…

Computer Vision and Pattern Recognition · Computer Science 2024-03-26 Linglin Jing , Yiming Ding , Yunpeng Gao , Zhigang Wang , Xu Yan , Dong Wang , Gerald Schaefer , Hui Fang , Bin Zhao , Xuelong Li

Speech separation (SS) has advanced significantly with neural network-based methods, showing improved performance on signal-level metrics. However, these methods often struggle to maintain speech intelligibility in the separated signals,…

Sound · Computer Science 2026-01-28 Tianhua Li , Chenda Li , Wei Wang , Xin Zhou , Xihui Chen , Jianqing Gao , Yanmin Qian