English
Related papers

Related papers: LibriVAD: A Scalable Open Dataset with Deep Learni…

200 papers

Scaling laws guide large language model training by relating compute to cross-entropy loss, and recent work further extends them to predict downstream benchmark performance. However, prior approaches face generalization limitations from two…

Computation and Language · Computer Science 2026-05-28 Quanen Sun , Changxin Tian , Ke Shi , Cai Chen , Cunyin Peng , Jia Liu , Kunlong Chen , Zhiqiang Zhang

Open-vocabulary detection (OVD) is a challenging task to detect and classify objects from an unrestricted set of categories, including those unseen during training. Existing open-vocabulary detectors are limited by complex visual-textual…

Computer Vision and Pattern Recognition · Computer Science 2025-02-27 Caixiong Li , Xiongwei Zhao , Jinhang Zhang , Xing Zhang , Qihao Sun , Zhou Wu

Keyword spotting (KWS) and speaker verification (SV) have been studied independently although it is known that acoustic and speaker domains are complementary. In this paper, we propose a multi-task network that performs KWS and SV…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-10 Myunghun Jung , Youngmoon Jung , Jahyun Goo , Hoirin Kim

Open-vocabulary semantic segmentation (OVSS) underpins many vision and robotics tasks that require generalizable semantic understanding. Existing approaches either rely on limited segmentation training data, which hinders generalization, or…

Computer Vision and Pattern Recognition · Computer Science 2026-04-13 Omar Alama , Darshil Jariwala , Avigyan Bhattacharya , Seungchan Kim , Wenshan Wang , Sebastian Scherer

3D object detection plays a crucial role in autonomous systems, yet existing methods are limited by closed-set assumptions and struggle to recognize novel objects and their attributes in real-world scenarios. We propose OVODA, a novel…

Computer Vision and Pattern Recognition · Computer Science 2025-08-26 Xinhao Xiang , Kuan-Chuan Peng , Suhas Lohit , Michael J. Jones , Jiawei Zhang

Current audio deepfake detection has achieved remarkable performance using diverse deep learning architectures such as ResNet, and has seen further improvements with the introduction of large models (LMs) like Wav2Vec. The success of large…

Sound · Computer Science 2026-03-27 Yupei Li , Shuaijie Shao , Manuel Milling , Björn Schuller

Despite its broad practical applications such as in fraud prevention, open-set speaker identification (OSI) has received less attention in the speaker recognition community compared to speaker verification (SV). OSI deals with determining…

Audio and Speech Processing · Electrical Eng. & Systems 2023-07-04 Raghuveer Peri , Seyed Omid Sadjadi , Daniel Garcia-Romero

Voice timbre attribute detection (vTAD) is the task of determining the relative intensity of timbre attributes between speech utterances. Voice timbre is a crucial yet inherently complex component of speech perception. While deep neural…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-06 Aemon Yat Fei Chiu , Yujia Xiao , Qiuqiang Kong , Tan Lee

Voice activity detection (VAD) plays a vital role in enabling applications such as speech recognition. We analyze the impact of window size on the accuracy of three VAD algorithms: Silero, WebRTC, and Root Mean Square (RMS) across a set of…

Sound · Computer Science 2026-01-27 Max McKinnon , Samir Khaki , Chandan KA Reddy , William Huang

Recent advances in AI-generated voices have intensified the challenge of detecting deepfake audio, posing risks for scams and the spread of disinformation. To tackle this issue, we establish the largest public voice dataset to date, named…

Thanks to the emerging of foundation models, the large language and vision models are integrated to acquire the multimodal ability of visual captioning, question answering, etc. Although existing multimodal models present impressive…

Computer Vision and Pattern Recognition · Computer Science 2023-12-29 Bo Zhao , Boya Wu , Muyang He , Tiejun Huang

Reverberation not only degrades the quality of speech for human perception, but also severely impacts the accuracy of automatic speech recognition. Prior work attempts to remove reverberation based on the audio modality only. Our idea is to…

Sound · Computer Science 2023-03-15 Changan Chen , Wei Sun , David Harwath , Kristen Grauman

The development of robust, multilingual speaker recognition systems is hindered by a lack of large-scale, publicly available and multilingual datasets, particularly for the read-speech style crucial for applications like anti-spoofing. To…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-26 Aref Farhadipour , Jan Marquenie , Srikanth Madikeri , Eleanor Chodroff

Retrieval-augmented models have proven to be effective in natural language processing tasks, yet there remains a lack of research on their optimization using variational inference. We introduce the Variational Open-Domain (VOD) framework…

Computation and Language · Computer Science 2023-06-01 Valentin Liévin , Andreas Geert Motzfeldt , Ida Riis Jensen , Ole Winther

Noise robustness in speech foundation models (SFMs) has been a critical challenge, as most models are primarily trained on clean data and experience performance degradation when the models are exposed to noisy speech. To address this issue,…

Sound · Computer Science 2025-08-19 Hyebin Ahn , Kangwook Jang , Hoirin Kim

Benefiting from the development of deep learning, text-to-speech (TTS) techniques using clean speech have achieved significant performance improvements. The data collected from real scenes often contains noise and generally needs to be…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-06 Qiushi Zhu , Yu Gu , Rilin Chen , Chao Weng , Yuchen Hu , Lirong Dai , Jie Zhang

Visual speech recognition (VSR), which decodes spoken words from video data, offers significant benefits, particularly when audio is unavailable. However, the high dimensionality of video data leads to prohibitive computational costs that…

Computer Vision and Pattern Recognition · Computer Science 2025-02-10 Iason Ioannis Panagos , Giorgos Sfikas , Christophoros Nikou

Current speaker diarization systems rely on an external voice activity detection model prior to speaker embedding extraction on the detected speech segments. In this paper, we establish that the attention system of a speaker embedding…

Audio and Speech Processing · Electrical Eng. & Systems 2024-05-16 Jenthe Thienpondt , Kris Demuynck

Recent vision-language pre-trained models (VL-PTMs) have shown remarkable success in open-vocabulary tasks. However, downstream use cases often involve further fine-tuning of VL-PTMs, which may distort their general knowledge and impair…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Lin Zhu , Yifeng Yang , Qinying Gu , Xinbing Wang , Chenghu Zhou , Nanyang Ye

Prosody diversity is essential for achieving naturalness and expressiveness in zero-shot text-to-speech (TTS). However, frequently used acoustic metrics capture only partial views of prosodic variation and correlate poorly with human…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-02 Yifan Yang , Bing Han , Hui Wang , Long Zhou , Wei Wang , Mingyu Cui , Xu Tan , Xie Chen
‹ Prev 1 3 4 5 6 7 10 Next ›