中文
相关论文

相关论文: DualVC 2: Dynamic Masked Convolution for Unified S…

200 篇论文

Speech Emotion Recognition (SER) systems often degrade in performance when exposed to the unpredictable acoustic interference found in real-world environments. Additionally, the opacity of deep learning models hinders their adoption in…

声音 · 计算机科学 2025-12-23 Sudip Chakrabarty , Pappu Bishwas , Rajdeep Chatterjee

Audio tokenization bridges continuous waveforms and multi-track music language models. In dual-track modeling, tokens should preserve three properties at once: high-fidelity reconstruction, strong predictability under a language model, and…

声音 · 计算机科学 2026-04-02 Rui Lin , Zhiyue Wu , Jiahe Le , Kangdi Wang , Weixiong Chen , Junyu Dai , Tao Jiang

Voice conversion (VC) aims to modify the speaker's identity while preserving the linguistic content. Commonly, VC methods use an encoder-decoder architecture, where disentangling the speaker's identity from linguistic information is…

音频与语音处理 · 电气工程与系统科学 2024-09-19 Philip H. Lee , Ismail Rasim Ulgen , Berrak Sisman

In this paper, we propose "SoundSpring", a cutting-edge error-resilient audio transceiver that marries the robustness benefits of joint source-channel coding (JSCC) while also being compatible with current digital communication systems.…

音频与语音处理 · 电气工程与系统科学 2025-01-23 Shengshi Yao , Jincheng Dai , Xiaoqi Qin , Sixian Wang , Siye Wang , Kai Niu , Ping Zhang

In this paper, we propose an end-to-end speech recognition network based on Nvidia's previous QuartzNet model. We try to promote the model performance, and design three components: (1) Multi-Resolution Convolution Module, replaces the…

音频与语音处理 · 电气工程与系统科学 2020-11-30 Jian Luo , Jianzong Wang , Ning Cheng , Guilin Jiang , Jing Xiao

It was shown recently that a combination of ASR and TTS models yield highly competitive performance on standard voice conversion tasks such as the Voice Conversion Challenge 2020 (VCC2020). To obtain good performance both models require…

音频与语音处理 · 电气工程与系统科学 2022-04-01 Mingjie Chen , Yanghao Zhou , Heyan Huang , Thomas Hain

Neural channel decoder, as a data-driven channel decoding strategy, has shown very promising improvement on error-correcting capability over the classical methods. However, the success of those deep learning-based decoder comes at the cost…

信息论 · 计算机科学 2026-05-20 Chengwei Zhang , Yifan Du , Siyu Liao

Speech separation (SS) seeks to disentangle a multi-talker speech mixture into single-talker speech streams. Although SS can be generally achieved using offline methods, such a processing paradigm is not suitable for real-time streaming…

声音 · 计算机科学 2025-04-04 Wupeng Wang , Zexu Pan , Xinke Li , Shuai Wang , Haizhou Li

Recent applications of Convolutional Neural Networks (ConvNets) for human action recognition in videos have proposed different solutions for incorporating the appearance and motion information. We study a number of ways of fusing ConvNet…

计算机视觉与模式识别 · 计算机科学 2016-09-27 Christoph Feichtenhofer , Axel Pinz , Andrew Zisserman

This paper studies deep network architectures to address the problem of video classification. A multi-stream framework is proposed to fully utilize the rich multimodal information in videos. Specifically, we first train three Convolutional…

计算机视觉与模式识别 · 计算机科学 2015-11-12 Zuxuan Wu , Yu-Gang Jiang , Xi Wang , Hao Ye , Xiangyang Xue , Jun Wang

Singing Voice Conversion (SVC) transfers a source singer's timbre to a target while keeping melody and lyrics. The key challenge in any-to-any SVC is adapting unseen speaker timbres to source audio without quality degradation. Existing…

声音 · 计算机科学 2025-08-11 Wei Chen , Binzhu Sha , Dan Luo , Jing Yang , Zhuo Wang , Fan Fan , Zhiyong Wu

Medical multimodal representation learning aims to integrate heterogeneous data into unified patient representations to support clinical outcome prediction. However, real-world medical datasets commonly contain systematic biases from…

机器学习 · 计算机科学 2026-05-19 Xiaoguang Zhu , Linxiao Gong , Lianlong Sun , Yang Liu , Haoyu Wang , Jing Liu

The practical deployment of diffusion-based Neural Video Compression (NVC) faces critical challenges, including severe information loss, prohibitive inference latency, and poor temporal consistency. To bridge this gap, we propose DiffVC-RT,…

计算机视觉与模式识别 · 计算机科学 2026-01-29 Wenzhuo Ma , Zhenzhong Chen

While both the data volume and heterogeneity of the digital music content is huge, it has become increasingly important and convenient to build a recommendation or search system to facilitate surfacing these content to the user or consumer…

Deep convolutional networks have achieved great success for object recognition in still images. However, for action recognition in videos, the improvement of deep convolutional networks is not so evident. We argue that there are two reasons…

计算机视觉与模式识别 · 计算机科学 2015-07-09 Limin Wang , Yuanjun Xiong , Zhe Wang , Yu Qiao

We introduce Diffusion-based Audio Captioning (DAC), a non-autoregressive diffusion model tailored for diverse and efficient audio captioning. Although existing captioning models relying on language backbones have achieved remarkable…

计算与语言 · 计算机科学 2025-06-03 Manjie Xu , Chenxing Li , Xinyi Tu , Yong Ren , Ruibo Fu , Wei Liang , Dong Yu

The parameters estimation of a system using indirect measurements over the same system is a problem that occurs in many fields of engineering, known as the inverse problem. It also happens in the field of underwater acoustic, especially in…

信号处理 · 电气工程与系统科学 2020-03-31 Marco Apolinario , Samuel Huaman Bustamante , Giorgio Morales , Joel Telles , Daniel Diaz

We study transfer learning in convolutional network architectures applied to the task of recognizing audio, such as environmental sound events and speech commands. Our key finding is that not only is it possible to transfer representations…

声音 · 计算机科学 2017-10-24 Brian McMahan , Delip Rao

A diffusion-based voice conversion (VC) model (e.g., VoiceGrad) can achieve high speech quality and speaker similarity; however, its conversion process is slow owing to iterative sampling. FastVoiceGrad overcomes this limitation by…

声音 · 计算机科学 2025-08-26 Takuhiro Kaneko , Hirokazu Kameoka , Kou Tanaka , Yuto Kondo

As parallel training data is scarce for one-shot voice conversion (VC) tasks, waveform reconstruction is typically performed by various VC systems. A typical one-shot VC system comprises a content encoder and a speaker encoder. However, two…

音频与语音处理 · 电气工程与系统科学 2025-06-03 Songjun Cao , Qinghua Wu , Jie Chen , Jin Li , Long Ma