中文
相关论文

相关论文: voc2vec: A Foundation Model for Non-Verbal Vocaliz…

200 篇论文

We present a novel approach to any-to-one (A2O) voice conversion (VC) in a sequence-to-sequence (seq2seq) framework. A2O VC aims to convert any speaker, including those unseen during training, to a fixed target speaker. We utilize…

音频与语音处理 · 电气工程与系统科学 2020-10-26 Wen-Chin Huang , Yi-Chiao Wu , Tomoki Hayashi , Tomoki Toda

Traditional pet emotion recognition from vocalizations, based on discrete classification, struggles with ambiguity and capturing intensity variations. We propose a continuous Valence-Arousal (VA) model that represents emotions in a…

声音 · 计算机科学 2025-10-16 Junyao Huang , Rumin Situ

Dialogue models falter in noisy, multi-speaker environments, often producing irrelevant responses and awkward turn-taking. We present AV-Dialog, the first multimodal dialog framework that uses both audio and visual cues to track the target…

计算与语言 · 计算机科学 2025-11-17 Tuochao Chen , Bandhav Veluri , Hongyu Gong , Shyamnath Gollakota

This work details our approach to achieving a leading system with a 1.79% pooled equal error rate (EER) on the evaluation set of the Controlled Singing Voice Deepfake Detection (CtrSVDD). The rapid advancement of generative AI models…

音频与语音处理 · 电气工程与系统科学 2024-10-22 Anmol Guragain , Tianchi Liu , Zihan Pan , Hardik B. Sailor , Qiongqiong Wang

This survey paper provides a comprehensive overview of the recent advancements and challenges in applying large language models to the field of audio signal processing. Audio processing, with its diverse signal representations and a wide…

Recent work on intracranial brain-machine interfaces has demonstrated that spoken speech can be decoded with high accuracy, essentially by treating the problem as an instance of supervised learning and training deep neural networks to map…

神经元与认知 · 定量生物学 2024-05-30 Brian A. Yuan , Joseph G. Makin

We introduce a new approach for audio-visual speech separation. Given a video, the goal is to extract the speech associated with a face in spite of simultaneous background sounds and/or other human speakers. Whereas existing methods focus…

计算机视觉与模式识别 · 计算机科学 2021-04-07 Ruohan Gao , Kristen Grauman

While the use of deep neural networks has significantly boosted speaker recognition performance, it is still challenging to separate speakers in poor acoustic environments. Here speech enhancement methods have traditionally allowed improved…

音频与语音处理 · 电气工程与系统科学 2020-08-28 Yanpei Shi , Qiang Huang , Thomas Hain

Producing a large amount of annotated speech data for training ASR systems remains difficult for more than 95% of languages all over the world which are low-resourced. However, we note human babies start to learn the language by the sounds…

计算与语言 · 计算机科学 2018-10-31 Yi-Chen Chen , Chia-Hao Shen , Sung-Feng Huang , Hung-yi Lee , Lin-shan Lee

The rapid progress of Large Language Models (LLMs) has empowered omni models to act as voice assistants capable of understanding spoken dialogues. These models can process multimodal inputs beyond text, such as speech and visual data,…

Current state of the art acoustic models can easily comprise more than 100 million parameters. This growing complexity demands larger training datasets to maintain a decent generalization of the final decision function. An ideal dataset is…

音频与语音处理 · 电气工程与系统科学 2022-02-01 Philipp Klumpp , Tomás Arias-Vergara , Paula Andrea Pérez-Toro , Elmar Nöth , Juan Rafael Orozco-Arroyave

Modern speaker recognition system relies on abundant and balanced datasets for classification training. However, diverse defective datasets, such as partially-labelled, small-scale, and imbalanced datasets, are common in real-world…

音频与语音处理 · 电气工程与系统科学 2025-09-03 Ruijie Tao , Zhan Shi , Yidi Jiang , Tianchi Liu , Haizhou Li

Automatic emotion recognition is one of the central concerns of the Human-Computer Interaction field as it can bridge the gap between humans and machines. Current works train deep learning models on low-level data representations to solve…

音频与语音处理 · 电气工程与系统科学 2021-11-22 Mariana Rodrigues Makiuchi , Kuniaki Uto , Koichi Shinoda

Speech is known to carry health-related attributes, which has emerged as a novel venue for remote and long-term health monitoring. However, existing models are usually tailored for a specific type of disease, and have been shown to lack…

音频与语音处理 · 电气工程与系统科学 2024-10-11 Yi Zhu , Tiago Falk

Recent years have seen a surge in finding association between faces and voices within a cross-modal biometric application along with speaker recognition. Inspired from this, we introduce a challenging task in establishing association…

计算机视觉与模式识别 · 计算机科学 2021-04-23 Muhammad Saad Saeed , Shah Nawaz , Pietro Morerio , Arif Mahmood , Ignazio Gallo , Muhammad Haroon Yousaf , Alessio Del Bue

Speaker anonymization systems continue to improve their ability to obfuscate the original speaker characteristics in a speech signal, but often create processing artifacts and unnatural sounding voices as a tradeoff. Many of those systems…

音频与语音处理 · 电气工程与系统科学 2023-08-23 Ünal Ege Gaznepoglu , Nils Peters

We first propose a new task named Dialogue Description (Dial2Desc). Unlike other existing dialogue summarization tasks such as meeting summarization, we do not maintain the natural flow of a conversation but describe an object or an action…

计算与语言 · 计算机科学 2018-11-02 Haojie Pan , Junpei Zhou , Zhou Zhao , Yan Liu , Deng Cai , Min Yang

Voicebots have provided a new avenue for supporting the development of language skills, particularly within the context of second language learning. Voicebots, though, have largely been geared towards native adult speakers. We sought to…

计算与语言 · 计算机科学 2024-07-24 Simone Wills , Yu Bai , Cristian Tejedor-Garcia , Catia Cucchiarini , Helmer Strik

Voice conversion (VC) modifies voice characteristics while preserving linguistic content. This paper presents the Stepback network, a novel model for converting speaker identity using non-parallel data. Unlike traditional VC methods that…

声音 · 计算机科学 2025-01-28 Qian Yang , Calbert Graham

Voice activity detection (VAD) is an important pre-processing step for speech technology applications. The task consists of deriving segment boundaries of audio signals which contain voicing information. In recent years, it has been shown…

音频与语音处理 · 电气工程与系统科学 2023-03-28 Eklavya Sarkar , RaviShankar Prasad , Mathew Magimai. -Doss