中文
相关论文

相关论文: A Scalable Pipeline for Enabling Non-Verbal Speech…

200 篇论文

In this work, we address the problem of generating speech from silent lip videos for any speaker in the wild. In stark contrast to previous works, our method (i) is not restricted to a fixed number of speakers, (ii) does not explicitly…

计算机视觉与模式识别 · 计算机科学 2022-09-02 Sindhu B Hegde , K R Prajwal , Rudrabha Mukhopadhyay , Vinay P Namboodiri , C. V. Jawahar

Humans naturally perform audiovisual speech recognition (AVSR), enhancing the accuracy and robustness by integrating auditory and visual information. Spiking neural networks (SNNs), which mimic the brain's information-processing mechanisms,…

多媒体 · 计算机科学 2025-08-28 Qianhui Liu , Jiadong Wang , Yang Wang , Xin Yang , Gang Pan , Haizhou Li

Naturalistic recordings capture audio in real-world environments where participants behave naturally without interference from researchers or experimental protocols. Naturalistic long-form recordings extend this concept by capturing…

音频与语音处理 · 电气工程与系统科学 2025-09-24 Jialu Li , Marvin Lavechin , Xulin Fan , Nancy L. McElwain , Alejandrina Cristia , Paola Garcia-Perera , Mark Hasegawa-Johnson

The key challenge of generative Visual Dialogue (VD) systems is to respond to human queries with informative answers in natural and contiguous conversation flow. Traditional Maximum Likelihood Estimation (MLE)-based methods only learn from…

计算机视觉与模式识别 · 计算机科学 2019-08-15 Heming Zhang , Shalini Ghosh , Larry Heck , Stephen Walsh , Junting Zhang , Jie Zhang , C. -C. Jay Kuo

A natural language interface exploits the conceptual simplicity and naturalness of the language to create a high-level user-friendly communication channel between humans and machines. One of the promising applications of such interfaces is…

计算与语言 · 计算机科学 2016-07-05 Kaveh Hassani , Won-Sook Lee

Currently, a common approach in many speech processing tasks is to leverage large scale pre-trained models by fine-tuning them on in-domain data for a particular application. Yet obtaining even a small amount of such data can be…

音频与语音处理 · 电气工程与系统科学 2024-08-20 Samuele Cornell , Jordan Darefsky , Zhiyao Duan , Shinji Watanabe

The problem of synthetic speech detection has enjoyed considerable attention, with recent methods achieving low error rates across several established benchmarks. However, to what extent can low error rates on academic benchmarks translate…

音频与语音处理 · 电气工程与系统科学 2025-05-23 Ashi Garg , Zexin Cai , Lin Zhang , Henry Li Xinyuan , Leibny Paola García-Perera , Kevin Duh , Sanjeev Khudanpur , Matthew Wiesner , Nicholas Andrews

Spoken language models (SLMs) have emerged as a unified paradigm for speech understanding and generation, enabling natural human machine interaction. However, while most progress has focused on semantic accuracy and instruction following,…

The evaluation of audio fingerprinting at a realistic scale is limited by the scarcity of large public music databases. We present an audio-free approach that synthesises latent fingerprints which approximate the distribution of real…

声音 · 计算机科学 2025-09-24 Aditya Bhattacharjee , Marco Pasini , Emmanouil Benetos

To build an interpretable neural text classifier, most of the prior work has focused on designing inherently interpretable models or finding faithful explanations. A new line of work on improving model interpretability has just started, and…

计算与语言 · 计算机科学 2020-11-20 Hanjie Chen , Yangfeng Ji

For speaker recognition, it is difficult to extract an accurate speaker representation from speech because of its mixture of speaker traits and content. This paper proposes a disentanglement framework that simultaneously models speaker…

音频与语音处理 · 电气工程与系统科学 2023-11-02 Tianchi Liu , Kong Aik Lee , Qiongqiong Wang , Haizhou Li

Sarcasm fundamentally alters meaning through tone and context, yet detecting it in speech remains a challenge due to data scarcity. In addition, existing detection systems often rely on multimodal data, limiting their applicability in…

计算与语言 · 计算机科学 2026-04-21 Zhu Li , Yuqing Zhang , Xiyuan Gao , Shekhar Nayak , Matt Coler

Various applications of voice synthesis have been developed independently despite the fact that they generate "voice" as output in common. In addition, the majority of voice synthesis models currently rely on annotated audio data, but it is…

音频与语音处理 · 电气工程与系统科学 2023-05-31 Rongjie Huang , Chunlei Zhang , Yongqi Wang , Dongchao Yang , Luping Liu , Zhenhui Ye , Ziyue Jiang , Chao Weng , Zhou Zhao , Dong Yu

We describe a neural network-based system for text-to-speech (TTS) synthesis that is able to generate speech audio in the voice of many different speakers, including those unseen during training. Our system consists of three independently…

Laughter serves as a multifaceted communicative signal in human interaction, yet its identification within dialogue presents a significant challenge for conversational AI systems. This study addresses this challenge by annotating laughable…

计算与语言 · 计算机科学 2025-03-19 Koji Inoue , Mikey Elmers , Divesh Lala , Tatsuya Kawahara

In this paper, we propose three methods for generating synthetic samples to train and evaluate multimodal large language models capable of processing both text and speech inputs. Addressing the scarcity of samples containing both…

音频与语音处理 · 电气工程与系统科学 2024-06-21 Vahid Noroozi , Zhehuai Chen , Somshubra Majumdar , Steve Huang , Jagadeesh Balam , Boris Ginsburg

Computational methods to aid journalists in the task often require adapting a model to specific domains and generating explanations. However, most automated fact-checking methods rely on three-class datasets, which do not accurately reflect…

计算与语言 · 计算机科学 2024-10-08 Jing Yang , Anderson Rocha

Recent advances in Large Language Models (LLMs) have shown promise in automating discourse annotation for conversations. While manually designing tree annotation schemes significantly improves annotation quality for humans and models, their…

计算与语言 · 计算机科学 2025-06-04 Kseniia Petukhova , Ekaterina Kochmar

Speech synthesis has recently seen significant improvements in fidelity, driven by the advent of neural vocoders and neural prosody generators. However, these systems lack intuitive user controls over prosody, making them unable to rectify…

音频与语音处理 · 电气工程与系统科学 2020-08-13 Max Morrison , Zeyu Jin , Justin Salamon , Nicholas J. Bryan , Gautham J. Mysore

In this paper, we propose a novel text-based talking-head video generation framework that synthesizes high-fidelity facial expressions and head motions in accordance with contextual sentiments as well as speech rhythm and pauses. To be…

计算机视觉与模式识别 · 计算机科学 2021-05-10 Lincheng Li , Suzhen Wang , Zhimeng Zhang , Yu Ding , Yixing Zheng , Xin Yu , Changjie Fan