中文
相关论文

相关论文: VoiceCoach: Interactive Evidence-based Training fo…

200 篇论文

In this work, we present a novel perspective on cognitive impairment classification from speech by integrating speech foundation models that explicitly recognize speech dialects. Our motivation is based on the observation that individuals…

声音 · 计算机科学 2026-01-14 Tiantian Feng , Anfeng Xu , Jinkook Lee , Shrikanth Narayanan

In text-to-speech synthesis, the ability to control voice characteristics is vital for various applications. By leveraging thriving text prompt-based generation techniques, it should be possible to enhance the nuanced control of voice…

Speech foundation models have demonstrated exceptional capabilities in speech-related tasks. Nevertheless, these models often struggle with non-verbal audio data, such as vocalizations, baby crying, etc., which are critical for various…

音频与语音处理 · 电气工程与系统科学 2025-02-25 Alkis Koudounas , Moreno La Quatra , Marco Sabato Siniscalchi , Elena Baralis

One of the long-term goals of artificial intelligence is to build an agent that can communicate intelligently with human in natural language. Most existing work on natural language learning relies heavily on training over a pre-collected…

计算与语言 · 计算机科学 2017-05-30 Haichao Zhang , Haonan Yu , Wei Xu

Mental manipulation, the strategic use of language to covertly influence or exploit others, is a newly emerging task in computational social reasoning. Prior work has focused exclusively on textual conversations, overlooking how…

计算与语言 · 计算机科学 2026-01-14 Run Chen , Wen Liang , Ziwei Gong , Lin Ai , Julia Hirschberg

The growing prevalence of speech deepfakes has raised serious concerns, particularly in real-world scenarios such as telephone fraud and identity theft. While many anti-spoofing systems have demonstrated promising performance on…

音频与语音处理 · 电气工程与系统科学 2026-05-12 Tong Zhang , Yihuan Huang , Yanzhen Ren

Training dialog policies for speech-based virtual assistants requires a plethora of conversational data. The data collection phase is often expensive and time consuming due to human involvement. To address this issue, a common solution is…

计算与语言 · 计算机科学 2019-11-11 Maryam Fazel-Zarandi , Longshaokan Wang , Aditya Tiwari , Spyros Matsoukas

In the field of Geriatronics, enabling effective and transparent communication between humans and robots is crucial for enhancing the acceptance and performance of assistive robots. Our early-stage research project investigates the…

Unsupervised Zero-Shot Voice Conversion (VC) aims to modify the speaker characteristic of an utterance to match an unseen target speaker without relying on parallel training data. Recently, self-supervised learning of speech representation…

声音 · 计算机科学 2022-02-14 Trung Dang , Dung Tran , Peter Chin , Kazuhito Koishida

Sound is one of the most informative and abundant modalities in the real world while being robust to sense without contacts by small and cheap sensors that can be placed on mobile devices. Although deep learning is capable of extracting…

机器人学 · 计算机科学 2022-08-05 Xufeng Zhao , Cornelius Weber , Muhammad Burhan Hafez , Stefan Wermter

We propose a method for training language models in an interactive setting inspired by child language acquisition. In our setting, a speaker attempts to communicate some information to a listener in a single-turn dialogue and receives a…

计算与语言 · 计算机科学 2025-05-12 Lennart Stöpler , Rufat Asadli , Mitja Nikolaus , Ryan Cotterell , Alex Warstadt

Human tutoring interventions play a crucial role in supporting student learning, improving academic performance, and promoting personal growth. This paper focuses on analyzing mathematics tutoring discourse using talk moves - a framework of…

One problem that every presenter faces when delivering a public discourse is how to hold the listeners' attentions or to keep them involved. Therefore, many studies in conversation analysis work on this issue and suggest qualitatively…

计算与语言 · 计算机科学 2017-04-18 Zhe Liu , Anbang Xu , Mengdi Zhang , Jalal Mahmud , Vibha Sinha

Audio-visual speaker extraction has attracted increasing attention, as it removes the need for pre-registered speech and leverages the visual modality as a complement to audio. Although existing methods have achieved impressive performance,…

多媒体 · 计算机科学 2026-03-03 Jiadong Wang , Ke Zhang , Xinyuan Qian , Ruijie Tao , Haizhou Li , Björn Schuller

Collaborative robots must quickly adapt to their partner's intent and preferences to proactively identify helpful actions. This is especially true in situated settings where human partners can continually teach robots new high-level…

机器人学 · 计算机科学 2025-06-17 Jennifer Grannen , Siddharth Karamcheti , Blake Wulfe , Dorsa Sadigh

In speech processing pipelines, improving the quality and intelligibility of real-world recordings is crucial. While supervised regression is the primary method for speech enhancement, audio tokenization is emerging as a promising…

声音 · 计算机科学 2025-07-18 Luca Della Libera , Cem Subakan , Mirco Ravanelli

Emotional voice conversion (EVC) aims to change the emotional state of an utterance while preserving the linguistic content and speaker identity. In this paper, we propose a novel 2-stage training strategy for sequence-to-sequence emotional…

计算与语言 · 计算机科学 2021-06-10 Kun Zhou , Berrak Sisman , Haizhou Li

Diffusion models have revolutionized the field of talking head generation, yet still face challenges in expressiveness, controllability, and stability in long-time generation. In this research, we propose an EmotiveTalk framework to address…

计算机视觉与模式识别 · 计算机科学 2024-12-17 Haotian Wang , Yuzhe Weng , Yueyan Li , Zilu Guo , Jun Du , Shutong Niu , Jiefeng Ma , Shan He , Xiaoyan Wu , Qiming Hu , Bing Yin , Cong Liu , Qingfeng Liu

Acoustics-to-word models are end-to-end speech recognizers that use words as targets without relying on pronunciation dictionaries or graphemes. These models are notoriously difficult to train due to the lack of linguistic knowledge. It is…

音频与语音处理 · 电气工程与系统科学 2018-11-14 Hao Tang , James Glass

The performance of speaker verification systems degrades significantly under language mismatch, a critical challenge exacerbated by the field's reliance on English-centric data. To address this, we propose the TidyVoice Challenge for…

音频与语音处理 · 电气工程与系统科学 2026-01-30 Aref Farhadipour , Jan Marquenie , Srikanth Madikeri , Teodora Vukovic , Volker Dellwo , Kathy Reid , Francis M. Tyers , Ingo Siegert , Eleanor Chodroff