中文
相关论文

相关论文: FireRedChat: A Pluggable, Full-Duplex Voice Intera…

200 篇论文

Spoken dialogue generation is crucial for applications like podcasts, dynamic commentary, and entertainment content, but poses significant challenges compared to single-utterance text-to-speech (TTS). Key requirements include accurate…

We introduce a dialogue policy based on a transformer architecture, where the self-attention mechanism operates over the sequence of dialogue turns. Recent work has used hierarchical recurrent neural networks to encode multiple utterances…

计算与语言 · 计算机科学 2020-05-04 Vladimir Vlasov , Johannes E. M. Mosig , Alan Nichol

The capacity for highly complex, evidence-based, and strategically adaptive persuasion remains a formidable great challenge for artificial intelligence. Previous work, like IBM Project Debater, focused on generating persuasive speeches in…

计算与语言 · 计算机科学 2025-11-25 Allen Roush , Devin Gonier , John Hines , Judah Goldfeder , Philippe Martin Wyder , Sanjay Basu , Ravid Shwartz Ziv

Turn-taking is a crucial aspect of human-robot interaction, directly influencing conversational fluidity and user engagement. While previous research has explored turn-taking models in controlled environments, their robustness in real-world…

机器人学 · 计算机科学 2025-07-15 Koji Inoue , Yuki Okafuji , Jun Baba , Yoshiki Ohira , Katsuya Hyodo , Tatsuya Kawahara

With recent advances in autonomous driving, Voice Control Systems have become increasingly adopted as human-vehicle interaction methods. This technology enables drivers to use voice commands to control the vehicle and will be soon available…

机器学习 · 计算机科学 2021-12-03 Jiwei Guan , Xi Zheng , Chen Wang , Yipeng Zhou , Alireza Jolfa

High-quality, large-scale audio captioning is crucial for advancing audio understanding, yet current automated methods often generate captions that lack fine-grained detail and contextual accuracy, primarily due to their reliance on limited…

声音 · 计算机科学 2025-06-03 Shunian Chen , Xinyuan Xie , Zheshu Chen , Liyan Zhao , Owen Lee , Zhan Su , Qilin Sun , Benyou Wang

Spoken language understanding is typically based on pipeline architectures including speech recognition and natural language understanding steps. These components are optimized independently to allow usage of available data, but the overall…

音频与语音处理 · 电气工程与系统科学 2020-08-13 Pavel Denisov , Ngoc Thang Vu

Task-oriented dialogue systems have been plagued by the difficulties of obtaining large-scale and high-quality annotated conversations. Furthermore, most of the publicly available datasets only include written conversations, which are…

Humans naturally process real-world multimodal information in a full-duplex manner. In artificial intelligence, replicating this capability is essential for advancing model development and deployment, particularly in embodied contexts. The…

人工智能 · 计算机科学 2025-06-03 Yiqun Yao , Xiang Li , Xin Jiang , Xuezhi Fang , Naitong Yu , Aixin Sun , Yequan Wang

With the recent advancements in reasoning capabilities, tool calling using MCP servers and Audio Language Models (ALMs), development and integration of multi-modal agents (with voice and text support) has come to the industry forefront.…

声音 · 计算机科学 2026-03-03 Anupam Purwar , Aditya Choudhary

We propose a separation guided speaker diarization (SGSD) approach by fully utilizing a complementarity of speech separation and speaker clustering. Since the conventional clustering-based speaker diarization (CSD) approach cannot well…

音频与语音处理 · 电气工程与系统科学 2021-07-07 Shu-Tong Niu , Jun Du , Lei Sun , Chin-Hui Lee

Interactions with virtual assistants typically start with a predefined trigger phrase followed by the user command. To make interactions with the assistant more intuitive, we explore whether it is feasible to drop the requirement that users…

The adoption of multimodal interactions by Voice Assistants (VAs) is growing rapidly to enhance human-computer interactions. Smartwatches have now incorporated trigger-less methods of invoking VAs, such as Raise To Speak (RTS), where the…

Recent advancements in zero-shot text-to-speech (TTS) modeling have led to significant strides in generating high-fidelity and diverse speech. However, dialogue generation, along with achieving human-like naturalness in speech, continues to…

音频与语音处理 · 电气工程与系统科学 2024-12-17 Leying Zhang , Yao Qian , Long Zhou , Shujie Liu , Dongmei Wang , Xiaofei Wang , Midia Yousefi , Yanmin Qian , Jinyu Li , Lei He , Sheng Zhao , Michael Zeng

Cascaded LLM systems coordinate models of varying sizes with human experts to balance accuracy, cost, and abstention under uncertainty. However, single-model tiers at each stage often struggle with ambiguous queries, triggering premature…

计算与语言 · 计算机科学 2026-04-15 Raeyoung Chang , Dongwook Kwon , Jisoo Lee , Nikhil Verma

With the rapid explosion of the World Wide Web, it is becoming increasingly possible to easily acquire a wide variety of information such as flight schedules, yellow pages, used car prices, current stock prices, entertainment event…

cmp-lg · 计算机科学 2008-02-03 Rajeev Agarwal

Multi-party open-ended conversation remains a major challenge in human-robot interaction, particularly when robots must recognise speakers, allocate turns, and respond coherently under overlapping or rapidly shifting dialogue. This paper…

人机交互 · 计算机科学 2025-12-15 Giulio Antonio Abbo , Maria Jose Pinto-Bernal , Martijn Catrycke , Tony Belpaeme

Existing 3D visual grounding methods rely on precise text prompts to locate objects within 3D scenes. Speech, as a natural and intuitive modality, offers a promising alternative. Real-world speech inputs, however, often suffer from…

计算机视觉与模式识别 · 计算机科学 2025-06-18 Yu Qi , Lipeng Gu , Honghua Chen , Liangliang Nan , Mingqiang Wei

The pipeline for multi-participant audiobook production primarily consists of three stages: script analysis, character voice timbre selection, and speech synthesis. Among these, script analysis can be automated with high accuracy using NLP…

音频与语音处理 · 电气工程与系统科学 2025-09-22 Ziqi Dai , Yiting Chen , Jiacheng Xu , Liufei Xie , Yuchen Wang , Zhenchuan Yang , Bingsong Bai , Yangsheng Gao , Wenjiang Zhou , Weifeng Zhao , Ruohua Zhou

There has been considerable progress made towards conversational models that generate coherent and fluent responses; however, this often involves training large language models on large dialogue datasets, such as Reddit. These large…

计算与语言 · 计算机科学 2020-10-12 Andrea Madotto , Etsuko Ishii , Zhaojiang Lin , Sumanth Dathathri , Pascale Fung