中文
相关论文

相关论文: MOSS-VoiceGenerator: Create Realistic Voices with …

200 篇论文

Text-to-Audio (TTA) generation is an emerging area within AI-generated content (AIGC), where audio is created from natural language descriptions. Despite growing interest, developing robust TTA models remains challenging due to the scarcity…

音频与语音处理 · 电气工程与系统科学 2025-01-29 Xinfa Zhu , Wenjie Tian , Xinsheng Wang , Lei He , Xi Wang , Sheng Zhao , Lei Xie

Speech synthesis has come a long way as current text-to-speech (TTS) models can now generate natural human-sounding speech. However, most of the TTS research focuses on using adult speech data and there has been very limited work done on…

声音 · 计算机科学 2022-04-05 Rishabh Jain , Mariam Yiwere , Dan Bigioi , Peter Corcoran , Horia Cucu

Speech is a natural interface for humans to interact with robots. Yet, aligning a robot's voice to its appearance is challenging due to the rich vocabulary of both modalities. Previous research has explored a few labels to describe robots…

人机交互 · 计算机科学 2024-02-09 Pol van Rijn , Silvan Mertes , Kathrin Janowski , Katharina Weitz , Nori Jacoby , Elisabeth André

We introduce Moshi, a speech-text foundation model and full-duplex spoken dialogue framework. Current systems for spoken dialogue rely on pipelines of independent components, namely voice activity detection, speech recognition, textual…

音频与语音处理 · 电气工程与系统科学 2024-10-03 Alexandre Défossez , Laurent Mazaré , Manu Orsini , Amélie Royer , Patrick Pérez , Hervé Jégou , Edouard Grave , Neil Zeghidour

Designing high-quality indoor 3D scenes is important in many practical applications, such as room planning or game development. Conventionally, this has been a time-consuming process which requires both artistic skill and familiarity with…

计算机视觉与模式识别 · 计算机科学 2024-07-31 Başak Melis Öcal , Maxim Tatarchenko , Sezer Karaoglu , Theo Gevers

Recent studies have outlined the accessibility challenges faced by blind or visually impaired, and less-literate people, in interacting with social networks, in-spite of facilitating technologies such as monotone text-to-speech (TTS) screen…

社会与信息网络 · 计算机科学 2024-10-28 Suparna De , Ionut Bostan , Nishanth Sastry

Developing AI agents powered by large language models (LLMs) faces significant challenges in achieving true Turing completeness and adaptive, code-driven evolution. Current approaches often generate code independently of its runtime…

软件工程 · 计算机科学 2024-09-25 Ming Zhu , Yi Zhou

User Simulators are one of the major tools that enable offline training of task-oriented dialogue systems. For this task the Agenda-Based User Simulator (ABUS) is often used. The ABUS is based on hand-crafted rules and its output is in…

计算与语言 · 计算机科学 2018-05-21 Florian Kreyssig , Inigo Casanueva , Pawel Budzianowski , Milica Gasic

Application of formal models provides many benefits for the software and system development, however, the learning curve of formal languages could be a critical factor for an industrial project. Thus, a natural language specification that…

软件工程 · 计算机科学 2016-12-07 Phan Vo Thu Nhat , Maria Spichkova

The generation of realistic and contextually relevant co-speech gestures is a challenging yet increasingly important task in the creation of multimodal artificial agents. Prior methods focused on learning a direct correspondence between…

人机交互 · 计算机科学 2023-05-09 Hendric Voß , Stefan Kopp

Text-based audio generation models have limitations as they cannot encompass all the information in audio, leading to restricted controllability when relying solely on text. To address this issue, we propose a novel model that enhances the…

声音 · 计算机科学 2023-12-29 Zhifang Guo , Jianguo Mao , Rui Tao , Long Yan , Kazushige Ouchi , Hong Liu , Xiangdong Wang

Text-to-audio generation (TTA) produces audio from a text description, learning from pairs of audio samples and hand-annotated text. However, commercializing audio generation is challenging as user-input prompts are often under-specified…

Although contemporary text-to-image generation models have achieved remarkable breakthroughs in producing visually appealing images, their capacity to generate precise and flexible typographic elements, especially non-Latin alphabets,…

计算机视觉与模式识别 · 计算机科学 2025-04-29 Haofan Wang , Yujia Xu , Yimeng Li , Junchen Li , Chaowei Zhang , Jing Wang , Kejia Yang , Zhibo Chen

Creating a virtual avatar with semantically coherent gestures that are aligned with speech is a challenging task. Existing gesture generation research mainly focused on generating rhythmic beat gestures, neglecting the semantic context of…

计算机视觉与模式识别 · 计算机科学 2025-07-28 Lanmiao Liu , Esam Ghaleb , Aslı Özyürek , Zerrin Yumak

In text-to-speech synthesis, the ability to control voice characteristics is vital for various applications. By leveraging thriving text prompt-based generation techniques, it should be possible to enhance the nuanced control of voice…

Text-to-speech (TTS) has advanced from generating natural-sounding speech to enabling fine-grained control over attributes like emotion, timbre, and style. Driven by rising industrial demand and breakthroughs in deep learning, e.g.,…

计算与语言 · 计算机科学 2025-08-26 Tianxin Xie , Yan Rong , Pengfei Zhang , Wenwu Wang , Li Liu

Socially competent robots should be equipped with the ability to perceive the world that surrounds them and communicate about it in a human-like manner. Representative skills that exhibit such ability include generating image descriptions…

机器人学 · 计算机科学 2021-02-01 Ting Han , Sina Zarrieß

Generating 3D human gestures and speech from a text script is critical for creating realistic talking avatars. One solution is to leverage separate pipelines for text-to-speech (TTS) and speech-to-gesture (STG), but this approach suffers…

多媒体 · 计算机科学 2024-09-26 Zixin Guo , Jian Zhang

Prior work in style-controlled text generation has focused on tasks such as emulating the style of prolific literary authors, producing formal or informal text, and mitigating toxicity of generated text. Plentiful demonstrations of these…

计算与语言 · 计算机科学 2024-03-05 Aleem Khan , Andrew Wang , Sophia Hager , Nicholas Andrews

The rapid growth of voice assistants powered by large language models (LLM) has highlighted a need for speech instruction data to train these systems. Despite the abundance of speech recognition data, there is a notable scarcity of speech…

音频与语音处理 · 电气工程与系统科学 2025-08-26 Alan Dao , Dinh Bach Vu , Huy Hoang Ha , Tuan Le Duc Anh , Shreyas Gopal , Yue Heng Yeo , Warren Keng Hoong Low , Eng Siong Chng , Jia Qi Yip