中文
相关论文

相关论文: MultiActor-Audiobook: Zero-Shot Audiobook Generati…

200 篇论文

Audiobook generation aims to create rich, immersive listening experiences from multimodal inputs, but current approaches face three critical challenges: (1) the lack of synergistic generation of diverse audio types (e.g., speech, sound…

声音 · 计算机科学 2025-08-13 Yan Rong , Shan Yang , Chenxing Li , Dong Yu , Li Liu

Existing text-to-speech systems predominantly focus on single-sentence synthesis and lack adequate contextual modeling as well as fine-grained performance control capabilities for generating coherent multicast audiobooks. To address these…

音频与语音处理 · 电气工程与系统科学 2025-09-23 Min Liu , JingJing Yin , Xiang Zhang , Siyu Hao , Yanni Hu , Bin Lin , Yuan Feng , Hongbin Zhou , Jianhao Ye

The rapid advancement of large language models (LLMs) and artificial intelligence-generated content (AIGC) has accelerated AI-native applications, such as AI-based storybooks that automate engaging story production for children. However,…

计算与语言 · 计算机科学 2025-03-10 Xuenan Xu , Jiahao Mei , Chenliang Li , Yuning Wu , Ming Yan , Shaopeng Lai , Ji Zhang , Mengyue Wu

Multimodality-to-Multiaudio (MM2MA) generation faces significant challenges in synthesizing diverse and contextually aligned audio types (e.g., sound effects, speech, music, and songs) from multimodal inputs (e.g., video, text, images),…

声音 · 计算机科学 2025-08-06 Yan Rong , Jinting Wang , Guangzhi Lei , Shan Yang , Li Liu

Generating natural-sounding, multi-speaker dialogue is crucial for applications such as podcast creation, virtual agents, and multimedia content generation. However, existing systems struggle to maintain speaker consistency, model…

We introduce Audio-Agent, a multimodal framework for audio generation, editing and composition based on text or video inputs. Conventional approaches for text-to-audio (TTA) tasks often make single-pass inferences from text descriptions.…

声音 · 计算机科学 2025-01-15 Zixuan Wang , Chi-Keung Tang , Yu-Wing Tai

This research introduces an innovative AI-driven multi-agent framework specifically designed for creating immersive audiobooks. Leveraging neural text-to-speech synthesis with FastSpeech 2 and VALL-E for expressive narration and…

声音 · 计算机科学 2025-05-09 Shaja Arul Selvamani , Nia D'Souza Ganapathy

An audiobook can dramatically improve a work of literature's accessibility and improve reader engagement. However, audiobooks can take hundreds of hours of human effort to create, edit, and publish. In this work, we present a system that…

Large-scale multimodal generative modeling has created milestones in text-to-image and text-to-video generation. Its application to audio still lags behind for two main reasons: the lack of large-scale datasets with high-quality text-audio…

While modern Text-to-Speech (TTS) systems achieve high fidelity for read-style speech, they struggle to generate Autonomous Sensory Meridian Response (ASMR), a specialized, low-intensity speech style essential for relaxation. The inherent…

声音 · 计算机科学 2026-01-23 Leying Zhang , Tingxiao Zhou , Haiyang Sun , Mengxiao Bi , Yanmin Qian

Recognizing characters and predicting speakers of dialogue are critical for comic processing tasks, such as voice generation or translation. However, because characters vary by comic title, supervised learning approaches like training…

多媒体 · 计算机科学 2024-09-06 Yingxuan Li , Ryota Hinami , Kiyoharu Aizawa , Yusuke Matsui

Existing Existing automatic audio generation methods struggle to generate podcast-like audio programs effectively. The key challenges lie in in-depth content generation, appropriate and expressive voice production. This paper proposed…

声音 · 计算机科学 2025-03-04 Yujia Xiao , Lei He , Haohan Guo , Fenglong Xie , Tan Lee

Joint audio-video generation aims to synthesize synchronized multisensory content, yet current unified models struggle with fine-grained acoustic control, particularly for identity-preserving speech. Existing approaches either suffer from…

声音 · 计算机科学 2026-01-09 Chunyu Qiang , Jun Wang , Xiaopeng Wang , Kang Yin , Yuxin Guo

The pipeline for multi-participant audiobook production primarily consists of three stages: script analysis, character voice timbre selection, and speech synthesis. Among these, script analysis can be automated with high accuracy using NLP…

音频与语音处理 · 电气工程与系统科学 2025-09-22 Ziqi Dai , Yiting Chen , Jiacheng Xu , Liufei Xie , Yuchen Wang , Zhenchuan Yang , Bingsong Bai , Yangsheng Gao , Wenjiang Zhou , Weifeng Zhao , Ruohua Zhou

Recent advancements in zero-shot text-to-speech (TTS) modeling have led to significant strides in generating high-fidelity and diverse speech. However, dialogue generation, along with achieving human-like naturalness in speech, continues to…

音频与语音处理 · 电气工程与系统科学 2024-12-17 Leying Zhang , Yao Qian , Long Zhou , Shujie Liu , Dongmei Wang , Xiaofei Wang , Midia Yousefi , Yanmin Qian , Jinyu Li , Lei He , Sheng Zhao , Michael Zeng

Audio-driven human animation methods, such as talking head and talking body generation, have made remarkable progress in generating synchronized facial movements and appealing visual quality videos. However, existing methods primarily focus…

计算机视觉与模式识别 · 计算机科学 2025-05-29 Zhe Kong , Feng Gao , Yong Zhang , Zhuoliang Kang , Xiaoming Wei , Xunliang Cai , Guanying Chen , Wenhan Luo

With the advent of the big data and large language model era, zero-shot personalized rapid customization has emerged as a significant trend. In this report, we introduce Takin AudioLLM, a series of techniques and models, mainly including…

Language model (LM) based audio generation frameworks, e.g., AudioLM, have recently achieved new state-of-the-art performance in zero-shot audio generation. In this paper, we explore the feasibility of LMs for zero-shot voice conversion. An…

音频与语音处理 · 电气工程与系统科学 2023-08-22 Zhichao Wang , Yuanzhe Chen , Lei Xie , Qiao Tian , Yuping Wang

The task of audio captioning is similar in essence to tasks such as image and video captioning. However, it has received much less attention. We propose three desiderata for captioning audio -- (i) fluency of the generated text, (ii)…

声音 · 计算机科学 2023-09-08 Tal Shaharabany , Ariel Shaulov , Lior Wolf

Audiobook interpretations are attracting increasing attention, as they provide accessible and in-depth analyses of books that offer readers practical insights and intellectual inspiration. However, their manual creation process remains…

计算与语言 · 计算机科学 2025-12-30 Minjiang Huang , Jipeng Qiang , Yi Zhu , Chaowei Zhang , Xiangyu Zhao , Kui Yu
‹ 上一页 1 2 3 10 下一页 ›