中文
相关论文

相关论文: AlignDiT: Multimodal Aligned Diffusion Transformer…

200 篇论文

We present VoiceDiT, a multi-modal generative model for producing environment-aware speech and audio from text and visual prompts. While aligning speech with text is crucial for intelligible speech, achieving this alignment in noisy…

音频与语音处理 · 电气工程与系统科学 2024-12-30 Jaemin Jung , Junseok Ahn , Chaeyoung Jung , Tan Dat Nguyen , Youngjoon Jang , Joon Son Chung

Text-to-video (T2V) diffusion models have recently achieved impressive visual quality, yet most systems still generate silent clips and treat audio as a secondary concern. Existing audio-video generation pipelines typically decompose the…

Movie dubbing aims to synthesize speech that preserves the vocal identity of a reference audio while synchronizing with the lip movements in a target video. Existing methods fail to achieve precise lip-sync and lack naturalness due to…

声音 · 计算机科学 2026-04-15 Gaoxiang Cong , Liang Li , Jiaxin Ye , Zhedong Zhang , Hongming Shan , Yuankai Qi , Qingming Huang

Audio-driven talking head generation is critical for applications such as virtual assistants, video games, and films, where natural lip movements are essential. Despite progress in this field, challenges remain in producing both consistent…

计算机视觉与模式识别 · 计算机科学 2025-07-08 Yucheng Wang , Dan Xu

Generating realistic conversational gestures are essential for achieving natural, socially engaging interactions with digital humans. However, existing methods typically map a single audio stream to a single speaker's motion, without…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Yichen Peng , Jyun-Ting Song , Siyeol Jung , Ruofan Liu , Haiyang Liu , Xuangeng Chu , Ruicong Liu , Erwin Wu , Hideki Koike , Kris Kitani

Sounding Video Generation (SVG) remains a challenging task due to the inherent structural misalignment between audio and video, as well as the high computational cost of multimodal data processing. In this paper, we introduce ProAV-DiT, a…

多媒体 · 计算机科学 2025-11-18 Jiahui Sun , Weining Wang , Mingzhen Sun , Yirong Yang , Xinxin Zhu , Jing Liu

We present AudioGen-Omni - a unified approach based on multimodal diffusion transformers (MMDit), capable of generating high-fidelity audio, speech, and song coherently synchronized with the input video. AudioGen-Omni introduces a novel…

声音 · 计算机科学 2025-08-08 Le Wang , Jun Wang , Chunyu Qiang , Feng Deng , Chen Zhang , Di Zhang , Kun Gai

Recent multimodal face generation models address the spatial control limitations of text-to-image diffusion models by augmenting text-based conditioning with spatial priors such as segmentation masks, sketches, or edge maps. This multimodal…

计算机视觉与模式识别 · 计算机科学 2026-04-01 Bharath Krishnamurthy , Ajita Rattani

Recent Diffusion Transformers (DiTs) have shown impressive capabilities in generating high-quality single-modality content, including images, videos, and audio. However, it is still under-explored whether the transformer-based diffuser can…

计算机视觉与模式识别 · 计算机科学 2024-06-13 Kai Wang , Shijian Deng , Jing Shi , Dimitrios Hatzinakos , Yapeng Tian

This paper introduces JavisDiT, a novel Joint Audio-Video Diffusion Transformer designed for synchronized audio-video generation (JAVG). Based on the powerful Diffusion Transformer (DiT) architecture, JavisDiT simultaneously generates…

计算机视觉与模式识别 · 计算机科学 2026-02-24 Kai Liu , Wei Li , Lai Chen , Shengqiong Wu , Yanhao Zheng , Jiayi Ji , Fan Zhou , Jiebo Luo , Ziwei Liu , Hao Fei , Tat-Seng Chua

We propose a novel talking head synthesis pipeline called "DiT-Head", which is based on diffusion transformers and uses audio as a condition to drive the denoising process of a diffusion model. Our method is scalable and can generalise to…

人工智能 · 计算机科学 2023-12-12 Aaron Mir , Eduardo Alonso , Esther Mondragón

The Text-to-Video (T2V) model aims to generate dynamic and expressive videos from textual prompts. The generation pipeline typically involves multiple modules, such as language encoder, Diffusion Transformer (DiT), and Variational…

分布式、并行与集群计算 · 计算机科学 2025-06-17 Heyang Huang , Cunchen Hu , Jiaqi Zhu , Ziyuan Gao , Liangliang Xu , Yizhou Shan , Yungang Bao , Sun Ninghui , Tianwei Zhang , Sa Wang

Whole-body audio-driven avatar pose and expression generation is a critical task for creating lifelike digital humans and enhancing the capabilities of interactive virtual agents, with wide-ranging applications in virtual reality, digital…

声音 · 计算机科学 2025-10-15 Tianbao Zhang , Jian Zhao , Yuer Li , Zheng Zhu , Ping Hu , Zhaoxin Fan , Wenjun Wu , Xuelong Li

Controllable text-to-audio generation aims to synthesize audio from textual descriptions while satisfying user-specified constraints, including event types, temporal sequences, and onset and offset timestamps. This enables precise control…

声音 · 计算机科学 2026-02-10 Yisu Liu , Chenxing Li , Wanqian Zhang , Wenfu Wang , Meng Yu , Ruibo Fu , Zheng Lin , Weiping Wang , Dong Yu

Talking head synthesis, also known as speech-to-lip synthesis, reconstructs the facial motions that align with the given audio tracks. The synthesized videos are evaluated on mainly two aspects, lip-speech synchronization and image…

机器学习 · 计算机科学 2025-03-18 Xulin Fan , Heting Gao , Ziyi Chen , Peng Chang , Mei Han , Mark Hasegawa-Johnson

We propose Kling-Foley, a large-scale multimodal Video-to-Audio generation model that synthesizes high-quality audio synchronized with video content. In Kling-Foley, we introduce multimodal diffusion transformers to model the interactions…

The intrinsic link between facial motion and speech is often overlooked in generative modeling, where talking head synthesis and text-to-speech (TTS) are typically addressed as separate tasks. This paper introduces JAM-Flow, a unified…

计算机视觉与模式识别 · 计算机科学 2026-05-18 Mingi Kwon , Joonghyuk Shin , Jaeseok Jung , Jaesik Park , Youngjung Uh

Talking head synthesis is a promising approach for the video production industry. Recently, a lot of effort has been devoted in this research area to improve the generation quality or enhance the model generalization. However, there are few…

计算机视觉与模式识别 · 计算机科学 2023-04-21 Shuai Shen , Wenliang Zhao , Zibin Meng , Wanhua Li , Zheng Zhu , Jie Zhou , Jiwen Lu

Although significant progress has been made in audio-driven talking head generation, text-driven methods remain underexplored. In this work, we present OmniTalker, a unified framework that jointly generates synchronized talking audio-video…

计算机视觉与模式识别 · 计算机科学 2025-06-04 Zhongjian Wang , Peng Zhang , Jinwei Qi , Guangyuan Wang , Chaonan Ji , Sheng Xu , Bang Zhang , Liefeng Bo

Recent advancements in the field of Diffusion Transformers have substantially improved the generation of high-quality 2D images, 3D videos, and 3D shapes. However, the effectiveness of the Transformer architecture in the domain of co-speech…

计算机视觉与模式识别 · 计算机科学 2024-08-07 Xiaofeng Mao , Zhengkai Jiang , Qilin Wang , Chencan Fu , Jiangning Zhang , Jiafu Wu , Yabiao Wang , Chengjie Wang , Wei Li , Mingmin Chi
‹ 上一页 1 2 3 10 下一页 ›