English
Related papers

Related papers: Archon: A Unified Multimodal Model for Holistic Di…

200 papers

In this work, we propose Mutual Forcing, a framework for fast autoregressive audio-video generation with long-horizon audio-video synchronization. Our approach addresses two key challenges: joint audio-video modeling and fast autoregressive…

Computer Vision and Pattern Recognition · Computer Science 2026-04-29 Yupeng Zhou , Lianghua Huang , Zhifan Wu , Jiabao Wang , Yupeng Shi , Biao Jiang , Daquan Zhou , Yu Liu , Ming-Ming Cheng , Qibin Hou

The rising demand for creating lifelike avatars in the digital realm has led to an increased need for generating high-quality human videos guided by textual descriptions and poses. We propose Dancing Avatar, designed to fabricate human…

Computer Vision and Pattern Recognition · Computer Science 2023-08-16 Bosheng Qin , Wentao Ye , Qifan Yu , Siliang Tang , Yueting Zhuang

Real-world multimodal applications often require any-to-any capabilities, enabling both understanding and generation across modalities including text, image, audio, and video. However, integrating the strengths of autoregressive language…

Machine Learning · Computer Science 2025-08-15 Jiulin Li , Ping Huang , Yexin Li , Shuo Chen , Juewen Hu , Ye Tian

We present UniMotion, to our knowledge the first unified framework for simultaneous understanding and generation of human motion, natural language, and RGB images within a single architecture. Existing unified models handle only restricted…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Ziyi Wang , Xinshun Wang , Shuang Chen , Yang Cong , Mengyuan Liu

Recent text-to-image models produce high-quality results but still struggle with precise visual control, balancing multimodal inputs, and requiring extensive training for complex multimodal image generation. To address these limitations, we…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Haozhe Zhao , Zefan Cai , Shuzheng Si , Liang Chen , Jiuxiang Gu , Wen Xiao , Minjia Zhang , Junjie Hu

Synthesizing high-fidelity head avatars is a central problem for computer vision and graphics. While head avatar synthesis algorithms have advanced rapidly, the best ones still face great obstacles in real-world scenarios. One of the vital…

Computer Vision and Pattern Recognition · Computer Science 2023-05-24 Dongwei Pan , Long Zhuo , Jingtan Piao , Huiwen Luo , Wei Cheng , Yuxin Wang , Siming Fan , Shengqi Liu , Lei Yang , Bo Dai , Ziwei Liu , Chen Change Loy , Chen Qian , Wayne Wu , Dahua Lin , Kwan-Yee Lin

The recently developed discrete diffusion models perform extraordinarily well in the text-to-image task, showing significant promise for handling the multi-modality signals. In this work, we harness these traits and present a unified…

Computer Vision and Pattern Recognition · Computer Science 2022-11-29 Minghui Hu , Chuanxia Zheng , Heliang Zheng , Tat-Jen Cham , Chaoyue Wang , Zuopeng Yang , Dacheng Tao , Ponnuthurai N. Suganthan

Diffusion models have shown impressive performance in many visual generation and manipulation tasks. Many existing methods focus on training a model for a specific task, especially, text-to-video (T2V) generation, while many other works…

Computer Vision and Pattern Recognition · Computer Science 2026-02-06 Ruibin Li , Tao Yang , Yangming Shi , Weiguo Feng , Shilei Wen , Bingyue Peng , Lei Zhang

World models aim to endow AI systems with the ability to represent, generate, and interact with dynamic environments in a coherent and temporally consistent manner. While recent video generation models have demonstrated impressive visual…

Despite the remarkable process of talking-head-based avatar-creating solutions, directly generating anchor-style videos with full-body motions remains challenging. In this study, we propose Make-Your-Anchor, a novel system necessitating…

Computer Vision and Pattern Recognition · Computer Science 2024-03-26 Ziyao Huang , Fan Tang , Yong Zhang , Xiaodong Cun , Juan Cao , Jintao Li , Tong-Yee Lee

Recent advances in omni-modal large language models have enabled remarkable progress in joint vision-audio understanding. However, prevailing architectures rely on modality-specific encoders with a \emph{video-coarse, audio-dense} design --…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Detao Bai , Shimin Yao , Weixuan Chen , Chengen Lai , Yuanming Li , Zhiheng Ma , Xihan Wei

We introduce a novel framework for 3D human avatar generation and personalization, leveraging text prompts to enhance user engagement and customization. Central to our approach are key innovations aimed at overcoming the challenges in…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Armand Comas-Massagué , Di Qiu , Menglei Chai , Marcel Bühler , Amit Raj , Ruiqi Gao , Qiangeng Xu , Mark Matthews , Paulo Gotardo , Octavia Camps , Sergio Orts-Escolano , Thabo Beeler

Human video generation is a dynamic and rapidly evolving task that aims to synthesize 2D human body video sequences with generative models given control conditions such as text, audio, and pose. With the potential for wide-ranging…

Computer Vision and Pattern Recognition · Computer Science 2024-07-12 Wentao Lei , Jinting Wang , Fengji Ma , Guanjie Huang , Li Liu

In this paper, we propose ARCH (Animatable Reconstruction of Clothed Humans), a novel end-to-end framework for accurate reconstruction of animation-ready 3D clothed humans from a monocular image. Existing approaches to digitize 3D humans…

Graphics · Computer Science 2020-04-14 Zeng Huang , Yuanlu Xu , Christoph Lassner , Hao Li , Tony Tung

We present a novel approach for generating 360-degree high-quality, spatio-temporally coherent human videos from a single image. Our framework combines the strengths of diffusion transformers for capturing global correlations across…

Computer Vision and Pattern Recognition · Computer Science 2024-09-25 Ruizhi Shao , Youxin Pang , Zerong Zheng , Jingxiang Sun , Yebin Liu

Human-Centric Video Generation (HCVG) methods seek to synthesize human videos from multimodal inputs, including text, image, and audio. Existing methods struggle to effectively coordinate these heterogeneous modalities due to two…

Computer Vision and Pattern Recognition · Computer Science 2025-09-11 Liyang Chen , Tianxiang Ma , Jiawei Liu , Bingchuan Li , Zhuowei Chen , Lijie Liu , Xu He , Gen Li , Qian He , Zhiyong Wu

Addressing missing modalities presents a critical challenge in multimodal learning. Current approaches focus on developing models that can handle modality-incomplete inputs during inference, assuming that the full set of modalities are…

Computer Vision and Pattern Recognition · Computer Science 2024-06-05 Yunpeng Zhao , Cheng Chen , Qing You Pang , Quanzheng Li , Carol Tang , Beng-Ti Ang , Yueming Jin

Contemporary Video Object Segmentation (VOS) approaches typically consist stages of feature extraction, matching, memory management, and multiple objects aggregation. Recent advanced models either employ a discrete modeling for these…

Computer Vision and Pattern Recognition · Computer Science 2024-03-14 Wanyun Li , Pinxue Guo , Xinyu Zhou , Lingyi Hong , Yangji He , Xiangyu Zheng , Wei Zhang , Wenqiang Zhang

Visual question answering by using information from multiple modalities has attracted more and more attention in recent years. However, it is a very challenging task, as the visual content and natural language have quite different…

Computer Vision and Pattern Recognition · Computer Science 2020-03-13 Zhaoquan Yuan , Siyuan Sun , Lixin Duan , Xiao Wu , Changsheng Xu

Talking face generation is a novel and challenging generation task, aiming at synthesizing a vivid speaking-face video given a specific audio. To fulfill emotion-controllable talking face generation, current methods need to overcome two…

Computer Vision and Pattern Recognition · Computer Science 2025-08-21 Ziqi Zhang , Cheng Deng