中文
相关论文

相关论文: OmniCamera: A Unified Framework for Multi-task Vid…

200 篇论文

Filmmaking and animation production often require sophisticated techniques for coordinating camera transitions and object movements, typically involving labor-intensive real-world capturing. Despite advancements in generative AI for video…

计算机视觉与模式识别 · 计算机科学 2024-06-24 Yaowei Li , Xintao Wang , Zhaoyang Zhang , Zhouxia Wang , Ziyang Yuan , Liangbin Xie , Yuexian Zou , Ying Shan

Unified understanding and generation is a highly appealing research direction in multimodal learning. There exist two approaches: one trains a transformer via an auto-regressive paradigm, and the other adopts a two-stage scheme connecting…

计算机视觉与模式识别 · 计算机科学 2025-12-09 Shihao Zhao , Yitong Chen , Zeyinzi Jiang , Bojia Zi , Shaozhe Hao , Yu Liu , Chaojie Mao , Kwan-Yee K. Wong

Face forgery detection encompasses multiple critical tasks, including identifying forged images and videos and localizing manipulated regions and temporal segments. Current approaches typically employ task-specific models with independent…

计算机视觉与模式识别 · 计算机科学 2025-12-02 Haotian Liu , Haoyu Chen , Chenhui Pan , You Hu , Guoying Zhao , Xiaobai Li

High-fidelity motion tracking serves as the ultimate litmus test for generalizable, human-level motor skills. However, current policies often hit a "generality barrier": as motion libraries scale in diversity, tracking fidelity inevitably…

We present UniModel, a unified generative model that jointly supports visual understanding and visual generation within a single pixel-to-pixel diffusion framework. Our goal is to achieve unification along three axes: the model, the tasks,…

计算机视觉与模式识别 · 计算机科学 2025-11-24 Chi Zhang , Jiepeng Wang , Youming Wang , Yuanzhi Liang , Xiaoyan Yang , Zuoxin Li , Haibin Huang , Xuelong Li

Video Generation is a relatively new and yet popular subject in machine learning due to its vast variety of potential applications and its numerous challenges. Current methods in Video Generation provide the user with little or no control…

计算机视觉与模式识别 · 计算机科学 2021-11-22 Bahman Rouhani , Mohammad Rahmati

The emergence of Diffusion Transformers (DiT) has brought significant advancements to video generation, especially in text-to-video and image-to-video tasks. Although video generation is widely applied in various fields, most existing…

计算机视觉与模式识别 · 计算机科学 2025-06-03 Sen Liang , Zhentao Yu , Zhengguang Zhou , Teng Hu , Hongmei Wang , Yi Chen , Qin Lin , Yuan Zhou , Xin Li , Qinglin Lu , Zhibo Chen

We present a novel approach named OmniControl for incorporating flexible spatial control signals into a text-conditioned human motion generation model based on the diffusion process. Unlike previous methods that can only control the pelvis…

计算机视觉与模式识别 · 计算机科学 2024-04-16 Yiming Xie , Varun Jampani , Lei Zhong , Deqing Sun , Huaizu Jiang

Generating visual instructions in a given context is essential for developing interactive world simulators. While prior works address this problem through either text-guided image manipulation or video prediction, these tasks are typically…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Yujiang Pu , Zhanbo Huang , Vishnu Boddeti , Yu Kong

We introduce OneDiffusion, a versatile, large-scale diffusion model that seamlessly supports bidirectional image synthesis and understanding across diverse tasks. It enables conditional generation from inputs such as text, depth, pose,…

计算机视觉与模式识别 · 计算机科学 2025-06-16 Duong H. Le , Tuan Pham , Sangho Lee , Christopher Clark , Aniruddha Kembhavi , Stephan Mandt , Ranjay Krishna , Jiasen Lu

Achieving machine autonomy and human control often represent divergent objectives in the design of interactive AI systems. Visual generative foundation models such as Stable Diffusion show promise in navigating these goals, especially when…

计算机视觉与模式识别 · 计算机科学 2023-11-03 Can Qin , Shu Zhang , Ning Yu , Yihao Feng , Xinyi Yang , Yingbo Zhou , Huan Wang , Juan Carlos Niebles , Caiming Xiong , Silvio Savarese , Stefano Ermon , Yun Fu , Ran Xu

Achieving precise camera control in video generation remains challenging, as existing methods often rely on camera pose annotations that are difficult to scale to large and dynamic datasets and are frequently inconsistent with depth…

计算机视觉与模式识别 · 计算机科学 2025-12-23 Zelin Zhao , Xinyu Gong , Bangya Liu , Ziyang Song , Jun Zhang , Suhui Wu , Yongxin Chen , Hao Zhang

The field of text-to-image (T2I) generation has made significant progress in recent years, largely driven by advancements in diffusion models. Linguistic control enables effective content creation, but struggles with fine-grained control…

计算机视觉与模式识别 · 计算机科学 2024-07-19 Yanan Sun , Yanchen Liu , Yinhao Tang , Wenjie Pei , Kai Chen

Video Diffusion Models have been developed for video generation, usually integrating text and image conditioning to enhance control over the generated content. Despite the progress, ensuring consistency across frames remains a challenge,…

计算机视觉与模式识别 · 计算机科学 2024-11-12 Tian Xia , Xuweiyi Chen , Sihan Xu

Recent advances in text-to-image generation with diffusion models present transformative capabilities in image quality. However, user controllability of the generated image, and fast adaptation to new tasks still remains an open challenge,…

计算机视觉与模式识别 · 计算机科学 2023-02-17 Omer Bar-Tal , Lior Yariv , Yaron Lipman , Tali Dekel

Instruction-guided image editing methods have demonstrated significant potential by training diffusion models on automatically synthesized or manually annotated image editing pairs. However, these methods remain far from practical,…

计算机视觉与模式识别 · 计算机科学 2025-04-29 Cong Wei , Zheyang Xiong , Weiming Ren , Xinrun Du , Ge Zhang , Wenhu Chen

Federated learning (FL) has become a promising paradigm for collaborative medical image analysis, yet existing frameworks remain tightly coupled to task-specific backbones and are fragile under heterogeneous imaging modalities. Such…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Meilin Liu , Jiaying Wang , Jing Shan

We tackle the dual challenges of video understanding and controllable video generation within a unified diffusion framework. Our key insights are two-fold: geometry-only cues (e.g., depth, edges) are insufficient: they specify layout but…

计算机视觉与模式识别 · 计算机科学 2025-11-27 Dianbing Xi , Jiepeng Wang , Yuanzhi Liang , Xi Qiu , Jialun Liu , Hao Pan , Yuchi Huo , Rui Wang , Haibin Huang , Chi Zhang , Xuelong Li

Despite tremendous recent progress, generative video models still struggle to capture real-world motion, dynamics, and physics. We show that this limitation arises from the conventional pixel reconstruction objective, which biases models…

计算机视觉与模式识别 · 计算机科学 2025-05-27 Hila Chefer , Uriel Singer , Amit Zohar , Yuval Kirstain , Adam Polyak , Yaniv Taigman , Lior Wolf , Shelly Sheynin

Existing mainstream video customization methods focus on generating identity-consistent videos based on given reference images and textual prompts. Benefiting from the rapid advancement of joint audio-video generation, this paper proposes a…

声音 · 计算机科学 2026-05-29 Maomao Li , Zhen Li , Kaipeng Zhang , Guosheng Yin , Zhifeng Li , Dong Xu