中文
相关论文

相关论文: MechVerse: Evaluating Physical Motion Consistency …

200 篇论文

State-of-the-art video generation models produce remarkable photorealism, but they lack the precise control required to align generated content with specific scene requirements. Furthermore, without an underlying explicit geometry, these…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Dana Cohen-Bar , Ido Sobol , Raphael Bensadoun , Shelly Sheynin , Oran Gafni , Or Patashnik , Daniel Cohen-Or , Amit Zohar

Recent advancements in video generation have significantly improved the ability to synthesize videos from text instructions. However, existing models still struggle with key challenges such as instruction misalignment, content…

计算机视觉与模式识别 · 计算机科学 2025-02-10 Haibo Tong , Zhaoyang Wang , Zhaorun Chen , Haonian Ji , Shi Qiu , Siwei Han , Kexin Geng , Zhongkai Xue , Yiyang Zhou , Peng Xia , Mingyu Ding , Rafael Rafailov , Chelsea Finn , Huaxiu Yao

Evaluating text-to-image generative models remains a challenge, despite the remarkable progress being made in their overall performances. While existing metrics like CLIPScore work for coarse evaluations, they lack the sensitivity to…

计算机视觉与模式识别 · 计算机科学 2024-11-06 Georgia Gabriela Sampaio , Ruixiang Zhang , Shuangfei Zhai , Jiatao Gu , Josh Susskind , Navdeep Jaitly , Yizhe Zhang

Recent advances in text-to-video generation have produced increasingly realistic and diverse content, yet evaluating such videos remains a fundamental challenge due to their multi-faceted nature encompassing visual quality, semantic…

Generating video stories from text prompts is a complex task. In addition to having high visual quality, videos need to realistically adhere to a sequence of text prompts whilst being consistent throughout the frames. Creating a benchmark…

Co-speech gestures, if presented in the lively form of videos, can achieve superior visual effects in human-machine interaction. While previous works mostly generate structural human skeletons, resulting in the omission of appearance…

计算机视觉与模式识别 · 计算机科学 2024-04-03 Xu He , Qiaochu Huang , Zhensong Zhang , Zhiwei Lin , Zhiyong Wu , Sicheng Yang , Minglei Li , Zhiyi Chen , Songcen Xu , Xiaofei Wu

Driving video generation has achieved much progress in controllability, video resolution, and length, but fails to support fine-grained object-level controllability for diverse driving videos, while preserving the spatiotemporal…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Li-Heng Chen , Ke Cheng , Yahui Liu , Lei Shi , Shi-Sheng Huang , Hongbo Fu

Text-to-video and image-to-video generation have made rapid progress in visual quality, but they remain limited in controlling the precise timing of motion. In contrast, audio provides temporal cues aligned with video motion, making it a…

计算机视觉与模式识别 · 计算机科学 2025-09-29 Jibin Song , Mingi Kwon , Jaeseok Jeong , Youngjung Uh

Synthesizing high-fidelity videos from real-world multi-view input is challenging because of the complexities of real-world environments and highly dynamic motions. Previous works based on neural radiance fields have demonstrated…

计算机视觉与模式识别 · 计算机科学 2023-10-10 Feng Wang , Sinan Tan , Xinghang Li , Zeyue Tian , Yafei Song , Huaping Liu

Recent advances in text-to-video generation have harnessed the power of diffusion models to create visually compelling content conditioned on text prompts. However, they usually encounter high computational costs and often struggle to…

计算机视觉与模式识别 · 计算机科学 2024-04-24 Jiaxi Lv , Yi Huang , Mingfu Yan , Jiancheng Huang , Jianzhuang Liu , Yifan Liu , Yafei Wen , Xiaoxin Chen , Shifeng Chen

Video generation models have made significant progress in generating realistic content, enabling applications in simulation, gaming, and film making. However, current generated videos still contain visual artifacts arising from 3D…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Duolikun Danier , Ge Gao , Steven McDonagh , Changjian Li , Hakan Bilen , Oisin Mac Aodha

Synthesizing motion-rich and temporally consistent videos remains a challenge in artificial intelligence, especially when dealing with extended durations. Existing text-to-video (T2V) models commonly employ spatial cross-attention for text…

计算机视觉与模式识别 · 计算机科学 2025-08-18 Jiasong Feng , Ao Ma , Jing Wang , Ke Cao , Zhanjie Zhang

Many application areas ranging from serious games for health to learning by demonstration in robotics, could benefit from large body movement datasets extracted from textual instructions accompanied by images. The interpretation of…

人机交互 · 计算机科学 2020-06-09 Himangshu Sarma , Robert Porzel , Jan Smeddinck , Rainer Malaka

Text-to-video diffusion models have enabled high-quality video synthesis, yet often fail to generate temporally coherent and physically plausible motion. A key reason is the models' insufficient understanding of complex motions that natural…

计算机视觉与模式识别 · 计算机科学 2025-10-23 Aritra Bhowmik , Denis Korzhenkov , Cees G. M. Snoek , Amirhossein Habibian , Mohsen Ghafoorian

Reward-based fine-tuning of video diffusion models is an effective approach to improve the quality of generated videos, as it can fine-tune models without requiring real-world video datasets. However, it can sometimes be limited to specific…

计算机视觉与模式识别 · 计算机科学 2025-10-24 Takehiro Aoshima , Yusuke Shinohara , Byeongseon Park

Research on video generation has recently made tremendous progress, enabling high-quality videos to be generated from text prompts or images. Adding control to the video generation process is an important goal moving forward and recent…

计算机视觉与模式识别 · 计算机科学 2024-05-28 Zhengfei Kuang , Shengqu Cai , Hao He , Yinghao Xu , Hongsheng Li , Leonidas Guibas , Gordon Wetzstein

We present I2V3D, a novel framework for animating static images into dynamic videos with precise 3D control, leveraging the strengths of both 3D geometry guidance and advanced generative models. Our approach combines the precision of a…

计算机视觉与模式识别 · 计算机科学 2025-03-14 Zhiyuan Zhang , Dongdong Chen , Jing Liao

We introduce \textit{WonderVerse}, a simple but effective framework for generating extendable 3D scenes. Unlike existing methods that rely on iterative depth estimation and image inpainting, often leading to geometric distortions and…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Hao Feng , Zhi Zuo , Jia-Hui Pan , Ka-Hei Hui , Qi Dou , Jingyu Hu , Zhengzhe Liu

Although humans have the innate ability to imagine multiple possible actions from videos, it remains an extraordinary challenge for computers due to the intricate camera movements and montages. Most existing motion generation methods…

计算机视觉与模式识别 · 计算机科学 2024-08-14 Liangdong Qiu , Chengxing Yu , Yanran Li , Zhao Wang , Haibin Huang , Chongyang Ma , Di Zhang , Pengfei Wan , Xiaoguang Han

Video generation has recently made striking visual progress, but maintaining coherent object motion and interactions remains difficult. We trace two practical bottlenecks: (i) human-provided motion hints (e.g., small 2D maps) often collapse…

计算机视觉与模式识别 · 计算机科学 2025-10-21 Zhifei Chen , Tianshuo Xu , Leyi Wu , Luozhou Wang , Dongyu Yan , Zihan You , Wenting Luo , Guo Zhang , Yingcong Chen