中文
相关论文

相关论文: Point-to-Point: Sparse Motion Guidance for Control…

200 篇论文

We introduce a novel text-to-pose video editing method, ReimaginedAct. While existing video editing tasks are limited to changes in attributes, backgrounds, and styles, our method aims to predict open-ended human action changes in video.…

计算机视觉与模式识别 · 计算机科学 2024-03-13 Lan Wang , Vishnu Boddeti , Sernam Lim

Generative methods for image and video editing use generative models as priors to perform edits despite incomplete information, such as changing the composition of 3D objects shown in a single image. Recent methods have shown promising…

计算机视觉与模式识别 · 计算机科学 2025-03-28 Juil Koo , Paul Guerrero , Chun-Hao Paul Huang , Duygu Ceylan , Minhyuk Sung

To serve the intricate and varied demands of image editing, precise and flexible manipulation in image content is indispensable. Recently, Drag-based editing methods have gained impressive performance. However, these methods predominantly…

计算机视觉与模式识别 · 计算机科学 2024-08-06 Pengyang Ling , Lin Chen , Pan Zhang , Huaian Chen , Yi Jin , Jinjin Zheng

Visual autoregressive (VAR) models have recently emerged as a promising family of generative models, enabling a wide range of downstream vision tasks such as text-guided image editing. By shifting the editing paradigm from noise…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Tao Xia , Jiawei Liu , Yukun Zhang , Ting Liu , Wei Wang , Lei Zhang

Video representation learning has seen tremendous progress in recent years. This has been driven by many factors, including the scale of training and the success of visual models trained contrastively with language. While these factors have…

计算机视觉与模式识别 · 计算机科学 2026-05-25 Mantas Skackauskas , Xinyue Hao , Laura Sevilla-Lara

There has been significant progresses for image object detection in recent years. Nevertheless, video object detection has received little attention, although it is more challenging and more important in practical scenarios. Built upon the…

计算机视觉与模式识别 · 计算机科学 2017-12-01 Xizhou Zhu , Jifeng Dai , Lu Yuan , Yichen Wei

Multimodal pre-training demonstrates strong generalization performance, but this paradigm is often impractical in domains where paired data are scarce. A promising alternative is post-hoc multimodal alignment, which aligns separately…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Shiwon Kim , Yu Rang Park

This paper addresses the long-standing challenge of reconstructing 3D structures from videos with dynamic content. Current approaches to this problem were not designed to operate on casual videos recorded by standard cameras or require a…

计算机视觉与模式识别 · 计算机科学 2024-06-28 Yoni Kasten , Wuyue Lu , Haggai Maron

While Transformers have rapidly gained popularity in various computer vision applications, post-hoc explanations of their internal mechanisms remain largely unexplored. Vision Transformers extract visual information by representing image…

计算机视觉与模式识别 · 计算机科学 2024-03-22 Junyi Wu , Bin Duan , Weitai Kang , Hao Tang , Yan Yan

Recent years have seen substantial progress in diffusion-based controllable video generation. However, achieving precise control in complex scenarios, including fine-grained object parts, sophisticated motion trajectories, and coherent…

计算机视觉与模式识别 · 计算机科学 2025-01-07 Haitao Zhou , Chuang Wang , Rui Nie , Jinlin Liu , Dongdong Yu , Qian Yu , Changhu Wang

Motion-centric video editing remains difficult for large generative video models, which often respond well to appearance changes but struggle to produce specific, localized actions or state transitions in an existing clip. We introduce…

Point tracking is a challenging task in computer vision, aiming to establish point-wise correspondence across long video sequences. Recent advancements have primarily focused on temporal modeling techniques to improve local feature…

计算机视觉与模式识别 · 计算机科学 2024-07-31 Zikai Song , Ying Tang , Run Luo , Lintao Ma , Junqing Yu , Yi-Ping Phoebe Chen , Wei Yang

Tracking Any Point (TAP) in a video is a challenging computer vision problem with many demonstrated applications in robotics, video editing, and 3D reconstruction. Existing methods for TAP rely heavily on complex tracking-specific inductive…

计算机视觉与模式识别 · 计算机科学 2025-04-15 Artem Zholus , Carl Doersch , Yi Yang , Skanda Koppula , Viorica Patraucean , Xu Owen He , Ignacio Rocco , Mehdi S. M. Sajjadi , Sarath Chandar , Ross Goroshin

Video-to-video synthesis poses significant challenges in maintaining character consistency, smooth temporal transitions, and preserving visual quality during fast motion. While recent fully cross-frame self-attention mechanisms have…

计算机视觉与模式识别 · 计算机科学 2024-11-12 Tanvir Mahmud , Mustafa Munir , Radu Marculescu , Diana Marculescu

Machine learning is transforming the video editing industry. Recent advances in computer vision have leveled-up video editing tasks such as intelligent reframing, rotoscoping, color grading, or applying digital makeups. However, most of the…

计算机视觉与模式识别 · 计算机科学 2022-07-22 Dawit Mureja Argaw , Fabian Caba Heilbron , Joon-Young Lee , Markus Woodson , In So Kweon

Multi-person total motion capture is extremely challenging when it comes to handle severe occlusions, different reconstruction granularities from body to face and hands, drastically changing observation scales and fast body movements. To…

计算机视觉与模式识别 · 计算机科学 2021-08-25 Yuxiang Zhang , Zhe Li , Liang An , Mengcheng Li , Tao Yu , Yebin Liu

Learning robotic manipulation from human videos is a promising solution to the data bottleneck in robotics, but the distribution shift between humans and robots remains a critical challenge. Existing approaches often produce entangled…

机器人学 · 计算机科学 2026-05-06 Zhiyuan Li , Wenyan Yang , Wenshuai Zhao , Yue Ma , Yuanpeng Tu , Pekka Marttinen , Joni Pajarinen

Recent advances in diffusion-based text-to-video models, particularly those built on the diffusion transformer architecture, have achieved remarkable progress in generating high-quality and temporally coherent videos. However, transferring…

计算机视觉与模式识别 · 计算机科学 2026-04-08 Zhexin Zhang , Yangyang Xu , Yifeng Zhu , Long Chen , Yong Du , Shengfeng He , Jun Yu

In this paper, we tackle the problem of human motion transfer, where we synthesize novel motion video for a target person that imitates the movement from a reference video. It is a video-to-video translation task in which the estimated…

计算机视觉与模式识别 · 计算机科学 2020-04-08 Jian Ren , Menglei Chai , Sergey Tulyakov , Chen Fang , Xiaohui Shen , Jianchao Yang

This paper improves upon the Pix2Seq object detector by extending it for videos. In the process, it introduces a new way to perform end-to-end video object detection that improves upon existing video detectors in two key ways. First, by…

计算机视觉与模式识别 · 计算机科学 2025-08-21 Abhineet Singh , Nilanjan Ray