English
Related papers

Related papers: SMRABooth: Subject and Motion Representation Align…

200 papers

Training-free consistent text-to-image generation depicting the same subjects across different images is a topic of widespread recent interest. Existing works in this direction predominantly rely on cross-frame self-attention; which…

Computer Vision and Pattern Recognition · Computer Science 2025-04-09 Jaskirat Singh , Junshen Kevin Chen , Jonas Kohler , Michael Cohen

Self-supervised learning of visual representations has been focusing on learning content features, which do not capture object motion or location, and focus on identifying and differentiating objects in images and videos. On the other hand,…

Computer Vision and Pattern Recognition · Computer Science 2023-07-25 Adrien Bardes , Jean Ponce , Yann LeCun

Existing video generation models predominantly emphasize appearance fidelity while exhibiting limited ability to synthesize complex human motions, such as whole-body movements, long-range dynamics, and fine-grained human-environment…

Computer Vision and Pattern Recognition · Computer Science 2026-02-25 Haoyu Wang , Hao Tang , Donglin Di , Zhilu Zhang , Wangmeng Zuo , Feng Gao , Siwei Ma , Shiliang Zhang

We introduce AvatarBooth, a novel method for generating high-quality 3D avatars using text prompts or specific images. Unlike previous approaches that can only synthesize avatars based on simple text descriptions, our method enables the…

Computer Vision and Pattern Recognition · Computer Science 2023-06-19 Yifei Zeng , Yuanxun Lu , Xinya Ji , Yao Yao , Hao Zhu , Xun Cao

Image customization has been extensively studied in text-to-image (T2I) diffusion models, leading to impressive outcomes and applications. With the emergence of text-to-video (T2V) diffusion models, its temporal counterpart, motion…

Computer Vision and Pattern Recognition · Computer Science 2024-08-29 Yixuan Ren , Yang Zhou , Jimei Yang , Jing Shi , Difan Liu , Feng Liu , Mingi Kwon , Abhinav Shrivastava

The essence of a video lies in its dynamic motions, including character actions, object movements, and camera movements. While text-to-video generative diffusion models have recently advanced in creating diverse contents, controlling…

Computer Vision and Pattern Recognition · Computer Science 2024-01-04 Yuxin Zhang , Fan Tang , Nisha Huang , Haibin Huang , Chongyang Ma , Weiming Dong , Changsheng Xu

Despite the promising progress in subject-driven image generation, current models often deviate from the reference identities and struggle in complex scenes with multiple subjects. To address this challenge, we introduce OpenSubject, a…

Computer Vision and Pattern Recognition · Computer Science 2025-12-11 Yexin Liu , Manyuan Zhang , Yueze Wang , Hongyu Li , Dian Zheng , Weiming Zhang , Changsheng Lu , Xunliang Cai , Yan Feng , Peng Pei , Harry Yang

A central goal in AI is to represent scenes as compositions of discrete objects, enabling fine-grained, controllable image and video generation. Yet leading diffusion models treat images holistically and rely on text conditioning, creating…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Adil Kaan Akan

This report presents our method for Single Object Tracking (SOT), which aims to track a specified object throughout a video sequence. We employ the LoRAT method. The essence of the work lies in adapting LoRA, a technique that fine-tunes a…

Computer Vision and Pattern Recognition · Computer Science 2024-10-23 Zhiqiang Zhong , Yang Yang , Fengqiang Wan , Henglu Wei , Xiangyang Ji

Recent advancements in personalized Text-to-Video (T2V) generation have made significant strides in synthesizing character-specific content. However, these methods face a critical limitation: the inability to perform fine-grained control…

Computer Vision and Pattern Recognition · Computer Science 2026-01-06 Haopeng Fang , Di Qiu , Binjie Mao , He Tang

The development of Text-to-Video (T2V) generation has made motion transfer possible, enabling the control of video motion based on existing footage. However, current methods have two limitations: 1) struggle to handle multi-subjects videos,…

Computer Vision and Pattern Recognition · Computer Science 2025-04-04 Jiayi Gao , Zijin Yin , Changcheng Hua , Yuxin Peng , Kongming Liang , Zhanyu Ma , Jun Guo , Yang Liu

Rotary Position Embedding (RoPE) has shown strong performance in text-based Large Language Models (LLMs), but extending it to video remains a challenge due to the intricate spatiotemporal structure of video frames. Existing adaptations,…

Artificial Intelligence · Computer Science 2025-11-03 Zikang Liu , Longteng Guo , Yepeng Tang , Tongtian Yue , Junxian Cai , Kai Ma , Qingbin Liu , Xi Chen , Jing Liu

Recently, breakthroughs in the video diffusion transformer have shown remarkable capabilities in diverse motion generations. As for the motion-transfer task, current methods mainly use two-stage Low-Rank Adaptations (LoRAs) finetuning to…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Yue Ma , Yulong Liu , Qiyuan Zhu , Ayden Yang , Kunyu Feng , Xinhua Zhang , Zexuan Yan , Zhifeng Li , Sirui Han , Chenyang Qi , Qifeng Chen

Research on diffusion model-based video generation has advanced rapidly. However, limitations in object fidelity and generation length hinder its practical applications. Additionally, specific domains like animated wallpapers require…

Computer Vision and Pattern Recognition · Computer Science 2024-04-17 Fanyi Wang , Peng Liu , Haotian Hu , Dan Meng , Jingwen Su , Jinjin Xu , Yanhao Zhang , Xiaoming Ren , Zhiwang Zhang

In this paper, we consider the task of unsupervised object discovery in videos. Previous works have shown promising results via processing optical flows to segment objects. However, taking flow as input brings about two drawbacks. First,…

Computer Vision and Pattern Recognition · Computer Science 2022-10-04 Shuangrui Ding , Weidi Xie , Yabo Chen , Rui Qian , Xiaopeng Zhang , Hongkai Xiong , Qi Tian

Recent advancements in image generation models have enabled personalized image creation with both user-defined subjects (content) and styles. Prior works achieved personalization by merging corresponding low-rank adapters (LoRAs) through…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Donald Shenaj , Ondrej Bohdal , Mete Ozay , Pietro Zanuttigh , Umberto Michieli

Multi-object tracking (MOT) is a fundamental task in computer vision that requires continuously tracking multiple targets while maintaining consistent identities across frames. However, most existing approaches primarily rely on…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Yanchao Wang , Dawei Zhang , Chengzhuan Yang , Wei Liu , Minglu Li , Hua Wang , Zhonglong Zheng , Ming-Hsuan Yang

Visual object tracking plays a critical role in visual-based autonomous systems, as it aims to estimate the position and size of the object of interest within a live video. Despite significant progress made in this field, state-of-the-art…

Computer Vision and Pattern Recognition · Computer Science 2024-04-10 Jianlang Chen , Xuhong Ren , Qing Guo , Felix Juefei-Xu , Di Lin , Wei Feng , Lei Ma , Jianjun Zhao

Subject-driven image generation aims to synthesize new images that preserve the identity of the given subject while following textual instructions. Existing approaches often encode text and reference images separately. This limits…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Shuhong Zheng , Aashish Kumar Misraa , Yu-Teng Li , Yu-Jhe Li , Igor Gilitschenski

Storyboard synthesis plays a crucial role in visual storytelling, aiming to generate coherent shot sequences that visually narrate cinematic events with consistent characters, scenes, and transitions. However, existing approaches are mostly…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Junjia Huang , Binbin Yang , Pengxiang Yan , Jiyang Liu , Bin Xia , Zhao Wang , Yitong Wang , Liang Lin , Guanbin Li