English
Related papers

Related papers: SMRABooth: Subject and Motion Representation Align…

200 papers

Recent text-to-video diffusion models have achieved impressive progress. In practice, users often desire the ability to control object motion and camera movement independently for customized video creation. However, current methods lack the…

Computer Vision and Pattern Recognition · Computer Science 2024-05-07 Shiyuan Yang , Liang Hou , Haibin Huang , Chongyang Ma , Pengfei Wan , Di Zhang , Xiaodong Chen , Jing Liao

Generating motion-controlled videos--where user-specified actions drive physically plausible scene dynamics under freely chosen viewpoints--demands two capabilities: (1) disentangled motion control, allowing users to separately control the…

Computer Vision and Pattern Recognition · Computer Science 2026-04-09 Shaowei Liu , Xuanchi Ren , Tianchang Shen , Huan Ling , Saurabh Gupta , Shenlong Wang , Sanja Fidler , Jun Gao

Music is both an auditory and an embodied phenomenon, closely linked to human motion and naturally expressed through dance. However, most existing audio representations neglect this embodied dimension, limiting their ability to capture…

Sound · Computer Science 2026-01-30 Xuanchen Wang , Heng Wang , Weidong Cai

With the rapid advancement of diffusion-based generative models, portrait image animation has achieved remarkable results. However, it still faces challenges in temporally consistent video generation and fast sampling due to its iterative…

Computer Vision and Pattern Recognition · Computer Science 2025-09-22 Taekyung Ki , Dongchan Min , Gyeongsu Chae

Existing person video generation methods either lack the flexibility in controlling both the appearance and motion, or fail to preserve detailed appearance and temporal consistency. In this paper, we tackle the problem of motion transfer…

Computer Vision and Pattern Recognition · Computer Science 2019-08-13 Kun Cheng , Hao-Zhi Huang , Chun Yuan , Lingyiqing Zhou , Wei Liu

Recent studies have explored the combination of multiple LoRAs to simultaneously generate user-specified subjects and styles. However, most existing approaches fuse LoRA weights using static statistical heuristics that deviate from LoRA's…

Computer Vision and Pattern Recognition · Computer Science 2026-02-18 Qinglong Cao , Yuntian Chen , Chao Ma , Xiaokang Yang

Customized video generation aims to produce videos featuring specific subjects under flexible user-defined conditions, yet existing methods often struggle with identity consistency and limited input modalities. In this paper, we propose…

Computer Vision and Pattern Recognition · Computer Science 2025-05-09 Teng Hu , Zhentao Yu , Zhengguang Zhou , Sen Liang , Yuan Zhou , Qin Lin , Qinglin Lu

Existing mainstream video customization methods focus on generating identity-consistent videos based on given reference images and textual prompts. Benefiting from the rapid advancement of joint audio-video generation, this paper proposes a…

Sound · Computer Science 2026-05-29 Maomao Li , Zhen Li , Kaipeng Zhang , Guosheng Yin , Zhifeng Li , Dong Xu

Co-Speech Gesture Video Generation aims to generate vivid speech videos from audio-driven still images, which is challenging due to the diversity of body parts in terms of motion amplitude, audio relevance, and detailed features. Relying…

Computer Vision and Pattern Recognition · Computer Science 2026-03-16 Siyuan Wang , Jiawei Liu , Wei Wang , Yeying Jin , Jinsong Du , Zhi Han

Recent proprietary models such as Sora2 demonstrate promising progress in generating multi-shot videos conditioned on multiple reference characters. However, academic research on this problem remains limited. We study this task and identify…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Binyuan Huang , Yuning Lu , Weinan Jia , Hualiang Wang , Mu Liu , Daiqing Yang

Despite advancements in Text-to-Video (T2V) generation, producing videos with realistic motion remains challenging. Current models often yield static or minimally dynamic outputs, failing to capture complex motions described by text. This…

Computer Vision and Pattern Recognition · Computer Science 2024-11-01 Penghui Ruan , Pichao Wang , Divya Saxena , Jiannong Cao , Yuhui Shi

Subject-driven image generation plays a crucial role in applications such as virtual try-on and poster design. Existing approaches typically fine-tune pretrained generative models or apply LoRA-based adaptations for individual subjects.…

Computer Vision and Pattern Recognition · Computer Science 2025-07-03 Peng Zheng , Ye Wang , Rui Ma , Zuxuan Wu

Object-centric slot attention is a powerful framework for unsupervised learning of structured and explainable representations that can support reasoning about objects and actions, including in surgical videos. While conventional…

Image and Video Processing · Electrical Eng. & Systems 2026-03-04 Guiqiu Liao , Matjaz Jogan , Marcel Hussing , Kenta Nakahashi , Kazuhiro Yasufuku , Amin Madani , Eric Eaton , Daniel A. Hashimoto

Controllable video generation has emerged as a versatile tool for autonomous driving, enabling realistic synthesis of traffic scenarios. However, existing methods depend on control signals at inference time to guide the generative model…

Computer Vision and Pattern Recognition · Computer Science 2026-02-06 Mirlan Karimov , Teodora Spasojevic , Markus Braun , Julian Wiederer , Vasileios Belagiannis , Marc Pollefeys

We introduce an approach for augmenting text-to-video generation models with customized motions, extending their capabilities beyond the motions depicted in the original training data. By leveraging a few video samples demonstrating…

Computer Vision and Pattern Recognition · Computer Science 2024-12-11 Joanna Materzynska , Josef Sivic , Eli Shechtman , Antonio Torralba , Richard Zhang , Bryan Russell

Recent approaches in text-to-image customization have primarily focused on preserving the identity of the input subject, but often fail to control the spatial location and size of objects. We introduce GroundingBooth, which achieves…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Zhexiao Xiong , Wei Xiong , Jing Shi , He Zhang , Yizhi Song , Nathan Jacobs

Self supervised representation learning has recently attracted a lot of research interest for both the audio and visual modalities. However, most works typically focus on a particular modality or feature alone and there has been very…

Audio and Speech Processing · Electrical Eng. & Systems 2020-02-21 Abhinav Shukla , Konstantinos Vougioukas , Pingchuan Ma , Stavros Petridis , Maja Pantic

Methods for finetuning generative models for concept-driven personalization generally achieve strong results for subject-driven or style-driven generation. Recently, low-rank adaptations (LoRA) have been proposed as a parameter-efficient…

Computer Vision and Pattern Recognition · Computer Science 2026-02-05 Viraj Shah , Nataniel Ruiz , Forrester Cole , Erika Lu , Svetlana Lazebnik , Yuanzhen Li , Varun Jampani

Significant progress has been made in audio-driven human animation, while most existing methods focus mainly on facial movements, limiting their ability to create full-body animations with natural synchronization and fluidity. They also…

Computer Vision and Pattern Recognition · Computer Science 2025-06-24 Qijun Gan , Ruizi Yang , Jianke Zhu , Shaofei Xue , Steven Hoi

Referring Multi-Object Tracking (RMOT) extends conventional multi-object tracking (MOT) by introducing natural language references for multi-modal fusion tracking. RMOT benchmarks only describe the object's appearance, relative positions,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Weiyi Lv , Ning Zhang , Hanyang Sun , Haoran Jiang , Kai Zhao , Jing Xiao , Dan Zeng
‹ Prev 1 3 4 5 6 7 10 Next ›