English
Related papers

Related papers: MG-Former: A Transformer-Based Framework for Music…

200 papers

Existing music-driven dance generation approaches have achieved strong realism and effective audio-motion alignment. However, they generally lack semantic controllability, making it difficult to guide specific movements through natural…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Xinran Liu , Diptesh Kanojia , Wenwu Wang , Zhenhua Feng

3D hand pose is an underexplored modality for action recognition. Poses are compact yet informative and can greatly benefit applications with limited compute budgets. However, poses alone offer an incomplete understanding of actions, as…

Computer Vision and Pattern Recognition · Computer Science 2024-08-15 Md Salman Shamil , Dibyadip Chatterjee , Fadime Sener , Shugao Ma , Angela Yao

Aligning multiple modalities in a latent space, such as images and texts, has shown to produce powerful semantic visual representations, fueling tasks like image captioning, text-to-image generation, or image grounding. In the context of…

Computer Vision and Pattern Recognition · Computer Science 2024-09-11 Ginger Delmas , Philippe Weinzaepfel , Francesc Moreno-Noguer , Grégory Rogez

For human-like agents, including virtual avatars and social robots, making proper gestures while speaking is crucial in human--agent interaction. Co-speech gestures enhance interaction experiences and make the agents look alive. However, it…

Graphics · Computer Science 2020-09-07 Youngwoo Yoon , Bok Cha , Joo-Haeng Lee , Minsu Jang , Jaeyeon Lee , Jaehong Kim , Geehyuk Lee

Recent advancements in large language models (LLMs) have significantly propelled the development of large multi-modal models (LMMs), highlighting the potential for general and intelligent assistants. However, most LMMs model visual and…

Computation and Language · Computer Science 2025-03-20 Rui Yang , Lin Song , Yicheng Xiao , Runhui Huang , Yixiao Ge , Ying Shan , Hengshuang Zhao

Existing keyframe-based motion synthesis mainly focuses on the generation of cyclic actions or short-term motion, such as walking, running, and transitions between close postures. However, these methods will significantly degrade the…

Computer Vision and Pattern Recognition · Computer Science 2022-07-14 Junjun Pan , Siyuan Wang , Junxuan Bai , Ju Dai

Advancements in language foundation models have primarily fueled the recent surge in artificial intelligence. In contrast, generative learning of non-textual modalities, especially videos, significantly trails behind language modeling. This…

Computer Vision and Pattern Recognition · Computer Science 2024-05-28 Lijun Yu

Pedestrian trajectory prediction, vital for selfdriving cars and socially-aware robots, is complicated due to intricate interactions between pedestrians, their environment, and other Vulnerable Road Users. This paper presents GSGFormer, an…

Computer Vision and Pattern Recognition · Computer Science 2023-12-08 Zhongchang Luo , Marion Robin , Pavan Vasishta

This paper presents an innovative application of Transformer-XL for long sequence tasks in robotic learning from demonstrations (LfD). The proposed framework effectively integrates multi-modal sensor inputs, including RGB-D images, LiDAR,…

Robotics · Computer Science 2025-12-16 Gao Tianci

Recent advances in diffusion models have demonstrated exceptional capabilities in image and video generation, further improving the effectiveness of 4D synthesis. Existing 4D generation methods can generate high-quality 4D objects or scenes…

Computer Vision and Pattern Recognition · Computer Science 2024-10-10 Bohan Zeng , Ling Yang , Siyu Li , Jiaming Liu , Zixiang Zhang , Juanxi Tian , Kaixin Zhu , Yongzhen Guo , Fu-Yun Wang , Minkai Xu , Stefano Ermon , Wentao Zhang

Data-driven and controllable human motion synthesis and prediction are active research areas with various applications in interactive media and social robotics. Challenges remain in these fields for generating diverse motions given past…

Computer Vision and Pattern Recognition · Computer Science 2023-04-11 Wenjie Yin , Ruibo Tu , Hang Yin , Danica Kragic , Hedvig Kjellström , Mårten Björkman

Lack of audio-video synchronization is a common problem during television broadcasts and video conferencing, leading to an unsatisfactory viewing experience. A widely accepted paradigm is to create an error detection mechanism that…

Computer Vision and Pattern Recognition · Computer Science 2023-03-22 Akash Gupta , Rohun Tripathi , Wondong Jang

Audio-driven human gesture synthesis is a crucial task with broad applications in virtual avatars, human-computer interaction, and creative content generation. Despite notable progress, existing methods often produce gestures that are…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Xukun Zhou , Fengxin Li , Ming Chen , Yan Zhou , Pengfei Wan , Di Zhang , Yeying Jin , Zhaoxin Fan , Hongyan Liu , Jun He

Recent techniques for text-to-4D generation synthesize dynamic 3D scenes using supervision from pre-trained text-to-video models. However, existing representations for motion, such as deformation models or time-dependent neural…

Computer Vision and Pattern Recognition · Computer Science 2024-10-16 Sherwin Bahmani , Xian Liu , Wang Yifan , Ivan Skorokhodov , Victor Rong , Ziwei Liu , Xihui Liu , Jeong Joon Park , Sergey Tulyakov , Gordon Wetzstein , Andrea Tagliasacchi , David B. Lindell

The development of multimodal models has significantly advanced multimodal sentiment analysis and emotion recognition. However, in real-world applications, the presence of various missing modality cases often leads to a degradation in the…

Computation and Language · Computer Science 2024-07-09 Zirun Guo , Tao Jin , Zhou Zhao

Driving 3D characters to dance following a piece of music is highly challenging due to the spatial constraints applied to poses by choreography norms. In addition, the generated dance sequence also needs to maintain temporal coherency with…

Sound · Computer Science 2022-03-28 Li Siyao , Weijiang Yu , Tianpei Gu , Chunze Lin , Quan Wang , Chen Qian , Chen Change Loy , Ziwei Liu

How to automatically synthesize natural-looking dance movements based on a piece of music is an incrementally popular yet challenging task. Most existing data-driven approaches require hard-to-get paired training data and fail to generate…

Computer Vision and Pattern Recognition · Computer Science 2023-03-30 Bin Feng , Tenglong Ao , Zequn Liu , Wei Ju , Libin Liu , Ming Zhang

The recent success in StyleGAN demonstrates that pre-trained StyleGAN latent space is useful for realistic video generation. However, the generated motion in the video is usually not semantically meaningful due to the difficulty of…

Computer Vision and Pattern Recognition · Computer Science 2022-10-24 Seung Hyun Lee , Gyeongrok Oh , Wonmin Byeon , Chanyoung Kim , Won Jeong Ryoo , Sang Ho Yoon , Hyunjun Cho , Jihyun Bae , Jinkyu Kim , Sangpil Kim

Music relies heavily on repetition to build structure and meaning. Self-reference occurs on multiple timescales, from motifs to phrases to reusing of entire sections of music, such as in pieces with ABA structure. The Transformer (Vaswani…

Predicting pedestrian behavior is a crucial task for intelligent driving systems. Accurate predictions require a deep understanding of various contextual elements that potentially impact the way pedestrians behave. To address this…

Computer Vision and Pattern Recognition · Computer Science 2022-10-17 Amir Rasouli , Iuliia Kotseruba