English
Related papers

Related papers: CustomDancer: Customized Dance Recommendation by T…

200 papers

Text-to-video (T2V) diffusion models have shown promising capabilities in synthesizing realistic videos from input text prompts. However, the input text description alone provides limited control over the precise objects movements and…

Computer Vision and Pattern Recognition · Computer Science 2025-02-20 Yen-Siang Wu , Chi-Pin Huang , Fu-En Yang , Yu-Chiang Frank Wang

Video retrieval using natural language queries requires learning semantically meaningful joint embeddings between the text and the audio-visual input. Often, such joint embeddings are learnt using pairwise (or triplet) contrastive loss…

Information Retrieval · Computer Science 2021-03-10 Jayaprakash A , Abhishek , Rishabh Dabral , Ganesh Ramakrishnan , Preethi Jyothi

Creating high-dynamic videos such as motion-rich actions and sophisticated visual effects poses a significant challenge in the field of artificial intelligence. Unfortunately, current state-of-the-art video generation methods, primarily…

Computer Vision and Pattern Recognition · Computer Science 2023-11-21 Yan Zeng , Guoqiang Wei , Jiani Zheng , Jiaxin Zou , Yang Wei , Yuchen Zhang , Hang Li

Synthesizing human motion with a global structure, such as a choreography, is a challenging task. Existing methods tend to concentrate on local smooth pose transitions and neglect the global context or the theme of the motion. In this work,…

Human motion recovered from monocular videos often appears overly smooth or dynamically inconsistent, even when joint positions are numerically accurate. We observe that this limitation stems from the absence of reliable high-order temporal…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Dingkun Wei , Zehong Shen , Yan Xia , Georgios Pavlakos , Yujun Shen , Xiaowei Zhou

Videos are more informative than images because they capture the dynamics of the scene. By representing motion in videos, we can capture dynamic activities. In this work, we introduce GPT-4 generated motion descriptions that capture…

Computer Vision and Pattern Recognition · Computer Science 2024-06-10 Chinmaya Devaraj , Cornelia Fermuller , Yiannis Aloimonos

Learning a good representation of text is key to many recommendation applications. Examples include news recommendation where texts to be recommended are constantly published everyday. However, most existing recommendation techniques, such…

Information Retrieval · Computer Science 2017-06-27 Ting Chen , Liangjie Hong , Yue Shi , Yizhou Sun

Advances in Deep Learning have recently made it possible to recover full 3D meshes of human poses from individual images. However, extension of this notion to videos for recovering temporally coherent poses still remains unexplored. A major…

Computer Vision and Pattern Recognition · Computer Science 2019-07-02 Jian Liu , Naveed Akhtar , Ajmal Mian

Our objective in this work is long range understanding of the narrative structure of movies. Instead of considering the entire movie, we propose to learn from the `key scenes' of the movie, providing a condensed look at the full storyline.…

Computer Vision and Pattern Recognition · Computer Science 2020-10-26 Max Bain , Arsha Nagrani , Andrew Brown , Andrew Zisserman

Recent years have witnessed an increasing amount of dialogue/conversation on the web especially on social media. That inspires the development of dialogue-based retrieval, in which retrieving videos based on dialogue is of increasing…

Information Retrieval · Computer Science 2023-03-30 Chenyang Lyu , Manh-Duy Nguyen , Van-Tu Ninh , Liting Zhou , Cathal Gurrin , Jennifer Foster

Reactive dance generation (RDG), the task of generating a dance conditioned on a lead dancer's motion, holds significant promise for enhancing human-robot interaction and immersive digital entertainment. Despite progress in duet…

Computer Vision and Pattern Recognition · Computer Science 2026-03-06 Jingzhong Lin , Xinru Li , Yuanyuan Qi , Bohao Zhang , Wenxiang Liu , Kecheng Tang , Wenxuan Huang , Xiangfeng Xu , Bangyan Li , Changbo Wang , Gaoqi He

Animating stylized characters to match a reference motion sequence is a highly demanded task in film and gaming industries. Existing methods mostly focus on rigid deformations of characters' body, neglecting local deformations on the…

Computer Vision and Pattern Recognition · Computer Science 2024-07-17 Rong Wang , Wei Mao , Changsheng Lu , Hongdong Li

As sports training becomes more data-driven, traditional dart coaching based mainly on experience and visual observation is increasingly inadequate for high-precision, goal-oriented movements. Although prior studies have highlighted the…

Machine Learning · Computer Science 2026-04-09 Zhantao Chen , Dongyi He , Jin Fang , Xi Chen , Yishuo Liu , Xiaozhen Zhong , Xuejun Hu

Audio-Text retrieval takes a natural language query to retrieve relevant audio files in a database. Conversely, Text-Audio retrieval takes an audio file as a query to retrieve relevant natural language descriptions. Most of the literature…

Audio and Speech Processing · Electrical Eng. & Systems 2022-09-29 Soham Deshmukh , Benjamin Elizalde , Huaming Wang

Group dance generation from music has broad applications in film, gaming, and animation production. However, it requires synchronizing multiple dancers while maintaining spatial coordination. As the number of dancers and sequence length…

Artificial Intelligence · Computer Science 2025-07-31 Jing Xu , Weiqiang Wang , Cunjian Chen , Jun Liu , Qiuhong Ke

In this paper, we address the problem of spatio-temporal person retrieval from multiple videos using a natural language query, in which we output a tube (i.e., a sequence of bounding boxes) which encloses the person described by the query.…

Computer Vision and Pattern Recognition · Computer Science 2017-08-24 Masataka Yamaguchi , Kuniaki Saito , Yoshitaka Ushiku , Tatsuya Harada

Text-based person search (TBPS) aims at retrieving a target person from an image gallery with a descriptive text query. Solving such a fine-grained cross-modal retrieval task is challenging, which is further hampered by the lack of…

Computer Vision and Pattern Recognition · Computer Science 2021-10-22 Xiao Han , Sen He , Li Zhang , Tao Xiang

Dancing to music is one of human's innate abilities since ancient times. In machine learning research, however, synthesizing dance movements from music is a challenging problem. Recently, researchers synthesize human motion sequences…

Computer Vision and Pattern Recognition · Computer Science 2023-09-12 Ruozi Huang , Huang Hu , Wei Wu , Kei Sawada , Mi Zhang , Daxin Jiang

Existing dominant approaches for cross-modal video-text retrieval task are to learn a joint embedding space to measure the cross-modal similarity. However, these methods rarely explore long-range dependency inside video frames or textual…

Multimedia · Computer Science 2020-04-13 Rui Zhao , Kecheng Zheng , Zheng-jun Zha

Realistic human-centric rendering plays a key role in both computer vision and computer graphics. Rapid progress has been made in the algorithm aspect over the years, yet existing human-centric rendering datasets and benchmarks are rather…