English
Related papers

Related papers: CustomDancer: Customized Dance Recommendation by T…

200 papers

Large-scale noisy web image-text datasets have been proven to be efficient for learning robust vision-language models. However, when transferring them to the task of video retrieval, models still need to be fine-tuned on hand-curated paired…

Computer Vision and Pattern Recognition · Computer Science 2023-09-19 Nina Shvetsova , Anna Kukleva , Bernt Schiele , Hilde Kuehne

Fashion stylists have historically bridged the gap between consumers' desires and perfect outfits, which involve intricate combinations of colors, patterns, and materials. Although recent advancements in fashion recommendation systems have…

Computer Vision and Pattern Recognition · Computer Science 2023-11-07 Junkyu Jang , Eugene Hwang , Sung-Hyuk Park

We present CycleDance, a dance style transfer system to transform an existing motion clip in one dance style to a motion clip in another dance style while attempting to preserve motion context of the dance. Our method extends an existing…

Machine Learning · Computer Science 2023-04-04 Wenjie Yin , Hang Yin , Kim Baraka , Danica Kragic , Mårten Björkman

This paper introduces HarmonySet, a comprehensive dataset designed to advance video-music understanding. HarmonySet consists of 48,328 diverse video-music pairs, annotated with detailed information on rhythmic synchronization, emotional…

Computer Vision and Pattern Recognition · Computer Science 2025-03-05 Zitang Zhou , Ke Mei , Yu Lu , Tianyi Wang , Fengyun Rao

Social navigation and pedestrian behavior research has shifted towards machine learning-based methods and converged on the topic of modeling inter-pedestrian interactions and pedestrian-robot interactions. For this, large-scale datasets…

Computer Vision and Pattern Recognition · Computer Science 2024-10-28 Allan Wang , Daisuke Sato , Yasser Corzo , Sonya Simkin , Abhijat Biswas , Aaron Steinfeld

Dance improvisation is an active research topic in the arts. Motion analysis of improvised dance can be challenging due to its unique dynamics. Data-driven dance motion analysis, including recognition and generation, is often limited to…

Computer Vision and Pattern Recognition · Computer Science 2023-10-11 Jia Fu , Jiarui Tan , Wenjie Yin , Sepideh Pashami , Mårten Björkman

Motion synthesis for diverse object categories holds great potential for 3D content creation but remains underexplored due to two key challenges: (1) the lack of comprehensive motion datasets that include a wide range of high-quality…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Wonkwang Lee , Jongwon Jeong , Taehong Moon , Hyeon-Jong Kim , Jaehyeon Kim , Gunhee Kim , Byeong-Uk Lee

Group dance generation from music requires synchronizing multiple dancers while maintaining spatial coordination, making it highly relevant to applications such as film production, gaming, and animation. Recent group dance generation models…

Machine Learning · Computer Science 2026-03-25 Jing Xu , Weiqiang Wang , Cunjian Chen , Jun Liu , Qiuhong Ke

Linking human motion and natural language is of great interest for the generation of semantic representations of human activities as well as for the generation of robot activities based on natural language input. However, while there have…

Robotics · Computer Science 2018-08-10 Matthias Plappert , Christian Mandery , Tamim Asfour

Text-driven human motion synthesis has showcased its potential for revolutionizing motion design in the movie and game industry. Existing methods often rely on 3D motion capture data, which requires special setups, resulting in high costs…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Ruoxi Guo , Huaijin Pi , Zehong Shen , Qing Shuai , Zechen Hu , Zhumei Wang , Yajiao Dong , Ruizhen Hu , Taku Komura , Sida Peng , Xiaowei Zhou

In this paper, we introduce RoleMotion, a large-scale human motion dataset that encompasses a wealth of role-playing and functional motion data tailored to fit various specific scenes. Existing text datasets are mainly constructed…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Junran Peng , Yiheng Huang , Silei Shen , Zeji Wei , Jingwei Yang , Baojie Wang , Yonghao He , Chuanchen Luo , Man Zhang , Xucheng Yin , Wei Sui

Online video web content is richly multimodal: a single video blends vision, speech, ambient audio, and on-screen text. Retrieval systems typically treat these modalities as independent retrieval sources, which can lead to noisy and subpar…

Computer Vision and Pattern Recognition · Computer Science 2025-06-09 David Wan , Han Wang , Elias Stengel-Eskin , Jaemin Cho , Mohit Bansal

In recent years, large language models (LLMs) have made rapid progress in information retrieval, yet existing research has mainly focused on text or static multimodal settings. Open-domain video shot retrieval, which involves richer…

The task of retrieving video content relevant to natural language queries plays a critical role in effectively handling internet-scale datasets. Most of the existing methods for this caption-to-video retrieval problem do not fully exploit…

Computer Vision and Pattern Recognition · Computer Science 2020-07-22 Valentin Gabeur , Chen Sun , Karteek Alahari , Cordelia Schmid

Studies of object detection and localization, particularly pedestrian detection have received considerable attention in recent times due to its several prospective applications such as surveillance, driving assistance, autonomous cars, etc.…

Computer Vision and Pattern Recognition · Computer Science 2019-12-24 Sudip Das , Partha Sarathi Mukherjee , Ujjwal Bhattacharya

Well-coordinated, music-aligned holistic dance enhances emotional expressiveness and audience engagement. However, generating such dances remains challenging due to the scarcity of holistic 3D dance datasets, the difficulty of achieving…

Multimedia · Computer Science 2025-07-30 Xiaojie Li , Ronghui Li , Shukai Fang , Shuzhao Xie , Xiaoyang Guo , Jiaqing Zhou , Junkun Peng , Zhi Wang

We introduce TV show Retrieval (TVR), a new multimodal retrieval dataset. TVR requires systems to understand both videos and their associated subtitle (dialogue) texts, making it more realistic. The dataset contains 109K queries collected…

Computer Vision and Pattern Recognition · Computer Science 2020-08-19 Jie Lei , Licheng Yu , Tamara L. Berg , Mohit Bansal

Text-video retrieval is a critical multi-modal task to find the most relevant video for a text query. Although pretrained models like CLIP have demonstrated impressive potential in this area, the rising cost of fully finetuning these models…

Computer Vision and Pattern Recognition · Computer Science 2024-01-22 Xiangpeng Yang , Linchao Zhu , Xiaohan Wang , Yi Yang

What we appreciate in dance is the ability of people to sponta- neously improvise new movements and choreographies, sur- rendering to the music rhythm, being inspired by the cur- rent perceptions and sensations and by previous experiences,…

Artificial Intelligence · Computer Science 2017-08-02 Agnese Augello , Emanuele Cipolla , Ignazio Infantino , Adriano Manfre , Giovanni Pilato , Filippo Vella

Cross-modal retrieval between videos and texts has gained increasing research interest due to the rapid emergence of videos on the web. Generally, a video contains rich instance and event information and the query text only describes a part…

Computer Vision and Pattern Recognition · Computer Science 2022-09-28 Chengzhi Lin , Ancong Wu , Junwei Liang , Jun Zhang , Wenhang Ge , Wei-Shi Zheng , Chunhua Shen
‹ Prev 1 3 4 5 6 7 10 Next ›