English
Related papers

Related papers: Audio Matters Too! Enhancing Markerless Motion Cap…

200 papers

Despite the central role that melody plays in music perception, it remains an open challenge in music information retrieval to reliably detect the notes of the melody present in an arbitrary music recording. A key challenge in melody…

Sound · Computer Science 2022-12-06 Chris Donahue , John Thickstun , Percy Liang

Recent focus in video captioning has been on designing architectures that can consume both video and text modalities, and using large-scale video datasets with text transcripts for pre-training, such as HowTo100M. Though these approaches…

Computer Vision and Pattern Recognition · Computer Science 2023-06-23 Yuhan Shen , Linjie Yang , Longyin Wen , Haichao Yu , Ehsan Elhamifar , Heng Wang

Recently, an event-based end-to-end model (SEDT) has been proposed for sound event detection (SED) and achieves competitive performance. However, compared with the frame-based model, it requires more training data with temporal annotations…

Sound · Computer Science 2022-04-07 Zhirong Ye , Xiangdong Wang , Hong Liu , Yueliang Qian , Rui Tao , Long Yan , Kazushige Ouchi

Human motion generation aims to generate natural human pose sequences and shows immense potential for real-world applications. Substantial progress has been made recently in motion data collection technologies and generation methods, laying…

Computer Vision and Pattern Recognition · Computer Science 2023-11-16 Wentao Zhu , Xiaoxuan Ma , Dongwoo Ro , Hai Ci , Jinlu Zhang , Jiaxin Shi , Feng Gao , Qi Tian , Yizhou Wang

In the age of music streaming platforms, the task of automatically tagging music audio has garnered significant attention, driving researchers to devise methods aimed at enhancing performance metrics on standard datasets. Most recent…

Sound · Computer Science 2024-02-26 Vassilis Lyberatos , Spyridon Kantarelis , Edmund Dervakos , Giorgos Stamou

This paper presents a new continuous interaction strategy with visual feedback of hand pose and mid-air gesture recognition and control for a smart music speaker, which utilizes only 2 video frames to recognize gestures. Frame-based hand…

Human-Computer Interaction · Computer Science 2023-03-01 Songpei Xu , Chaitanya Kaul , Xuri Ge , Roderick Murray-Smith

For many music analysis problems, we need to know the presence of instruments for each time frame in a multi-instrument musical piece. However, such a frame-level instrument recognition task remains difficult, mainly due to the lack of…

Sound · Computer Science 2019-02-19 Yun-Ning Hung , Yi-An Chen , Yi-Hsuan Yang

We propose a new task named Audio-driven Per-formance Video Generation (APVG), which aims to synthesizethe video of a person playing a certain instrument guided bya given music audio clip. It is a challenging task to gener-ate the…

Computer Vision and Pattern Recognition · Computer Science 2020-11-06 Hao Zhu , Yi Li , Feixia Zhu , Aihua Zheng , Ran He

In this work, we present a marker-based multi-view spine tracking method that is specifically adjusted to the requirements for movements in sports. A maximal focus is on the accurate detection of markers and fast usage of the system. For…

Computer Vision and Pattern Recognition · Computer Science 2023-06-06 Hendrik Hachmann , Bodo Rosenhahn

We study the merit of transfer learning for two sound recognition problems, i.e., audio tagging and sound event detection. Employing feature fusion, we adapt a baseline system utilizing only spectral acoustic inputs to also make use of…

Audio and Speech Processing · Electrical Eng. & Systems 2022-09-27 Wim Boes , Hugo Van hamme

We tackle the problem of highly-accurate, holistic performance capture for the face, body and hands simultaneously. Motion-capture technologies used in film and game production typically focus only on face, body or hand capture…

Human movement analysis is a key area of research in robotics, biomechanics, and data science. It encompasses tracking, posture estimation, and movement synthesis. While numerous methodologies have evolved over time, a systematic and…

Robotics · Computer Science 2023-05-11 Brenda Elizabeth Olivas-Padilla , Alina Glushkova , Sotiris Manitsaris

We present a simple lightweight markerless facial performance capture framework using just a monocular video input that combines Active Appearance Models for feature tracking and prior constraints on 3D shapes into an integrated objective…

Computer Vision and Pattern Recognition · Computer Science 2019-01-17 Shridhar Ravikumar

Tracking 3D human motion in real-time is crucial for numerous applications across many fields. Traditional approaches involve attaching artificial fiducial objects or sensors to the body, limiting their usability and comfort-of-use and…

Computer Vision and Pattern Recognition · Computer Science 2023-04-03 Luca Fortini , Mattia Leonori , Juan M. Gandarias , Elena de Momi , Arash Ajoudani

Hand gesture recognition attracts great attention for interaction since it is intuitive and natural to perform. In this paper, we explore a novel method for interaction by using bone-conducted sound generated by finger movements while…

Human-Computer Interaction · Computer Science 2021-12-14 Bing Zhou , Matias Aiskovich , Sinem Guven

Markerless motion capture enables the tracking of human motion without requiring physical markers or suits, offering increased flexibility and reduced costs compared to traditional systems. However, these advantages often come at the…

Computer Vision and Pattern Recognition · Computer Science 2025-06-24 David Tolpin , Sefy Kagarlitsky

This paper presents a neural network model to generate virtual violinist's 3-D skeleton movements from music audio. Improved from the conventional recurrent neural network models for generating 2-D skeleton data in previous works, the…

Multimedia · Computer Science 2020-09-18 Hsuan-Kai Kao , Li Su

The objective of this paper is to perform audio-visual sound source separation, i.e.~to separate component audios from a mixture based on the videos of sound sources. Moreover, we aim to pinpoint the source location in the input video…

Computer Vision and Pattern Recognition · Computer Science 2021-04-20 Lingyu Zhu , Esa Rahtu

Generating 3D dances from music is an emerged research task that benefits a lot of applications in vision and graphics. Previous works treat this task as sequence generation, however, it is challenging to render a music-aligned long-term…

Artificial Intelligence · Computer Science 2023-07-28 Buyu Li , Yongchi Zhao , Zhelun Shi , Lu Sheng

Human motion video generation has garnered significant research interest due to its broad applications, enabling innovations such as photorealistic singing heads or dynamic avatars that seamlessly dance to music. However, existing surveys…

Computer Vision and Pattern Recognition · Computer Science 2025-09-05 Haiwei Xue , Xiangyang Luo , Zhanghao Hu , Xin Zhang , Xunzhi Xiang , Yuqin Dai , Jianzhuang Liu , Zhensong Zhang , Minglei Li , Jian Yang , Fei Ma , Zhiyong Wu , Changpeng Yang , Zonghong Dai , Fei Richard Yu