English
Related papers

Related papers: Audio Matters Too! Enhancing Markerless Motion Cap…

200 papers

Recent deep learning approaches have achieved impressive performance on visual sound separation tasks. However, these approaches are mostly built on appearance and optical flow like motion feature representations, which exhibit limited…

Computer Vision and Pattern Recognition · Computer Science 2020-04-21 Chuang Gan , Deng Huang , Hang Zhao , Joshua B. Tenenbaum , Antonio Torralba

Pose Estimation techniques rely on visual cues available through observations represented in the form of pixels. But the performance is bounded by the frame rate of the video and struggles from motion blur, occlusions, and temporal…

Computer Vision and Pattern Recognition · Computer Science 2022-03-25 Snehesh Shrestha , Cornelia Fermüller , Tianyu Huang , Pyone Thant Win , Adam Zukerman , Chethan M. Parameshwara , Yiannis Aloimonos

The multimodal nature of music performance has driven increasing interest in data beyond the audio domain within the music information retrieval (MIR) community. This paper introduces PianoVAM, a comprehensive piano performance dataset that…

Sound · Computer Science 2025-09-11 Yonghyun Kim , Junhyung Park , Joonhyung Bae , Kirak Kim , Taegyun Kwon , Alexander Lerch , Juhan Nam

Piano performance is a multimodal activity that intrinsically combines physical actions with the acoustic rendition. Despite growing research interest in analyzing the multimodal nature of piano performance, the laborious process of…

Sound · Computer Science 2025-09-19 Junhyung Park , Yonghyun Kim , Joonhyung Bae , Kirak Kim , Taegyun Kwon , Alexander Lerch , Juhan Nam

Marker-based motion capture (MoCap) systems have long been the gold standard for accurate 4D human modeling, yet their reliance on specialized hardware and markers limits scalability and real-world deployment. Advancing reliable markerless…

Computer Vision and Pattern Recognition · Computer Science 2026-04-15 Yeeun Park , Miqdad Naduthodi , Suryansh Kumar

We present a method to combine markerless motion capture and dense pose feature estimation into a single framework. We demonstrate that dense pose information can help for multiview/single-view motion capture, and multiview motion capture…

Computer Vision and Pattern Recognition · Computer Science 2018-12-12 Xiu Li , Yebin Liu , Hanbyul Joo , Qionghai Dai , Yaser Sheikh

Piano playing requires agile, precise, and coordinated hand control that stretches the limits of dexterity. Hand motion models with the sophistication to accurately recreate piano playing have a wide range of applications in character…

Graphics · Computer Science 2024-10-10 Ruocheng Wang , Pei Xu , Haochen Shi , Elizabeth Schumann , C. Karen Liu

Generating expressive conducting gestures from music is a challenging cross-modal motion synthesis problem: the output must follow long-range musical structure, preserve beat-level synchronization, and remain plausible as a fine-grained 3D…

Sound · Computer Science 2026-05-05 Ke Qiu , Yawen Qin , Tianzhi Jia , Xiaole Yang , Kaimin Wang , Kaixing Yang

While direction of arrival (DOA) of sound events is generally estimated from multichannel audio data recorded in a microphone array, sound events usually derive from visually perceptible source objects, e.g., sounds of footsteps come from…

Prior approaches to lead instrument detection primarily analyze mixture audio, limited to coarse classifications and lacking generalization ability. This paper presents a novel approach to lead instrument detection in multitrack music audio…

Sound · Computer Science 2025-03-06 Longshen Ou , Yu Takahashi , Ye Wang

Markerless motion capture is an active research in 3D virtualization. In proposed work we presented a system for markerless motion capture for 3D human character animation, paper presents a survey on motion and skeleton tracking techniques…

Graphics · Computer Science 2014-02-12 Ashish Shingade , Archana Ghotkar

Articulated hand pose tracking is an under-explored problem that carries the potential for use in an extensive number of applications, especially in the medical domain. With a robust and accurate tracking system on surgical videos, the…

Computer Vision and Pattern Recognition · Computer Science 2025-02-10 Nathan Louis , Luowei Zhou , Steven J. Yule , Roger D. Dias , Milisa Manojlovich , Francis D. Pagani , Donald S. Likosky , Jason J. Corso

The goal of score following is to track a musical performance, usually in the form of audio, in a corresponding score representation. Established methods mainly rely on computer-readable scores in the form of MIDI or MusicXML and achieve…

Machine Learning · Computer Science 2019-10-17 Florian Henkel , Rainer Kelz , Gerhard Widmer

We introduce Multimodal DuetDance (MDD), a diverse multimodal benchmark dataset designed for text-controlled and music-conditioned 3D duet dance motion generation. Our dataset comprises 620 minutes of high-quality motion capture data…

Graphics · Computer Science 2025-08-26 Prerit Gupta , Jason Alexander Fotso-Puepi , Zhengyuan Li , Jay Mehta , Aniket Bera

Recent advances in interactive technologies have highlighted the prominence of audio signals for semantic encoding. This paper explores a new task, where audio signals are used as conditioning inputs to generate motions that align with the…

Sound · Computer Science 2025-05-30 Zi-An Wang , Shihao Zou , Shiyao Yu , Mingyuan Zhang , Chao Dong

Conditional motion generation has been extensively studied in computer vision, yet two critical challenges remain. First, while masked autoregressive methods have recently outperformed diffusion-based approaches, existing masking models…

Computer Vision and Pattern Recognition · Computer Science 2025-03-13 Zeyu Zhang , Yiran Wang , Wei Mao , Danning Li , Rui Zhao , Biao Wu , Zirui Song , Bohan Zhuang , Ian Reid , Richard Hartley

Recognizing speaking in humans is a central task towards understanding social interactions. Ideally, speaking would be detected from individual voice recordings, as done previously for meeting scenarios. However, individual voice recordings…

Computer Vision and Pattern Recognition · Computer Science 2024-03-05 Jose Vargas Quiros , Chirag Raman , Stephanie Tan , Ekin Gedik , Laura Cabrera-Quiros , Hayley Hung

We construct the first markerless deformable interaction dataset recording interactive motions of the hands and deformable objects, called HMDO (Hand Manipulation with Deformable Objects). With our built multi-view capture system, it…

Computer Vision and Pattern Recognition · Computer Science 2023-01-19 Wei Xie , Zhipeng Yu , Zimeng Zhao , Binghui Zuo , Yangang Wang

Onset detection is the process of identifying the start points of musical note events within an audio recording. While the detection of percussive onsets is often considered a solved problem, soft onsets-as found in string instrument…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-17 Maciej Tomczak , Min Susan Li , Adrian Bradbury , Mark Elliott , Ryan Stables , Maria Witek , Tom Goodman , Diar Abdlkarim , Massimiliano Di Luca , Alan Wing , Jason Hockman

Motion-centric video editing remains difficult for large generative video models, which often respond well to appearance changes but struggle to produce specific, localized actions or state transitions in an existing clip. We introduce…

‹ Prev 1 2 3 10 Next ›