English
Related papers

Related papers: Co-Designing Multimodal Systems for Accessible Asy…

200 papers

Generating dance from music is crucial for advancing automated choreography. Current methods typically produce skeleton keypoint sequences instead of dance videos and lack the capability to make specific individuals dance, which reduces…

Computer Vision and Pattern Recognition · Computer Science 2024-12-02 Xuanchen Wang , Heng Wang , Dongnan Liu , Weidong Cai

We propose a novel task for generating 3D dance movements that simultaneously incorporate both text and music modalities. Unlike existing works that generate dance movements using a single modality such as music, our goal is to produce…

Computer Vision and Pattern Recognition · Computer Science 2023-10-03 Kehong Gong , Dongze Lian , Heng Chang , Chuan Guo , Zihang Jiang , Xinxin Zuo , Michael Bi Mi , Xinchao Wang

In the context of Audio Visual Question Answering (AVQA) tasks, the audio visual modalities could be learnt on three levels: 1) Spatial, 2) Temporal, and 3) Semantic. Existing AVQA methods suffer from two major shortcomings; the…

Computer Vision and Pattern Recognition · Computer Science 2023-10-31 Asmar Nadeem , Adrian Hilton , Robert Dawes , Graham Thomas , Armin Mustafa

To generate dance that temporally and aesthetically matches the music is a challenging problem, as the following factors need to be considered. First, the aesthetic styles and messages conveyed by the motion and music should be consistent.…

Multimedia · Computer Science 2022-07-18 Ho Yin Au , Jie Chen , Junkun Jiang , Yike Guo

Existing methods for human motion control in video generation typically rely on either 2D poses or explicit 3D parametric models (e.g., SMPL) as control signals. However, 2D poses rigidly bind motion to the driving viewpoint, precluding…

Computer Vision and Pattern Recognition · Computer Science 2026-02-17 Zhixue Fang , Xu He , Songlin Tang , Haoxian Zhang , Qingfeng Li , Xiaoqiang Liu , Pengfei Wan , Kun Gai

Augmented reality (AR) offers promising opportunities to support movement-based activities, such as personal training or physical therapy, with real-time, spatially-situated visual cues. While many approaches leverage AR to guide motion,…

Human-Computer Interaction · Computer Science 2025-10-02 Jade Kandel , Sriya Kasumarthi , Spiros Tsalikis , Chelsea Duppen , Daniel Szafir , Michael Lewek , Henry Fuchs , Danielle Szafir

The task of learning the piano has been a centuries-old challenge for novices, experts and technologists. Several innovations have been introduced to support proper posture, movement, and motivation, while sight-reading and improvisation…

Human-Computer Interaction · Computer Science 2022-11-14 Jordan Aiko Deja

Human intelligence is multimodal; we integrate visual, linguistic, and acoustic signals to maintain a holistic worldview. Most current pretraining methods, however, are limited to one or two modalities. We present i-Code, a self-supervised…

Despite impressive high-level video comprehension, multimodal language models struggle with spatial reasoning across time and space. While current spatial training approaches rely on real-world video data, obtaining diverse footage with…

Computer Vision and Pattern Recognition · Computer Science 2025-11-17 Ellis Brown , Arijit Ray , Ranjay Krishna , Ross Girshick , Rob Fergus , Saining Xie

Artificial intelligence (AI) and large language models (LLMs) are reshaping education, with virtual avatars emerging as digital teachers capable of enhancing engagement, sustaining attention, and addressing instructor shortages. Aligned…

Human-Computer Interaction · Computer Science 2026-01-27 Xiaokang Lei , Ching Christie Pang , Yuyang Jiang , Xin Tong , Pan Hui

Approximately 200 million individuals around the world suffer from varying degrees of visual impairment, making it crucial to leverage AI technology to offer walking assistance for these people. With the recent progress of vision-language…

Computer Vision and Pattern Recognition · Computer Science 2025-03-05 Zhiqiang Yuan , Ting Zhang , Ying Deng , Jiapei Zhang , Yeshuang Zhu , Zexi Jia , Jie Zhou , Jinchao Zhang

With the rapid development of computer technology, computer music has begun to appear in the laboratory. Many potential utility of computer music is gradually increasing. The purpose of this paper is attempted to analyze the possibility of…

Multimedia · Computer Science 2010-05-24 Gilbert Phuah Leong Siang , Nor Azman Ismail , Pang Yee Yong

Building 3-D models is challenging for blind and low-vision (BLV) users due to the inherent complexity of 3-D models and the lack of support for non-visual interaction in existing tools. To address this issue, we introduce A11yShape, a…

The motivation of this paper is to develop a smart system using multi-modal vision for next-generation mechanical assembly. It includes two phases where in the first phase human beings teach the assembly structure to a robot and in the…

Robotics · Computer Science 2016-01-27 Weiwei Wan , Feng Lu , Zepei Wu , Kensuke Harada

The ability to provide fine-grained control for generating and editing visual imagery has profound implications for computer vision and its applications. Previous works have explored extending controllability in two directions: instruction…

Computer Vision and Pattern Recognition · Computer Science 2024-10-18 Shufan Li , Harkanwar Singh , Aditya Grover

Information retrieval is an ever-evolving and crucial research domain. The substantial demand for high-quality human motion data especially in online acquirement has led to a surge in human motion research works. Prior works have mainly…

Computer Vision and Pattern Recognition · Computer Science 2024-03-04 Kangning Yin , Shihao Zou , Yuxuan Ge , Zheng Tian

Advances in Deep Learning have recently made it possible to recover full 3D meshes of human poses from individual images. However, extension of this notion to videos for recovering temporally coherent poses still remains unexplored. A major…

Computer Vision and Pattern Recognition · Computer Science 2019-07-02 Jian Liu , Naveed Akhtar , Ajmal Mian

Authors make their videos visually accessible by adding audio descriptions (AD), and auditorily accessible by adding closed captions (CC). However, creating AD and CC is challenging and tedious, especially for non-professional describers…

Human-Computer Interaction · Computer Science 2025-02-19 Xingyu "Bruce" Liu , Ruolin Wang , Dingzeyu Li , Xiang 'Anthony' Chen , Amy Pavel

We consider a sequence of related multivariate time series learning tasks, such as predicting failures for different instances of a machine from time series of multi-sensor data, or activity recognition tasks over different individuals from…

Machine Learning · Computer Science 2022-03-15 Vibhor Gupta , Jyoti Narwariya , Pankaj Malhotra , Lovekesh Vig , Gautam Shroff

Effective bipedal locomotion in dynamic environments, such as cluttered indoor spaces or uneven terrain, requires agile and adaptive movement in all directions. This necessitates omnidirectional terrain sensing and a controller capable of…

Robotics · Computer Science 2026-03-18 Mohitvishnu S. Gadde , Pranay Dugar , Ashish Malik , Alan Fern