English
Related papers

Related papers: Co-Designing Multimodal Systems for Accessible Asy…

200 papers

Several studies report that people diagnosed with autism spectrum disorder (ASD) have low employment rates and major difficulties in maintaining occupation (Hendricks, 2010). Adopting smart and interactive technology, promoted by Industry…

Human-Computer Interaction · Computer Science 2023-04-28 Carla Dei , Matteo Meregalli Falerni , Matteo Lavit Nicora , Mattia Chiappini , Matteo Malosio , Fabio Alexander Storm

Instructors play a pivotal role in integrating AI into education, yet their adoption of AI-powered tools remains inconsistent. Despite this, limited research explores how to design AI tools that support broader instructor adoption. This…

Human-Computer Interaction · Computer Science 2025-03-10 Si Chen , Reid Metoyer , Khiem Le , Adam Acunin , Izzy Molnar , Alex Ambrose , James Lang , Nitesh Chawla , Ronald Metoyer

Unified audio-visual generation is rapidly gaining industrial and creative relevance, enabling applications in virtual production and interactive media. However, when moving from general audio-video synthesis to music-dance co-generation,…

As robots enter the messy human world so the vital matter of safety takes on a fresh complexion with physical contact becoming inevitable and even desirable. We report on an artistic-exploration of how dancers, working as part of a…

We present a learning-based approach with pose perceptual loss for automatic music video generation. Our method can produce a realistic dance video that conforms to the beats and rhymes of almost any given music. To achieve this, we firstly…

Computer Vision and Pattern Recognition · Computer Science 2019-12-16 Xuanchi Ren , Haoran Li , Zijian Huang , Qifeng Chen

Guiding users through complex procedural plans is an inherently multimodal task in which having visually illustrated plan steps is crucial to deliver an effective plan guidance. However, existing works on plan-following language models…

Computer Vision and Pattern Recognition · Computer Science 2024-10-22 Diogo Glória-Silva , David Semedo , João Magalhães

We address the problem of text-guided video temporal grounding, which aims to identify the time interval of a certain event based on a natural language description. Different from most existing methods that only consider RGB images as…

Computer Vision and Pattern Recognition · Computer Science 2021-11-01 Yi-Wen Chen , Yi-Hsuan Tsai , Ming-Hsuan Yang

Multimodal machine learning is a vibrant multi-disciplinary research field that aims to design computer agents with intelligent capabilities such as understanding, reasoning, and learning through integrating multiple communicative…

Machine Learning · Computer Science 2023-02-21 Paul Pu Liang , Amir Zadeh , Louis-Philippe Morency

Dialog systems need to understand dynamic visual scenes in order to have conversations with users about the objects and events around them. Scene-aware dialog systems for real-world applications could be developed by integrating…

The Audio Description (AD) task aims to generate descriptions of visual elements for visually impaired individuals to help them access long-form video content, like movies. With video feature, text, character bank and context information as…

Computer Vision and Pattern Recognition · Computer Science 2025-04-16 Hanlin Wang , Zhan Tong , Kecheng Zheng , Yujun Shen , Limin Wang

Recent advances in multimodal LLMs, have led to several video-text models being proposed for critical video-related tasks. However, most of the previous works support visual input only, essentially muting the audio signal in the video. Few…

Computer Vision and Pattern Recognition · Computer Science 2024-07-23 Shivprasad Sagare , Hemachandran S , Kinshuk Sarabhai , Prashant Ullegaddi , Rajeshkumar SA

Active learning methods, like uncertainty sampling, combined with probabilistic prediction techniques have achieved success in various problems like image classification and text classification. For more complex multivariate prediction…

Machine Learning · Computer Science 2020-03-24 Sima Behpour

We describe a novel metric-based learning approach that introduces a multimodal framework and uses deep audio and geophone encoders in siamese configuration to design an adaptable and lightweight supervised model. This framework eliminates…

Sound · Computer Science 2021-11-16 Muhammad Shakeel , Katsutoshi Itoyama , Kenji Nishida , Kazuhiro Nakadai

Tutorial videos are a valuable resource for people looking to learn new tasks. People often learn these skills by viewing multiple tutorial videos to get an overall understanding of a task by looking at different approaches to achieve the…

Human-Computer Interaction · Computer Science 2025-03-28 Saelyne Yang , Anh Truong , Juho Kim , Dingzeyu Li

Comparing a user video to a reference how-to video is a key requirement for AR/VR technology delivering personalized assistance tailored to the user's progress. However, current approaches for language-based assistance can only answer…

Computer Vision and Pattern Recognition · Computer Science 2024-07-01 Tushar Nagarajan , Lorenzo Torresani

Providing timely, targeted, and multimodal feedback helps students quickly correct errors, build deep understanding and stay motivated, yet making it at scale remains a challenge. This study introduces a real-time AI-facilitated multimodal…

Human-Computer Interaction · Computer Science 2026-05-14 Chloe Qianhui Zhao , Jie Cao , Jionghao Lin , Kenneth R. Koedinger

Robotic guidance systems have shown promise in supporting blind and visually impaired (BVI) individuals with wayfinding and obstacle avoidance. However, most existing systems assume a clear path and do not support a critical aspect of…

Robotics · Computer Science 2026-03-17 Shaojun Cai , Nuwan Janaka , Ashwin Ram , Janidu Shehan , Yingjia Wan , Kotaro Hara , David Hsu

Generating coherent and diverse human dances from music signals has gained tremendous progress in animating virtual avatars. While existing methods support direct dance synthesis, they fail to recognize that enabling users to edit dance…

Graphics · Computer Science 2025-08-26 Hengyuan Zhang , Zhe Li , Xingqun Qi , Mengze Li , Muyi Sun , Man Zhang , Sirui Han

Visually impaired people encounter many challenges in their everyday life, especially when it comes to navigating and representing space. The issue of shopping is addressed mostly on the level of navigation and product detection, but…

Human-Computer Interaction · Computer Science 2020-10-12 Lancelot Dupont , Christophe Jouffrais , Simon T. Perrault

In this article, we present a Web-based System called M2LADS, which supports the integration and visualization of multimodal data recorded in learning sessions in a MOOC in the form of Web-based Dashboards. Based on the edBB platform, the…

Human-Computer Interaction · Computer Science 2023-05-23 Álvaro Becerra , Roberto Daza , Ruth Cobos , Aythami Morales , Mutlu Cukurova , Julian Fierrez
‹ Prev 1 8 9 10 Next ›