中文
相关论文

相关论文: Sanvaad: A Multimodal Accessibility Framework for …

200 篇论文

Sign language is a visual language that enhances communication between people and is frequently used as the primary form of communication by people with hearing loss. Even so, not many people with hearing loss use sign language, and they…

计算机视觉与模式识别 · 计算机科学 2023-05-25 Rupesh Kumar , Ayush Sinha , Ashutosh Bajpai , S. K Singh

We present, to our knowledge, the first sign language-driven Vision-Language-Action (VLA) framework for intuitive and inclusive human-robot interaction. Unlike conventional approaches that rely on gloss annotations as intermediate…

The integration of electric vehicles (EVs) into smart grids presents unique opportunities to enhance both transportation systems and energy networks. However, ensuring safe and interpretable interactions between drivers, vehicles, and the…

Virtual assistants (VAs) have become ubiquitous in daily life, integrated into smartphones and smart devices, sparking interest in AI companions that enhance user experiences and foster emotional connections. However, existing companions…

人机交互 · 计算机科学 2025-09-03 Xuetong Wang , Ching Christie Pang , Pan Hui

People with visual impairments often face significant challenges in locating and retrieving objects in their surroundings. Existing assistive technologies present a trade-off: systems that offer precise guidance typically require…

When a multimodal Transformer answers a visual question, is the prediction driven by visual evidence, linguistic reasoning, or genuinely fused cross-modal computation -- and how does this structure evolve across layers? We address this…

人工智能 · 计算机科学 2026-02-18 Hongxuan Wu , Yukun Zhang , Xueqing Zhou

Sign language is a natural and visual form of language that uses movements and expressions to convey meaning, serving as a crucial means of communication for individuals who are deaf or hard-of-hearing (DHH). However, the number of people…

机器人学 · 计算机科学 2026-02-27 Guanren Qiao , Sixu Lin , Ronglai Zuo , Zhizheng Wu , Kui Jia , Guiliang Liu

Audio-Visual Scene-Aware Dialog (AVSD) is a task to generate responses when chatting about a given video, which is organized as a track of the 8th Dialog System Technology Challenge (DSTC8). To solve the task, we propose a universal…

计算与语言 · 计算机科学 2020-02-04 Zekang Li , Zongjia Li , Jinchao Zhang , Yang Feng , Cheng Niu , Jie Zhou

Open-vocabulary keyword spotting (KWS) in continuous speech streams holds significant practical value across a wide range of real-world applications. While increasing attention has been paid to the role of different modalities in KWS, their…

声音 · 计算机科学 2025-12-18 Kewei Li , Yinan Zhong , Xiaotao Liang , Tianchi Dai , Shaofei Xue

Visually impaired people face significant challenges when attempting to interact with and understand complex environments, and traditional assistive technologies often struggle to quickly provide necessary contextual understanding and…

人机交互 · 计算机科学 2025-05-02 Bhanuja Ainary

With the recent advancements in Artificial Intelligence (AI), Intelligent Virtual Assistants (IVA) such as Alexa, Google Home, etc., have become a ubiquitous part of many homes. Currently, such IVAs are mostly audio-based, but going…

多媒体 · 计算机科学 2019-12-27 Shachi H Kumar , Eda Okur , Saurav Sahay , Jonathan Huang , Lama Nachman

The rapid development of Large Language Models (LLMs) creates an exciting potential for flexible, general knowledge-driven Human-Robot Interaction (HRI) systems for assistive robots. Existing HRI systems demonstrate great progress in…

机器人学 · 计算机科学 2025-07-22 Jens V. Rüppel , Andrey Rudenko , Tim Schreiter , Martin Magnusson , Achim J. Lilienthal

This paper proposes a novel direct Audio-Visual Speech to Audio-Visual Speech Translation (AV2AV) framework, where the input and output of the system are multimodal (i.e., audio and visual speech). With the proposed AV2AV, two key…

计算机视觉与模式识别 · 计算机科学 2024-03-27 Jeongsoo Choi , Se Jin Park , Minsu Kim , Yong Man Ro

The recent integration of artificial intelligence into medical imaging has driven remarkable advances in automated organ segmentation. However, most existing 3D segmentation frameworks rely exclusively on visual learning from large…

计算机视觉与模式识别 · 计算机科学 2025-12-30 Hasan Faraz Khan , Noor Fatima , Muzammil Behzad

Sign language is a fundamental means of communication for the deaf and hard-of-hearing (DHH) community, enabling nuanced expression through gestures, facial expressions, and body movements. Despite its critical role in facilitating…

计算机视觉与模式识别 · 计算机科学 2025-04-14 Alexander Brettmann , Jakob Grävinghoff , Marlene Rüschoff , Marie Westhues

Multimodal semantic segmentation has shown great potential in leveraging complementary information across diverse sensing modalities. However, existing approaches often rely on carefully designed fusion strategies that either use…

计算机视觉与模式识别 · 计算机科学 2026-04-06 Zelin Zhang , Kedi Li , Huiqi Liang , Tao Zhang , Chuanzhi Xu

Deaf people are more heavily affected by the digital divide than many would expect. Moreover, most accessibility guidelines addressing their needs just deal with captioning and audio-content transcription. However, this approach to the…

人机交互 · 计算机科学 2025-09-03 Fabrizio Borgia , Claudia S. Bianchini , Maria de Marsico

Screen readers can drive braille devices for allowing visually impaired users to access computer environments, by providing them the same information as sighted users. But in some cases, this view is not easy to use on a braille device. In…

人机交互 · 计算机科学 2007-05-23 Samuel Thibault , Sebastien Hinderer

Visual speech recognition (VSR), also known as lip reading, is the task of recognizing speech from silent video. Despite significant advancements in VSR over recent decades, most existing methods pay limited attention to real-world visual…

计算机视觉与模式识别 · 计算机科学 2025-09-29 Tianyue Wang , Shuang Yang , Shiguang Shan , Xilin Chen

This paper presents a wearable assistive device with the shape of a pair of eyeglasses that allows visually impaired people to navigate safely and quickly in unfamiliar environment, as well as perceive the complicated environment to…

计算机视觉与模式识别 · 计算机科学 2019-06-25 Jinqiang Bai , Zhaoxiang Liu , Yimin Lin , Ye Li , Shiguo Lian , Dijun Liu