English
Related papers

Related papers: Vinci: A Real-time Embodied Smart Assistant based …

200 papers

Embodied Visual Reasoning (EVR) seeks to follow complex, free-form instructions based on egocentric video, enabling semantic understanding and spatiotemporal reasoning in dynamic environments. Despite its promising potential, EVR encounters…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Kailing Li , Qi'ao Xu , Tianwen Qian , Yuqian Fu , Yang Jiao , Xiaoling Wang

Current upper extremity outcome measures for persons with cervical spinal cord injury (cSCI) lack the ability to directly collect quantitative information in home and community environments. A wearable first-person (egocentric) camera…

Human-Computer Interaction · Computer Science 2024-05-07 Jirapat Likitlersuang , Elizabeth R. Sumitro , Tianshi Cao , Ryan J. Visee , Sukhvinder Kalsi-Ryan , Jose Zariffa

Mobile virtual reality (VR) head mounted displays (HMD) have become popular among consumers in recent years. In this work, we demonstrate real-time egocentric hand gesture detection and localization on mobile HMDs. Our main contributions…

Computer Vision and Pattern Recognition · Computer Science 2017-12-15 Rohit Pandey , Marie White , Pavel Pidlypenskyi , Xue Wang , Christine Kaeser-Chen

The pursuit of artificial general intelligence (AGI) has placed embodied intelligence at the forefront of robotics research. Embodied intelligence focuses on agents capable of perceiving, reasoning, and acting within the physical world.…

We present a universal framework to model contextualized sentence representations with visual awareness that is motivated to overcome the shortcomings of the multimodal parallel data with manual annotations. For each sentence, we first…

Computation and Language · Computer Science 2019-11-12 Zhuosheng Zhang , Rui Wang , Kehai Chen , Masao Utiyama , Eiichiro Sumita , Hai Zhao

Recent end-to-end spoken dialogue models enable natural interaction. However, as user demands become increasingly complex, models that rely solely on conversational abilities often struggle to cope. Incorporating agentic capabilities is…

Sound · Computer Science 2026-04-20 Tianle Liang , Yifu Chen , Shengpeng Ji , Yijun Chen , Zhiyang Jia , Jingyu Lu , Fan Zhuo , Xueyi Pu , Yangzhuo Li , Zhou Zhao

Generating long, coherent egocentric videos is difficult, as hand-object interactions and procedural tasks require reliable long-term memory. Existing autoregressive models suffer from content drift, where object identity and scene…

Computer Vision and Pattern Recognition · Computer Science 2025-12-05 Liuzhou Zhang , Jiarui Ye , Yuanlei Wang , Ming Zhong , Mingju Cao , Wanke Xia , Bowen Zeng , Zeyu Zhang , Hao Tang

It is challenging for humans -- particularly those living with physical disabilities -- to control high-dimensional, dexterous robots. Prior work explores learning embedding functions that map a human's low-dimensional inputs (e.g., via a…

Robotics · Computer Science 2021-05-04 Siddharth Karamcheti , Albert J. Zhai , Dylan P. Losey , Dorsa Sadigh

Simulation has the potential to transform the development of robust algorithms for mobile agents deployed in safety-critical scenarios. However, the poor photorealism and lack of diverse sensor modalities of existing simulation engines…

In today's society, where independent living is becoming increasingly important, it can be extremely constricting for those who are blind. Blind and visually impaired (BVI) people face challenges because they need manual support to prompt…

Human-Computer Interaction · Computer Science 2023-03-15 Malay Joshi , Aditi Shukla , Jayesh Srivastava , Manya Rastogi

Vision-language models have shown impressive progress in recent years. However, existing models are largely limited to turn-based interactions, where each turn must be stepped (i.e., prompted) by the user. Open-ended, asynchronous…

In this paper, we introduce a new problem, Online-MMSI, where the model must perform multimodal social interaction understanding (MMSI) using only historical information. Given a recorded video and a multi-party dialogue, the AI assistant…

Computer Vision and Pattern Recognition · Computer Science 2026-04-02 Xinpeng Li , Shijian Deng , Bolin Lai , Weiguo Pian , James M. Rehg , Yapeng Tian

Visually impaired people face numerous challenges when it comes to transportation. Not only must they circumvent obstacles while navigating, but they also need access to essential information related to available public transport,…

Human-Computer Interaction · Computer Science 2017-03-08 Gourav G. Shenoy , Mangirish A. Wagle , Kay Connelly

Virtual Reality (VR) is inaccessible to blind people. While research has investigated many techniques to enhance VR accessibility, they require additional developer effort to integrate. As such, most mainstream VR apps remain inaccessible…

Human-Computer Interaction · Computer Science 2025-08-06 Daniel Killough , Justin Feng , Zheng Xue "ZX" Ching , Daniel Wang , Rithvik Dyava , Yapeng Tian , Yuhang Zhao

Wearable collaborative robots stand to assist human wearers who need fall prevention assistance or wear exoskeletons. Such a robot needs to be able to constantly adapt to the surrounding scene based on egocentric vision, and predict the ego…

Computer Vision and Pattern Recognition · Computer Science 2024-08-08 Weizhuo Wang , C. Karen Liu , Monroe Kennedy

We present ADVISER - an open-source, multi-domain dialog system toolkit that enables the development of multi-modal (incorporating speech, text and vision), socially-engaged (e.g. emotion recognition, engagement level prediction and…

Embodied AI models often employ off the shelf vision backbones like CLIP to encode their visual observations. Although such general purpose representations encode rich syntactic and semantic information about the scene, much of this…

Computer Vision and Pattern Recognition · Computer Science 2024-03-12 Ainaz Eftekhar , Kuo-Hao Zeng , Jiafei Duan , Ali Farhadi , Ani Kembhavi , Ranjay Krishna

Egocentric video generation with fine-grained control through body motion is a key requirement towards embodied AI agents that can simulate, predict, and plan actions. In this work, we propose EgoControl, a pose-controllable video diffusion…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Enrico Pallotta , Sina Mokhtarzadeh Azar , Lars Doorenbos , Serdar Ozsoy , Umar Iqbal , Juergen Gall

Smart space management can be done in many ways. On one hand, there are conversational assistants such as the Google Assistant or Amazon Alexa that enable users to comfortably interact with smart spaces with only their voice, but these have…

Human-Computer Interaction · Computer Science 2018-07-19 André Sousa Lago , Hugo Sereno Ferreira

This paper investigates the implementation of voice-enabled Google Assistant and Amazon Alexa on Raspberry Pi. Virtual Assistants are being a new trend in how we interact or do computations with physical devices. A voice-enabled system…

Computers and Society · Computer Science 2024-09-05 Shailesh D. Arya , Samir Patel
‹ Prev 1 8 9 10 Next ›