English
Related papers

Related papers: Vinci: A Real-time Embodied Smart Assistant based …

200 papers

Recent Omni-multimodal Large Language Models show promise in unified audio, vision, and text modeling. However, streaming audio-video understanding remains challenging, as existing approaches suffer from disjointed capabilities: they…

Computer Vision and Pattern Recognition · Computer Science 2026-01-16 Xueyun Tian , Wei Li , Bingbing Xu , Heng Dong , Yuanzhuo Wang , Huawei Shen

Mindfulness meditation is a validated means of helping people manage stress. Voice-based virtual assistants (VAs) in smart speakers, smartphones, and smart environments can assist people in carrying out mindfulness meditation through guided…

Human-Computer Interaction · Computer Science 2023-04-25 Bonhee Ku , Tatsuya Itagaki , Katie Seaborn

Recent advancements in Multi-modal Large Language Models (MLLMs) have opened new avenues for applications in Embodied AI. Building on previous work, EgoThink, we introduce VidEgoThink, a comprehensive benchmark for evaluating egocentric…

Computer Vision and Pattern Recognition · Computer Science 2024-10-16 Sijie Cheng , Kechen Fang , Yangyang Yu , Sicheng Zhou , Bohao Li , Ye Tian , Tingguang Li , Lei Han , Yang Liu

Vision-Language Models (VLMs) have shown great success as foundational models for downstream vision and natural language applications in a variety of domains. However, these models are limited to reasoning over objects and actions currently…

Robotics · Computer Science 2025-06-13 Zachary Chavis , Hyun Soo Park , Stephen J. Guy

We introduce a novel self-improving framework that enhances Embodied Visual Tracking (EVT) with Vision-Language Models (VLMs) to address the limitations of current active visual tracking systems in recovering from tracking failure. Our…

Computer Vision and Pattern Recognition · Computer Science 2025-05-29 Kui Wu , Shuhang Xu , Hao Chen , Churan Wang , Zhoujun Li , Yizhou Wang , Fangwei Zhong

Egocentric video-language pretraining is a crucial step in advancing the understanding of hand-object interactions in first-person scenarios. Despite successes on existing testbeds, we find that current EgoVLMs can be easily misled by…

Computer Vision and Pattern Recognition · Computer Science 2025-02-21 Boshen Xu , Ziheng Wang , Yang Du , Zhinan Song , Sipeng Zheng , Qin Jin

First-person video highlights a camera-wearer's activities in the context of their persistent environment. However, current video understanding approaches reason over visual features from short video clips that are detached from the…

Computer Vision and Pattern Recognition · Computer Science 2023-11-13 Tushar Nagarajan , Santhosh Kumar Ramakrishnan , Ruta Desai , James Hillis , Kristen Grauman

Assistive technologies for people with visual impairments (PVI) have made significant advancements, particularly with the integration of artificial intelligence (AI) and real-time sensor technologies. However, current solutions often…

Human-Computer Interaction · Computer Science 2024-10-08 He Zhang , Nicholas J. Falletta , Jingyi Xie , Rui Yu , Sooyeon Lee , Syed Masum Billah , John M. Carroll

In embodied AI, visual perception should be active rather than passive: the system must decide where to look and at what scale to sense to acquire maximally informative data under pixel and spatial budget constraints. Existing vision models…

Robotics · Computer Science 2026-04-06 Jiashu Yang , Yifan Han , Yucheng Xie , Ning Guo , Wenzhao Lian

Intelligent conversational agents and virtual assistants, such as chatbots and voice assistants, have the potential of augmenting health service capacity to screen symptoms and deliver healthcare interventions. In this paper, we developed…

Human-Computer Interaction · Computer Science 2022-02-07 Abdalsalam Almzayyen , Angel Vela de la Garza Evia , Nick Coronato , Mehdi Boukhechba

Passive tracking methods, such as phone and wearable sensing, have become dominant in monitoring human behaviors in modern ubiquitous computing studies. While there have been significant advances in machine-learning approaches to translate…

Human-Computer Interaction · Computer Science 2025-10-31 Jiachen Li , Xiwen Li , Justin Steinberg , Akshat Choube , Bingsheng Yao , Xuhai Xu , Dakuo Wang , Elizabeth Mynatt , Varun Mishra

The enhancement of generalization in robots by large vision-language models (LVLMs) is increasingly evident. Therefore, the embodied cognitive abilities of LVLMs based on egocentric videos are of great interest. However, current datasets…

Computer Vision and Pattern Recognition · Computer Science 2025-03-14 Ronghao Dang , Yuqian Yuan , Wenqi Zhang , Yifei Xin , Boqiang Zhang , Long Li , Liuyi Wang , Qinyang Zeng , Xin Li , Lidong Bing

Embodied AI is widely recognized as a cornerstone of artificial general intelligence (AGI) because it involves controlling embodied agents to perform tasks in the physical world. Building on the success of large language models (LLMs) and…

Robotics · Computer Science 2026-05-04 Yueen Ma , Zixing Song , Yuzheng Zhuang , Jianye Hao , Irwin King

Recent advancements in language and vision assistants have showcased impressive capabilities but suffer from a lack of transparency, limiting broader research and reproducibility. While open-source models handle general image tasks…

Computer Vision and Pattern Recognition · Computer Science 2024-10-08 Geewook Kim , Minjoon Seo

Intelligent assistive systems can navigate blind people, but most of them could only give non-intuitive cues or inefficient guidance. Based on computer vision and vibrotactile encoding, this paper presents an interactive system that…

Robotics · Computer Science 2022-06-22 Zhikai Wei , Xuhui Hu

As a long-term vision in the field of artificial intelligence, the core goal of embodied intelligence is to improve the perception, understanding, and interaction capabilities of agents and the environment. Vision-language navigation (VLN),…

Robotics · Computer Science 2024-03-21 Peng Gao , Peng Wang , Feng Gao , Fei Wang , Ruyue Yuan

We present EgoBlind, the first egocentric VideoQA dataset collected from blind individuals to evaluate the assistive capabilities of contemporary multimodal large language models (MLLMs). EgoBlind comprises 1,392 first-person videos from…

Computer Vision and Pattern Recognition · Computer Science 2025-11-04 Junbin Xiao , Nanxin Huang , Hao Qiu , Zhulin Tao , Xun Yang , Richang Hong , Meng Wang , Angela Yao

Brain-computer interface (BCI) systems have potential as assistive technologies for individuals with severe motor impairments. Nevertheless, individuals must first participate in many training sessions to obtain adequate data for optimizing…

Signal Processing · Electrical Eng. & Systems 2019-12-11 Behnam Reyhani-Masoleh , Tom Chau

We present Whole-Body Mobile Manipulation Interface (HoMMI), a data collection and policy learning framework that learns whole-body mobile manipulation directly from robot-free human demonstrations. We augment UMI interfaces with egocentric…

The growing adoption of augmented and virtual reality (AR and VR) technologies in industrial training and on-the-job assistance has created new opportunities for intelligent, context-aware support systems. As workers perform complex tasks…

Human-Computer Interaction · Computer Science 2025-11-18 Mahya Qorbani , Kamran Paynabar , Mohsen Moghaddam