中文
相关论文

相关论文: EgoLog: Ego-Centric Fine-Grained Daily Log with Ub…

200 篇论文

As the prevalence of wearable devices, learning egocentric motions becomes essential to develop contextual AI. In this work, we present EgoLM, a versatile framework that tracks and understands egocentric motions from multi-modal inputs,…

计算机视觉与模式识别 · 计算机科学 2024-09-27 Fangzhou Hong , Vladimir Guzov , Hyo Jin Kim , Yuting Ye , Richard Newcombe , Ziwei Liu , Lingni Ma

While on-body device-based human motion estimation is crucial for applications such as XR interaction, existing methods often suffer from poor wearability, expensive hardware, and cumbersome calibration, which hinder their adoption in daily…

计算机视觉与模式识别 · 计算机科学 2025-12-25 Siqi Zhu , Yixuan Li , Junfu Li , Qi Wu , Zan Wang , Haozhe Ma , Wei Liang

Combining different sensing modalities with multiple positions helps form a unified perception and understanding of complex situations such as human behavior. Hence, human activity recognition (HAR) benefits from combining redundant and…

机器学习 · 计算机科学 2024-04-26 Hymalai Bello

Human activity recognition (HAR) on smartglasses has various use cases, including health/fitness tracking and input for context-aware AI assistants. However, current approaches for egocentric activity recognition suffer from low performance…

计算机视觉与模式识别 · 计算机科学 2025-04-25 Akhil Padmanabha , Saravanan Govindarajan , Hwanmun Kim , Sergio Ortiz , Rahul Rajan , Doruk Senkal , Sneha Kadetotad

This work focuses on tracking and understanding human motion using consumer wearable devices, such as VR/AR headsets, smart glasses, cellphones, and smartwatches. These devices provide diverse, multi-modal sensor inputs, including…

计算机视觉与模式识别 · 计算机科学 2025-04-14 Jian Wang , Rishabh Dabral , Diogo Luvizon , Zhe Cao , Lingjie Liu , Thabo Beeler , Christian Theobalt

The integration of brain-computer interfaces (BCIs), in particular electroencephalography (EEG), with artificial intelligence (AI) has shown tremendous promise in decoding human cognition and behavior from neural signals. In particular, the…

人工智能 · 计算机科学 2025-10-15 Nie Lin , Yansen Wang , Dongqi Han , Weibang Jiang , Jingyuan Li , Ryosuke Furuta , Yoichi Sato , Dongsheng Li

Rich and context-aware activity logs facilitate user behavior analysis and health monitoring, making them a key research focus in ubiquitous computing. The remarkable semantic understanding and generation capabilities of Large Language…

人工智能 · 计算机科学 2025-07-21 Ye Tian , Xiaoyuan Ren , Zihao Wang , Onat Gungor , Xiaofan Yu , Tajana Rosing

All-day smart glasses are likely to emerge as platforms capable of continuous contextual sensing, uniquely positioning them for unprecedented assistance in our daily lives. Integrating the multi-modal AI agents required for human memory…

计算机视觉与模式识别 · 计算机科学 2025-08-05 Akshay Paruchuri , Sinan Hersek , Lavisha Aggarwal , Qiao Yang , Xin Liu , Achin Kulshrestha , Andrea Colaco , Henry Fuchs , Ishan Chatterjee

Multimodal Large Language Models (MLLMs) have recently achieved remarkable progress in vision-language understanding. Yet, human perception is inherently multisensory, integrating sight, sound, and motion to reason about the world. Among…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Bingwen Zhu , Yuqian Fu , Qiaole Dong , Guolei Sun , Tianwen Qian , Yuzheng Wu , Danda Pani Paudel , Xiangyang Xue , Yanwei Fu

The goal of creating intelligent, human-centered wearable systems for continuous activity understanding faces a fundamental trade-off: Egocentric video-based models capture rich semantic information and have demonstrated strong performance…

计算机视觉与模式识别 · 计算机科学 2026-04-22 Baiyu Chen , Wilson Wongso , Zechen Li , Yonchanok Khaokaew , Hao Xue , Flora Salim

Self-tracking has been long discussed, which can monitor daily activities and help users to recall previous experiences. Such data-capturing technique is no longer limited to photos, text messages, or personal diaries in recent years. With…

人机交互 · 计算机科学 2022-11-29 Jiyang Li , Ann Gina Konnayil , Adam Russell , Dingran Wang , Yincheng Jin , Seokmin Choi , Zhanpeng Jin

Human and environment sensing are two important topics in Computer Vision and Graphics. Human motion is often captured by inertial sensors, while the environment is mostly reconstructed using cameras. We integrate the two techniques…

计算机视觉与模式识别 · 计算机科学 2023-05-03 Xinyu Yi , Yuxiao Zhou , Marc Habermann , Vladislav Golyanik , Shaohua Pan , Christian Theobalt , Feng Xu

Despite the widespread integration of ambient light sensors (ALS) in smart devices commonly used for screen brightness adaptation, their application in human activity recognition (HAR), primarily through body-worn ALS, is largely…

人工智能 · 计算机科学 2024-08-23 Lala Shakti Swarup Ray , Daniel Geißler , Mengxi Liu , Bo Zhou , Sungho Suh , Paul Lukowicz

Human Activity Recognition (HAR) is a fundamental technology for numerous human - centered intelligent applications. Although deep learning methods have been utilized to accelerate feature extraction, issues such as multimodal data mixing,…

计算机视觉与模式识别 · 计算机科学 2026-02-06 Ying Yu , Siyao Li , Yixuan Jiang , Hang Xiao , Jingxi Long , Haotian Tang , Hanyu Liu , Chao Li

Understanding egocentric videos plays a vital role for embodied intelligence. Recent multi-modal large language models (MLLMs) can accept both visual and audio inputs. However, due to the challenge of obtaining text labels with coherent…

Emotion estimation in music listening is confronting challenges to capture the emotion variation of listeners. Recent years have witnessed attempts to exploit multimodality fusing information from musical contents and physiological signals…

人工智能 · 计算机科学 2016-12-01 Nattapong Thammasan , Ken-ichi Fukui , Masayuki Numao

We introduce EgoLife, a project to develop an egocentric life assistant that accompanies and enhances personal efficiency through AI-powered wearable glasses. To lay the foundation for this assistant, we conducted a comprehensive data…

Sensor data streams provide valuable information around activities and context for downstream applications, though integrating complementary information can be challenging. We show that large language models (LLMs) can be used for late…

The performance of speech emotion recognition (SER) is limited by the insufficient emotion information in unimodal systems and the feature alignment difficulties in multimodal systems. Recently, multimodal large language models (MLLMs) have…

声音 · 计算机科学 2025-09-22 Yiqing Yang , Man-Wai Mak

In the last years the pervasive use of sensors, as they exist in smart devices, e.g., phones, watches, medical devices, has increased dramatically the availability of personal data. However, existing research on data collection primarily…

人机交互 · 计算机科学 2025-03-26 Ivan Kayongo , Leonardo Malcotti , Haonan Zhao , Fausto Giunchiglia
‹ 上一页 1 2 3 10 下一页 ›