中文
相关论文

相关论文: A Large-Scale Multimodal Dataset and Benchmarks fo…

200 篇论文

In the realm of human activity recognition (HAR), the integration of explainable Artificial Intelligence (XAI) emerges as a critical necessity to elucidate the decision-making processes of complex models, fostering transparency and trust.…

人工智能 · 计算机科学 2024-08-22 Yiran Huang , Yexu Zhou , Haibin Zhao , Till Riedel , Michael Beigl

Human activity recognition (HAR) is a rapidly growing field that utilizes smart devices, sensors, and algorithms to automatically classify and identify the actions of individuals within a given environment. These systems have a wide range…

The combination of increased life expectancy and falling birth rates is resulting in an aging population. Wearable Sensor-based Human Activity Recognition (WSHAR) emerges as a promising assistive technology to support the daily lives of…

信号处理 · 电气工程与系统科学 2024-04-25 Jianyuan Ni , Hao Tang , Syed Tousiful Haque , Yan Yan , Anne H. H. Ngu

This work presents a novel architecture for context-aware interactions within smart environments, leveraging Large Language Models (LLMs) to enhance user experiences. Our system integrates user location data obtained through UWB tags and…

计算与语言 · 计算机科学 2025-02-21 Aurora Polo-Rodríguez , Laura Fiorini , Erika Rovini , Filippo Cavallo , Javier Medina-Quero

Multimodal systems have great potential to assist humans in procedural activities, where people follow instructions to achieve their goals. Despite diverse application scenarios, systems are typically evaluated on traditional classification…

Complex human activity recognition (CHAR) remains a pivotal challenge within ubiquitous computing, especially in the context of smart environments. Existing studies typically require meticulous labeling of both atomic and complex…

人工智能 · 计算机科学 2024-08-07 Yuan Sun , Navid Salami Pargoo , Taqiya Ehsan , Zhao Zhang , Jorge Ortiz

While Large Language Models (LLM) enable non-experts to specify open-world multi-robot tasks, the generated plans often lack kinematic feasibility and are not efficient, especially in long-horizon scenarios. Formal methods like Linear…

机器人学 · 计算机科学 2026-02-11 Shuyuan Hu , Tao Lin , Kai Ye , Yang Yang , Tianwei Zhang

Human Activity Recognition (HAR) plays a significant role in the everyday life of people because of its ability to learn extensive high-level information about human activity from wearable or stationary devices. A substantial amount of…

信号处理 · 电气工程与系统科学 2022-09-09 Md. Milon Islam , Sheikh Nooruddin , Fakhri Karray , Ghulam Muhammad

Human activity recognition (HAR) based on multimodal sensors has become a rapidly growing branch of biometric recognition and artificial intelligence. However, how to fully mine multimodal time series data and effectively learn accurate…

计算机视觉与模式识别 · 计算机科学 2022-05-25 Jialiang Wang , Haotian Wei , Yi Wang , Shu Yang , Chi Li

Benefiting from strong and efficient multi-modal alignment strategies, Large Visual Language Models (LVLMs) are able to simulate human visual and reasoning capabilities, such as solving CAPTCHAs. However, existing benchmarks based on visual…

人工智能 · 计算机科学 2025-12-15 Jianyi Zhang , Ziyin Zhou , Xu Ji , Shizhao Liu , Zhangchi Zhao

Automated co-located human-human interaction analysis has been addressed by the use of nonverbal communication as measurable evidence of social and psychological phenomena. We survey the computing studies (since 2010) detecting phenomena…

人机交互 · 计算机科学 2023-10-05 Cigdem Beyan , Alessandro Vinciarelli , Alessio Del Bue

Conversational human-likeness plays a central role in human-AI interaction, yet it has remained difficult to define, measure, and optimize. As a result, improvements in human-like behavior are largely driven by scale or broad supervised…

人工智能 · 计算机科学 2026-01-08 Masum Hasan , Junjie Zhao , Ehsan Hoque

Generation of images containing multiple humans, performing complex actions, while preserving their facial identities, is a significant challenge. A major factor contributing to this is the lack of a dedicated benchmark. To address this, we…

计算机视觉与模式识别 · 计算机科学 2026-01-23 Shubhankar Borse , Seokeon Choi , Sunghyun Park , Jeongho Kim , Shreya Kadambi , Risheek Garrepalli , Sungrack Yun , Munawar Hayat , Fatih Porikli

Spatial reasoning has emerged as a critical capability for Multimodal Large Language Models (MLLMs), drawing increasing attention and rapid advancement. However, existing benchmarks primarily focus on single-step perception-to-judgment…

计算机视觉与模式识别 · 计算机科学 2026-01-15 Rui Zhu , Xin Shen , Shuchen Wu , Chenxi Miao , Xin Yu , Yang Li , Weikang Li , Deguo Xia , Jizhou Huang

Wearable systems can recognize activities from IMU data but often fail to explain their underlying causes or contextual significance. To address this limitation, we introduce two large-scale resources: SensorCap, comprising 35,960…

计算与语言 · 计算机科学 2025-09-23 Sheikh Asif Imran , Mohammad Nur Hossain Khan , Subrata Biswas , Bashima Islam

Recent advances in multimodal AI have enabled progress in detecting synthetic and out-of-context content. However, existing efforts largely overlook the intent behind AI-generated images. To fill this gap, we introduce S-HArM, a multimodal…

计算机视觉与模式识别 · 计算机科学 2025-09-10 Anastasios Skoularikis , Stefanos-Iordanis Papadopoulos , Symeon Papadopoulos , Panagiotis C. Petrantonakis

It is expensive and time-consuming to collect sufficient labeled data to build human activity recognition (HAR) models. Training on existing data often makes the model biased towards the distribution of the training data, thus the model…

人工智能 · 计算机科学 2022-06-15 Wang Lu , Jindong Wang , Yiqiang Chen , Sinno Jialin Pan , Chunyu Hu , Xin Qin

Despite substantial progress in video understanding, most existing datasets are limited to Earth's gravitational conditions. However, microgravity alters human motion, interactions, and visual semantics, revealing a critical gap for…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Di Wen , Lei Qi , Kunyu Peng , Kailun Yang , Fei Teng , Ao Luo , Jia Fu , Yufan Chen , Ruiping Liu , Yitian Shi , M. Saquib Sarfraz , Rainer Stiefelhagen

Multimodal Large Language Models (MLLMs) have demonstrated significant advances across numerous vision-language tasks. MLLMs have shown promising capability in aligning visual and textual modalities, allowing them to process image-text…

计算与语言 · 计算机科学 2025-09-29 Xiaolong Wang , Zhaolu Kang , Wangyuxuan Zhai , Xinyue Lou , Yunghwei Lai , Ziyue Wang , Yawen Wang , Kaiyu Huang , Yile Wang , Peng Li , Yang Liu

Human activity recognition (HAR) by wearable sensor devices embedded in the Internet of things (IOT) can play a significant role in remote health monitoring and emergency notification, to provide healthcare of higher standards. The purpose…

机器学习 · 计算机科学 2022-01-24 M. Abid , A. Khabou , Y. Ouakrim , H. Watel , S. Chemkhi , A. Mitiche , A. Benazza-Benyahia , N. Mezghani