中文
相关论文

相关论文: Explainable Action Form Assessment by Exploiting M…

200 篇论文

Understanding human actions in visual data is tied to advances in complementary research areas including object recognition, human dynamics, domain adaptation and semantic segmentation. Over the last decade, human action analysis evolved…

计算机视觉与模式识别 · 计算机科学 2017-02-02 Samitha Herath , Mehrtash Harandi , Fatih Porikli

Action recognition is a fundamental ability for social species. Yet, its underlying computations are not well understood. Classical psychophysical studies using simplified stimuli have shown that humans can perceive body motion even under…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Prerana Kumar , Martin A. Giese

Wearable HAR has improved steadily, but most progress still relies on closed-set classification, which limits real-world use. In practice, human activity is open-ended, unscripted, personalized, and often compositional, unfolding as…

机器学习 · 计算机科学 2026-04-02 Lala Shakti Swarup Ray , Mengxi Liu , Alcina Pinto , Deepika Gurung , Daniel Geissler , Paul Lukowoicz , Bo Zhou

Vision-Language-Action models have emerged as essential generalist robot policies for diverse manipulation tasks, conventionally relying on directly translating multimodal inputs into actions via Vision-Language Model embeddings. Recent…

机器人学 · 计算机科学 2026-03-31 Linqing Zhong , Yi Liu , Yifei Wei , Ziyu Xiong , Maoqing Yao , Si Liu , Guanghui Ren

Chain-of-thought (CoT) reasoning has emerged as a powerful tool for multimodal large language models on video understanding tasks. However, its necessity and advantages over direct answering remain underexplored. In this paper, we first…

Evaluating human actions with clear and detailed feedback is important in areas such as sports, healthcare, and robotics, where decisions rely not only on final outcomes but also on interpretable reasoning. However, most existing methods…

计算机视觉与模式识别 · 计算机科学 2025-08-26 Junhao Wu , Xiuer Gu , Zhiying Li , Yeying Jin , Yunfeng Diao , Zhiyu Li , Zhenbo Song , Xiaomei Zhang , Zhaoxin Fan

In recent years, model explanation methods have been designed to interpret model decisions faithfully and intuitively so that users can easily understand them. In this paper, we propose a framework, Faithful Attention Explainer (FAE),…

计算与语言 · 计算机科学 2024-05-28 Yao Rong , David Scheerer , Enkelejda Kasneci

We make available to the community a new dataset to support action-recognition research. This dataset is different from prior datasets in several key ways. It is significantly larger. It contains streaming video with long segments…

计算机视觉与模式识别 · 计算机科学 2015-11-19 Daniel Paul Barrett , Ran Xu , Haonan Yu , Jeffrey Mark Siskind

Beyond possessing large enough size to feed data hungry machines (eg, transformers), what attributes measure the quality of a dataset? Assuming that the definitions of such attributes do exist, how do we quantify among their relative…

计算机视觉与模式识别 · 计算机科学 2022-04-19 Rajat Modi , Aayush Jung Rana , Akash Kumar , Praveen Tirupattur , Shruti Vyas , Yogesh Singh Rawat , Mubarak Shah

Existing Causal-Why Video Question Answering (VideoQA) models often struggle with higher-order reasoning, relying on opaque, monolithic pipelines that entangle video understanding, causal inference, and answer generation. These black-box…

计算机视觉与模式识别 · 计算机科学 2025-12-25 Paritosh Parmar , Eric Peh , Basura Fernando

In recent years, the widespread adoption of wearable devices has highlighted the growing importance of behavior analysis using IMU. While applications span diverse fields such as healthcare and robotics, recent studies have increasingly…

计算机视觉与模式识别 · 计算机科学 2025-06-05 Koki Matsuishi , Kosuke Ukita , Tsuyoshi Okita

Long-term Action Quality Assessment (AQA) evaluates the execution of activities in videos. However, the length presents challenges in fine-grained interpretability, with current AQA methods typically producing a single score by averaging…

计算机视觉与模式识别 · 计算机科学 2024-08-22 Xu Dong , Xinran Liu , Wanqing Li , Anthony Adeyemi-Ejeye , Andrew Gilbert

Explainable AI (XAI) is critical for building trust in complex machine learning models, yet mainstream attribution methods often provide an incomplete, static picture of a model's final state. By collapsing a feature's role into a single…

机器学习 · 计算机科学 2025-11-03 Hamed Najafi , Dongsheng Luo , Jason Liu

Understanding the decision process of neural networks is hard. One vital method for explanation is to attribute its decision to pivotal features. Although many algorithms are proposed, most of them solely improve the faithfulness to the…

人工智能 · 计算机科学 2022-09-07 Yuyou Gan , Yuhao Mao , Xuhong Zhang , Shouling Ji , Yuwen Pu , Meng Han , Jianwei Yin , Ting Wang

Identifying human behaviors is a challenging research problem due to the complexity and variation of appearances and postures, the variation of camera settings, and view angles. In this paper, we try to address the problem of human behavior…

计算机视觉与模式识别 · 计算机科学 2019-03-08 Eissa Jaber Alreshidi , Mohammad Bilal

Despite recent progress in video large language models (VideoLLMs), a key open challenge remains: how to equip models with chain-of-thought (CoT) reasoning abilities grounded in fine-grained object-level video understanding. Existing…

计算机视觉与模式识别 · 计算机科学 2025-07-21 Yanan Wang , Julio Vizcarra , Zhi Li , Hao Niu , Mori Kurokawa

Event cameras have recently been shown beneficial for practical vision tasks, such as action recognition, thanks to their high temporal resolution, power efficiency, and reduced privacy concerns. However, current research is hindered by 1)…

计算机视觉与模式识别 · 计算机科学 2024-03-20 Jiazhou Zhou , Xu Zheng , Yuanhuiyi Lyu , Lin Wang

Automated video surveillance with Large Vision-Language Models is limited by their inherent bias towards normality, often failing to detect crimes. While Chain-of-Thought reasoning strategies show significant potential for improving…

计算机视觉与模式识别 · 计算机科学 2025-12-24 Pedro Domingos , João Pereira , Vasco Lopes , João Neves , David Semedo

Nearly all existing Facial Action Coding System-based datasets that include facial action unit (AU) intensity information annotate the intensity values hierarchically using A--E levels. However, facial expressions change continuously and…

计算机视觉与模式识别 · 计算机科学 2023-01-31 Wei Gan , Jian Xue , Ke Lu , Yanfu Yan , Pengcheng Gao , Jiayi Lyu

In the context of fitness coaching or for rehabilitation purposes, the motor actions of a human participant must be observed and analyzed for errors in order to provide effective feedback. This task is normally carried out by human coaches,…

人工智能 · 计算机科学 2017-09-27 Felix Hülsmann , Stefan Kopp , Mario Botsch