中文
相关论文

相关论文: Go Beyond Earth: Understanding Human Actions and S…

200 篇论文

Distilling knowledge from human demonstrations is a promising way for robots to learn and act. Existing methods, which often rely on coarsely-aligned video pairs, are typically constrained to learning global or task-level features. As a…

机器人学 · 计算机科学 2025-11-18 Sicheng Xie , Haidong Cao , Zejia Weng , Zhen Xing , Haoran Chen , Shiwei Shen , Jiaqi Leng , Zuxuan Wu , Yu-Gang Jiang

Action Detection is a complex task that aims to detect and classify human actions in video clips. Typically, it has been addressed by processing fine-grained features extracted from a video classification backbone. Recently, thanks to the…

计算机视觉与模式识别 · 计算机科学 2021-03-02 Matteo Tomei , Lorenzo Baraldi , Simone Calderara , Simone Bronzin , Rita Cucchiara

Generating reasonable and high-quality human interactive motions in a given dynamic environment is crucial for understanding, modeling, transferring, and applying human behaviors to both virtual and physical robots. In this paper, we…

计算机视觉与模式识别 · 计算机科学 2025-03-04 Peishan Cong , Ziyi Wang , Yuexin Ma , Xiangyu Yue

We argue that progress in true multimodal intelligence calls for a shift from reactive, task-driven systems and brute-force long context towards a broader paradigm of supersensing. We frame spatial supersensing as four stages beyond…

计算机视觉与模式识别 · 计算机科学 2025-11-07 Shusheng Yang , Jihan Yang , Pinzhi Huang , Ellis Brown , Zihao Yang , Yue Yu , Shengbang Tong , Zihan Zheng , Yifan Xu , Muhan Wang , Daohan Lu , Rob Fergus , Yann LeCun , Li Fei-Fei , Saining Xie

Physical video understanding requires more than naming an event correctly. A model can answer a question about pouring, sliding, or collision from textual regularities while still failing to localize the event in time or space. We introduce…

计算机视觉与模式识别 · 计算机科学 2026-04-24 Alibay Osmanli , Zixu Cheng , Shaogang Gong

Characterizing and quantifying gender representation disparities in audiovisual storytelling contents is necessary to grasp how stereotypes may perpetuate on screen. In this article, we consider the high-level construct of objectification…

Generating dialogue grounded in videos requires a high level of understanding and reasoning about the visual scenes in the videos. However, existing large visual-language models are not effective due to their latent features and…

计算机视觉与模式识别 · 计算机科学 2023-11-23 Hongcheng Liu , Zhe Chen , Hui Li , Pingjie Wang , Yanfeng Wang , Yu Wang

When video reasoning requires external knowledge, many systems with large multimodal models (LMMs) adopt retrieval augmentation to supply the missing context. Appending textual or multi-clip evidence, however, forces heterogeneous signals…

计算机视觉与模式识别 · 计算机科学 2026-04-07 Songyuan Yang , Weijiang Yu , Ziyu Liu , Guijian Tang , Wenjing Yang , Huibin Tan , Nong Xiao

In this paper, we address the problem of referring expression comprehension in videos, which is challenging due to complex expression and scene dynamics. Unlike previous methods which solve the problem in multiple stages (i.e., tracking,…

计算机视觉与模式识别 · 计算机科学 2021-03-24 Sijie Song , Xudong Lin , Jiaying Liu , Zongming Guo , Shih-Fu Chang

Understanding human motion beyond surface kinematics is crucial for motion analysis, rehabilitation, and injury risk assessment. However, progress in this domain is limited by the lack of large-scale datasets with biomechanical annotations,…

计算机视觉与模式识别 · 计算机科学 2026-05-15 Yujun Huo , He Zhang , Chentao Song , Honglin Song , Zongyu Zuo , Tao Yu

In this work, we introduce a novel task - Humancentric Spatio-Temporal Video Grounding (HC-STVG). Unlike the existing referring expression tasks in images or videos, by focusing on humans, HC-STVG aims to localize a spatiotemporal tube of…

计算机视觉与模式识别 · 计算机科学 2021-06-03 Zongheng Tang , Yue Liao , Si Liu , Guanbin Li , Xiaojie Jin , Hongxu Jiang , Qian Yu , Dong Xu

Human motion video generation has garnered significant research interest due to its broad applications, enabling innovations such as photorealistic singing heads or dynamic avatars that seamlessly dance to music. However, existing surveys…

Accurate video understanding involves reasoning about the relationships between actors, objects and their environment, often over long temporal intervals. In this paper, we propose a message passing graph neural network that explicitly…

计算机视觉与模式识别 · 计算机科学 2021-03-30 Anurag Arnab , Chen Sun , Cordelia Schmid

Imitation learning from human demonstrations is a promising paradigm for teaching robots manipulation skills in the real world. However, learning complex long-horizon tasks often requires an unattainable amount of demonstrations. To reduce…

机器人学 · 计算机科学 2023-10-16 Chen Wang , Linxi Fan , Jiankai Sun , Ruohan Zhang , Li Fei-Fei , Danfei Xu , Yuke Zhu , Anima Anandkumar

A core capability towards general embodied intelligence lies in localizing task-relevant objects from an egocentric perspective, formulated as Spatio-Temporal Video Grounding (STVG). Despite recent progress, existing STVG studies remain…

计算机视觉与模式识别 · 计算机科学 2026-04-07 Qi'ao Xu , Tianwen Qian , Yuqian Fu , Kailing Li , Yang Jiao , Jiacheng Zhang , Xiaoling Wang , Liang He

Understanding how people interact with their surroundings and each other is essential for enabling robots to act in socially compliant and context-aware ways. While 3D Scene Graphs have emerged as a powerful semantic representation for…

计算机视觉与模式识别 · 计算机科学 2025-09-30 Ermanno Bartoli , Dennis Rotondi , Buwei He , Patric Jensfelt , Kai O. Arras , Iolanda Leite

Gait recognition is an emerging biometric technology that enables non-intrusive and hard-to-spoof human identification. However, most existing methods are confined to short-range, unimodal settings and fail to generalize to long-range and…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Zhiyang Lu , Wen Jiang , Tianren Wu , Zhichao Wang , Changwang Zhang , Siqi Shen , Ming Cheng

Human actions often involve complex interactions across several inter-related objects in the scene. However, existing approaches to fine-grained video understanding or visual relationship detection often rely on single object representation…

计算机视觉与模式识别 · 计算机科学 2018-03-22 Chih-Yao Ma , Asim Kadav , Iain Melvin , Zsolt Kira , Ghassan AlRegib , Hans Peter Graf

Understanding dynamic 4D scenes from an egocentric perspective-modeling changes in 3D spatial structure over time-is crucial for human-machine interaction, autonomous navigation, and embodied intelligence. While existing egocentric datasets…

计算机视觉与模式识别 · 计算机科学 2025-11-18 Junsheng Huang , Shengyu Hao , Bocheng Hu , Hongwei Wang , Gaoang Wang

Video captioning is a challenging task since it requires generating sentences describing various diverse and complex videos. Existing video captioning models lack adequate visual representation due to the neglect of the existence of gaps…

计算机视觉与模式识别 · 计算机科学 2021-10-14 Mingkang Tang , Zhanyu Wang , Zhenhua Liu , Fengyun Rao , Dian Li , Xiu Li