中文
相关论文

相关论文: MEVA: A Large-Scale Multiview, Multimodal Video Da…

200 篇论文

This paper presents question-answering on dense video events, a novel task that answers and grounds dense-event questions in long videos, thus challenging MLLMs to faithfully comprehend and reason about multiple events over extended periods…

计算机视觉与模式识别 · 计算机科学 2025-05-19 Hangyu Qin , Junbin Xiao , Angela Yao

Accurate state estimation in Unmanned Aerial Vehicles (UAVs) is crucial for ensuring reliable and safe operation, as anomalies occurring during mission execution may induce discrepancies between expected and observed system behaviors,…

机器人学 · 计算机科学 2026-02-17 Aykut Kabaoglu , Sanem Sariel

In Actor and Observer we introduced a dataset linking the first and third-person video understanding domains, the Charades-Ego Dataset. In this paper we describe the egocentric aspect of the dataset and present annotations for Charades-Ego…

计算机视觉与模式识别 · 计算机科学 2018-05-01 Gunnar A. Sigurdsson , Abhinav Gupta , Cordelia Schmid , Ali Farhadi , Karteek Alahari

Being able to detect and recognize human activities is essential for several applications, including personal assistive robotics. In this paper, we perform detection and recognition of unstructured human activity in unstructured…

机器人学 · 计算机科学 2014-06-25 Jaeyong Sung , Colin Ponce , Bart Selman , Ashutosh Saxena

Recent long-form video-language understanding benchmarks have driven progress in video large multimodal models (Video-LMMs). However, the scarcity of well-annotated long videos has left the training of hour-long Video-LMMs underexplored. To…

计算机视觉与模式识别 · 计算机科学 2025-12-03 Jingyang Lin , Jialian Wu , Ximeng Sun , Ze Wang , Jiang Liu , Yusheng Su , Xiaodong Yu , Hao Chen , Jiebo Luo , Zicheng Liu , Emad Barsoum

Video recognition has been advanced in recent years by benchmarks with rich annotations. However, research is still mainly limited to human action or sports recognition - focusing on a highly specific video understanding task and thus…

计算机视觉与模式识别 · 计算机科学 2020-12-16 Ali Diba , Mohsen Fayyaz , Vivek Sharma , Manohar Paluri , Jurgen Gall , Rainer Stiefelhagen , Luc Van Gool

This paper presents a method for indexing activities of daily living in videos obtained from wearable cameras. In the context of dementia diagnosis by doctors, the videos are recorded at patients' houses and later visualized by the medical…

High-quality human reconstruction and photo-realistic rendering of a dynamic scene is a long-standing problem in computer vision and graphics. Despite considerable efforts invested in developing various capture systems and reconstruction…

计算机视觉与模式识别 · 计算机科学 2024-04-03 Xiaoyun Zheng , Liwei Liao , Xufeng Li , Jianbo Jiao , Rongjie Wang , Feng Gao , Shiqi Wang , Ronggang Wang

Mobile robots are reaching unprecedented speeds, with platforms like Unitree B2, and Fraunhofer O3dyn achieving maximum speeds between 5 and 10 m/s. However, effectively utilizing such speeds remains a challenge due to the limitations of…

计算机视觉与模式识别 · 计算机科学 2025-06-04 Shrutarv Awasthi , Anas Gouda , Sven Franke , Jérôme Rutinowski , Frank Hoffmann , Moritz Roidl

Our lives can be seen as a complex weaving of activities; we switch from one activity to another, to maximise our achievements or in reaction to demands placed upon us. Observing a video of unscripted daily activities, we parse the video…

计算机视觉与模式识别 · 计算机科学 2022-04-05 Will Price , Carl Vondrick , Dima Damen

Event cameras, with their high dynamic range (HDR) and low latency, offer a promising alternative for robust depth estimation in challenging environments. However, many event-based depth estimation approaches are constrained by small-scale…

计算机视觉与模式识别 · 计算机科学 2025-11-06 Sadiq Layi Macaulay , Nimet Kaygusuz , Simon Hadfield

This paper investigates the modeling of automated machine description on sports video, which has seen much progress recently. Nevertheless, state-of-the-art approaches fall quite short of capturing how human experts analyze sports scenes.…

计算机视觉与模式识别 · 计算机科学 2022-08-10 Dekun Wu , He Zhao , Xingce Bao , Richard P. Wildes

Daily Activity Recordings for Artificial Intelligence (DARai, pronounced "Dahr-ree") is a multimodal, hierarchically annotated dataset constructed to understand human activities in real-world settings. DARai consists of continuous scripted…

计算机视觉与模式识别 · 计算机科学 2026-03-30 Ghazal Kaviani , Yavuz Yarici , Seulgi Kim , Mohit Prabhushankar , Ghassan AlRegib , Mashhour Solh , Ameya Patil

Human body actions are an important form of non-verbal communication in social interactions. This paper specifically focuses on a subset of body actions known as micro-actions, which are subtle, low-intensity body movements with promising…

计算机视觉与模式识别 · 计算机科学 2025-08-21 Kun Li , Pengyu Liu , Dan Guo , Fei Wang , Zhiliang Wu , Hehe Fan , Meng Wang

Natural human-computer interaction and audio-visual human behaviour sensing systems, which would achieve robust performance in-the-wild are more needed than ever as digital devices are increasingly becoming an indispensable part of our…

Video object segmentation (VOS) aims to segment specified target objects throughout a video. Although state-of-the-art methods have achieved impressive performance (e.g., 90+% J&F) on benchmarks such as DAVIS and YouTube-VOS, these datasets…

计算机视觉与模式识别 · 计算机科学 2025-09-23 Henghui Ding , Kaining Ying , Chang Liu , Shuting He , Xudong Jiang , Yu-Gang Jiang , Philip H. S. Torr , Song Bai

Spatio-temporal action detection is an important and challenging problem in video understanding. The existing action detection benchmarks are limited in aspects of small numbers of instances in a trimmed video or low-level atomic actions.…

计算机视觉与模式识别 · 计算机科学 2021-08-19 Yixuan Li , Lei Chen , Runyu He , Zhenzhi Wang , Gangshan Wu , Limin Wang

The paper provides a survey of the development of machine-learning techniques for video analysis. The survey provides a summary of the most popular deep learning methods used for human activity recognition. We discuss how popular…

计算机视觉与模式识别 · 计算机科学 2023-12-12 Marios S. Pattichis , Venkatesh Jatla , Alvaro E. Ullao Cerna

We develop a technique for generating smooth and accurate 3D human pose and motion estimates from RGB video sequences. Our method, which we call Motion Estimation via Variational Autoencoder (MEVA), decomposes a temporal sequence of human…

计算机视觉与模式识别 · 计算机科学 2020-10-07 Zhengyi Luo , S. Alireza Golestaneh , Kris M. Kitani

As robots transition from controlled settings to unstructured human environments, building generalist agents that can reliably follow natural language instructions remains a central challenge. Progress in robust mobile manipulation requires…