中文
相关论文

相关论文: GEB+: A Benchmark for Generic Event Boundary Capti…

200 篇论文

Understanding human behavior from complementary egocentric (ego) and exocentric (exo) points of view enables the development of systems that can support workers in industrial environments and enhance their safety. However, progress in this…

Contextual reasoning is essential to understand events in long untrimmed videos. In this work, we systematically explore different captioning models with various contexts for the dense-captioning events in video task, which aims to generate…

计算机视觉与模式识别 · 计算机科学 2019-07-12 Shizhe Chen , Yuqing Song , Yida Zhao , Qin Jin , Zhaoyang Zeng , Bei Liu , Jianlong Fu , Alexander Hauptmann

Pedestrian motion, due to its causal nature, is strongly influenced by domain gaps arising from discrepancies between training and testing data distributions. Focusing on 3D human pose estimation, this work presents a controllable human…

计算机视觉与模式识别 · 计算机科学 2026-05-13 Xinhao Hu , Yiyi Zhang , Liqing Zhang , Jianfu Zhang

Neural networks trained on datasets such as ImageNet have led to major advances in visual object classification. One obstacle that prevents networks from reasoning more deeply about complex scenes and situations, and from integrating visual…

Generating long form narratives such as stories and procedures from multiple modalities has been a long standing dream for artificial intelligence. In this regard, there is often crucial subtext that is derived from the surrounding…

计算与语言 · 计算机科学 2020-10-28 Khyathi Raghavi Chandu , Ruo-Ping Dong , Alan Black

The rapid advances in generative models have significantly lowered the barrier to producing convincing multimodal disinformation. Fabricated images and manipulated captions increasingly co-occur to create persuasive false narratives. While…

计算机视觉与模式识别 · 计算机科学 2025-10-01 Gagandeep Singh , Samudi Amarsinghe , Priyanka Singh , Xue Li

We introduce QuerYD, a new large-scale dataset for retrieval and event localisation in video. A unique feature of our dataset is the availability of two audio tracks for each video: the original audio, and a high-quality spoken description…

计算机视觉与模式识别 · 计算机科学 2021-02-18 Andreea-Maria Oncescu , João F. Henriques , Yang Liu , Andrew Zisserman , Samuel Albanie

Event camera, a novel neuromorphic vision sensor, records data with high temporal resolution and wide dynamic range, offering new possibilities for accurate visual representation in challenging scenarios. However, event data is inherently…

计算机视觉与模式识别 · 计算机科学 2025-08-08 Lin Zhu , Ruonan Liu , Xiao Wang , Lizhi Wang , Hua Huang

Evaluating video captioning systems is a challenging task as there are multiple factors to consider; for instance: the fluency of the caption, multiple actions happening in a single scene, and the human bias of what is considered important.…

计算机视觉与模式识别 · 计算机科学 2022-05-17 Luis Lebron , Yvette Graham , Kevin McGuinness , Konstantinos Kouramas , Noel E. O'Connor

Long-form video question answering remains challenging for modern vision-language models, which struggle to reason over hour-scale footage without exceeding practical token and compute budgets. Existing systems typically downsample frames…

计算机视觉与模式识别 · 计算机科学 2026-01-13 Aradhya Dixit , Tianxi Liang

Gait recognition enables non-intrusive, privacy-preserving identification but suffers in uncontrolled environments due to illumination and motion sensitivity of conventional cameras. In this work, we explore gait recognition using event…

计算机视觉与模式识别 · 计算机科学 2026-05-22 Senyan Xu , Shuai Chen , Chuanfu Shen , Kean Liu , Zhijing Sun , Chengzhi Cao , Xueyang Fu

We address the problem of text-guided video temporal grounding, which aims to identify the time interval of a certain event based on a natural language description. Different from most existing methods that only consider RGB images as…

计算机视觉与模式识别 · 计算机科学 2021-11-01 Yi-Wen Chen , Yi-Hsuan Tsai , Ming-Hsuan Yang

Modeling crowd behavior relies on accurate data of pedestrian movements at a high level of detail. Imaging sensors such as cameras provide a good basis for capturing such detailed pedestrian motion data. However, currently available…

计算机视觉与模式识别 · 计算机科学 2012-10-11 Stefan Seer , Norbert Brändle , Carlo Ratti

The analysis of events in dynamic environments poses a fundamental challenge in the development of intelligent agents and robots capable of interacting with humans. Current approaches predominantly utilize visual models. However, these…

计算机视觉与模式识别 · 计算机科学 2025-05-07 Sergey Linok , Vadim Semenov , Anastasia Trunova , Oleg Bulichev , Dmitry Yudin

Since the launch of the Microsoft Kinect, scores of RGBD datasets have been released. These have propelled advances in areas from reconstruction to gesture recognition. In this paper we explore the field, reviewing datasets across eight…

计算机视觉与模式识别 · 计算机科学 2016-04-14 Michael Firman

We present Dive Into the BoundarieS (DIBS), a novel pretraining framework for dense video captioning (DVC), that elaborates on improving the quality of the generated event captions and their associated pseudo event boundaries from unlabeled…

计算机视觉与模式识别 · 计算机科学 2024-04-04 Hao Wu , Huabin Liu , Yu Qiao , Xiao Sun

The convergence of 3D geometric perception and video synthesis has created an unprecedented demand for large-scale video data that is rich in both semantic and spatio-temporal information. While existing datasets have advanced either 3D…

计算机视觉与模式识别 · 计算机科学 2026-04-28 Yunnan Wang , Kecheng Zheng , Jianyuan Wang , Minghao Chen , David Novotny , Christian Rupprecht , Yinghao Xu , Xing Zhu , Wenjun Zeng , Xin Jin , Yujun Shen

Dense video captioning is a fine-grained video understanding task that involves two sub-problems: localizing distinct events in a long video stream, and generating captions for the localized events. We propose the Joint Event Detection and…

计算机视觉与模式识别 · 计算机科学 2018-12-27 Huijuan Xu , Boyang Li , Vasili Ramanishka , Leonid Sigal , Kate Saenko

Large-scale video generation models have demonstrated emergent physical coherence, positioning them as potential world models. However, a gap remains between contemporary "stateless" video architectures and classic state-centric world model…

计算机视觉与模式识别 · 计算机科学 2026-01-27 Luozhou Wang , Zhifei Chen , Yihua Du , Dongyu Yan , Wenhang Ge , Guibao Shen , Xinli Xu , Leyi Wu , Man Chen , Tianshuo Xu , Peiran Ren , Xin Tao , Pengfei Wan , Ying-Cong Chen

Artificial Intelligence makes great advances today and starts to bridge the gap between vision and language. However, we are still far from understanding, explaining and controlling explicitly the visual content from a linguistic…

人工智能 · 计算机科学 2023-09-19 Mihai Masala , Nicolae Cudlenco , Traian Rebedea , Marius Leordeanu
‹ 上一页 1 8 9 10 下一页 ›