中文
相关论文

相关论文: PANDA: A Gigapixel-level Human-centric Video Datas…

200 篇论文

Humans excel at constructing panoramic mental models of their surroundings, maintaining object permanence and inferring scene structure beyond visible regions. In contrast, current artificial vision systems struggle with persistent,…

计算机视觉与模式识别 · 计算机科学 2025-12-16 Finlay G. C. Hudson , James A. D. Gardner , William A. P. Smith

Map representations learned by expert demonstrations have shown promising research value. However, the field of visual navigation still faces challenges due to the lack of real-world human-navigation datasets that can support efficient,…

计算机视觉与模式识别 · 计算机科学 2025-02-25 Faith Johnson , Bryan Bo Cao , Kristin Dana , Shubham Jain , Ashwin Ashok

While feed-forward 3D reconstruction models have advanced rapidly, they still exhibit degraded performance on panoramas due to spherical distortions. Moreover, existing panoramic 3D datasets are predominantly collected with 360 cameras…

计算机视觉与模式识别 · 计算机科学 2026-04-27 Jing Ou , Zidong Cao , Yinrui Ren , Zhuoxiao Li , Jinjing Zhu , Tongyan Hua , Shuai Zhang , Hui Xiong , Wufan Zhao

People detection methods are highly sensitive to the perpetual occlusions among the targets. As multi-camera set-ups become more frequently encountered, joint exploitation of the across views information would allow for improved detection…

计算机视觉与模式识别 · 计算机科学 2017-07-31 Tatjana Chavdarova , Pierre Baqué , Stéphane Bouquet , Andrii Maksai , Cijo Jose , Louis Lettry , Pascal Fua , Luc Van Gool , François Fleuret

Along with the development of modern smart cities, human-centric video analysis has been encountering the challenge of analyzing diverse and complex events in real scenes. A complex event relates to dense crowds, anomalous individuals, or…

计算机视觉与模式识别 · 计算机科学 2023-07-14 Weiyao Lin , Huabin Liu , Shizhan Liu , Yuxi Li , Rui Qian , Tao Wang , Ning Xu , Hongkai Xiong , Guo-Jun Qi , Nicu Sebe

Autonomous trucking is a promising technology that can greatly impact modern logistics and the environment. Ensuring its safety on public roads is one of the main duties that requires an accurate perception of the environment. To achieve…

With the advent of portable 360{\deg} cameras, panorama has gained significant attention in applications like virtual reality (VR), virtual tours, robotics, and autonomous driving. As a result, wide-baseline panorama view synthesis has…

计算机视觉与模式识别 · 计算机科学 2025-03-25 Cheng Zhang , Haofei Xu , Qianyi Wu , Camilo Cruz Gambardella , Dinh Phung , Jianfei Cai

In video surveillance, pedestrian retrieval (also called person re-identification) is a critical task. This task aims to retrieve the pedestrian of interest from non-overlapping cameras. Recently, transformer-based models have achieved…

计算机视觉与模式识别 · 计算机科学 2022-04-07 Xianghao Zang , Ge Li , Wei Gao

The human hand is our primary interface to the physical world, yet egocentric perception rarely knows when, where, or how forcefully it makes contact. Robust wearable tactile sensors are scarce, and no existing in-the-wild datasets align…

Recently, NVS in human-object interaction scenes has received increasing attention. Existing human-object interaction datasets mainly consist of static data with limited views, offering only RGB images or videos, mostly containing…

计算机视觉与模式识别 · 计算机科学 2024-09-23 Shuai Guo , Houqiang Zhong , Qiuwen Wang , Ziyu Chen , Yijie Gao , Jiajing Yuan , Chenyu Zhang , Rong Xie , Li Song

We introduce PointOdyssey, a large-scale synthetic dataset, and data generation framework, for the training and evaluation of long-term fine-grained tracking algorithms. Our goal is to advance the state-of-the-art by placing emphasis on…

计算机视觉与模式识别 · 计算机科学 2023-07-28 Yang Zheng , Adam W. Harley , Bokui Shen , Gordon Wetzstein , Leonidas J. Guibas

In this work, we propose a novel Spatial-Temporal Attention (STA) approach to tackle the large-scale person re-identification task in videos. Different from the most existing methods, which simply compute representations of video clips…

计算机视觉与模式识别 · 计算机科学 2020-05-01 Yang Fu , Xiaoyang Wang , Yunchao Wei , Thomas Huang

Hands are the central means by which humans manipulate their world and being able to reliably extract hand state information from Internet videos of humans engaged in their hands has the potential to pave the way to systems that can learn…

计算机视觉与模式识别 · 计算机科学 2020-06-12 Dandan Shan , Jiaqi Geng , Michelle Shu , David F. Fouhey

Using drones to track multiple individuals simultaneously in their natural environment is a powerful approach for better understanding group primate behavior. Previous studies have demonstrated that it is possible to automate the…

We present a novel approach for generating 360-degree high-quality, spatio-temporally coherent human videos from a single image. Our framework combines the strengths of diffusion transformers for capturing global correlations across…

计算机视觉与模式识别 · 计算机科学 2024-09-25 Ruizhi Shao , Youxin Pang , Zerong Zheng , Jingxiang Sun , Yebin Liu

The ability to grasp objects, signal with gestures, and share emotion through touch all stem from the unique capabilities of human hands. Yet creating high-quality personalized hand avatars from images remains challenging due to complex…

计算机视觉与模式识别 · 计算机科学 2026-02-10 Zicong Fan , Edoardo Remelli , David Dimond , Fadime Sener , Liuhao Ge , Bugra Tekin , Cem Keskin , Shreyas Hampali

Synthesizing 3D human motion in a contextual, ecological environment is important for simulating realistic activities people perform in the real world. However, conventional optics-based motion capture systems are not suited for…

计算机视觉与模式识别 · 计算机科学 2023-04-03 Joao Pedro Araujo , Jiaman Li , Karthik Vetrivel , Rishi Agarwal , Deepak Gopinath , Jiajun Wu , Alexander Clegg , C. Karen Liu

Human pose estimation (HPE) with convolutional neural networks (CNNs) for indoor monitoring is one of the major challenges in computer vision. In contrast to HPE in perspective views, an indoor monitoring system can consist of an…

计算机视觉与模式识别 · 计算机科学 2023-04-18 Jingrui Yu , Tobias Scheck , Roman Seidel , Yukti Adya , Dipankar Nandi , Gangolf Hirtz

Medical ultrasound video analysis is challenging due to variable sequence lengths, subtle spatial cues, and the need for interpretable video-level assessment. We introduce GADA, a Graph Attention-based Detection Aggregation framework that…

图像与视频处理 · 电气工程与系统科学 2025-10-14 Li Chen , Naveen Balaraju , Jochen Kruecker , Balasundar Raju , Alvin Chen

The rapid advancement of deep learning has intensified the need for comprehensive data for use by autonomous driving algorithms. High-quality datasets are crucial for the development of effective data-driven autonomous driving solutions.…

计算机视觉与模式识别 · 计算机科学 2026-02-10 Lianqing Zheng , Long Yang , Qunshu Lin , Wenjin Ai , Minghao Liu , Shouyi Lu , Jianan Liu , Hongze Ren , Jingyue Mo , Xiaokai Bai , Jie Bai , Zhixiong Ma , Xichan Zhu