中文
相关论文

相关论文: PANDA: A Gigapixel-level Human-centric Video Datas…

200 篇论文

Understanding affective dynamics in real-world social systems is fundamental to modeling and analyzing human-human interactions in complex environments. Group affect emerges from intertwined human-human interactions, contextual influences,…

计算机视觉与模式识别 · 计算机科学 2026-04-20 Deepak Kumar , Abhishek Pratap Singh , Puneet Kumar , Xiaobai Li , Balasubramanian Raman

Human-centric 3D scene understanding has recently drawn increasing attention, driven by its critical impact on robotics. However, human-centric real-life scenarios are extremely diverse and complicated, and humans have intricate motions and…

计算机视觉与模式识别 · 计算机科学 2024-03-18 Yichen Yao , Zimo Jiang , Yujing Sun , Zhencai Zhu , Xinge Zhu , Runnan Chen , Yuexin Ma

This paper presents LAPA (Look Around and Pay Attention), a novel end-to-end transformer-based architecture for multi-camera point tracking that integrates appearance-based matching with geometric constraints. Traditional pipelines decouple…

计算机视觉与模式识别 · 计算机科学 2025-12-05 Bishoy Galoaa , Xiangyu Bai , Shayda Moezzi , Utsav Nandi , Sai Siddhartha Vivek Dhir Rangoju , Somaieh Amraee , Sarah Ostadabbas

Tracking dense 3D motion from monocular videos remains challenging, particularly when aiming for pixel-level precision over long sequences. We introduce DELTA, a novel method that efficiently tracks every pixel in 3D space, enabling…

计算机视觉与模式识别 · 计算机科学 2025-03-03 Tuan Duc Ngo , Peiye Zhuang , Chuang Gan , Evangelos Kalogerakis , Sergey Tulyakov , Hsin-Ying Lee , Chaoyang Wang

A comprehensive understanding of interested human-to-human interactions in video streams, such as queuing, handshaking, fighting and chasing, is of immense importance to the surveillance of public security in regions like campuses, squares…

计算机视觉与模式识别 · 计算机科学 2023-08-14 Zhenhua Wang , Kaining Ying , Jiajun Meng , Jifeng Ning

Recent years have seen remarkable advances in visual understanding. However, how to understand a story-based long video with artistic styles, e.g. movie, remains challenging. In this paper, we introduce MovieNet -- a holistic dataset for…

计算机视觉与模式识别 · 计算机科学 2020-07-22 Qingqiu Huang , Yu Xiong , Anyi Rao , Jiaze Wang , Dahua Lin

Foundation models, exemplified by GPT technology, are discovering new horizons in artificial intelligence by executing tasks beyond their designers' expectations. While the present generation provides fundamental advances in understanding…

计算机视觉与模式识别 · 计算机科学 2023-12-19 Maksim Makarenko , Qizhou Wang , Arturo Burguete-Lopez , Silvio Giancola , Bernard Ghanem , Luca Passone , Andrea Fratalocchi

Recently, pedestrian behavior research has shifted towards machine learning based methods and converged on the topic of modeling pedestrian interactions. For this, a large-scale dataset that contains rich information is needed. We propose a…

计算机视觉与模式识别 · 计算机科学 2023-10-02 Allan Wang , Abhijat Biswas , Henny Admoni , Aaron Steinfeld

Crowd analysis via computer vision techniques is an important topic in the field of video surveillance, which has wide-spread applications including crowd monitoring, public safety, space design and so on. Pixel-wise crowd understanding is…

计算机视觉与模式识别 · 计算机科学 2020-08-04 Qi Wang , Junyu Gao , Wei Lin , Yuan Yuan

For people affected by blindness and low vision (BLV), safe and independent navigation remains a major challenge, impacting over 2.2 billion individuals worldwide. Although multimodal large language models (MLLMs) offer new opportunities…

计算机视觉与模式识别 · 计算机科学 2026-05-01 Junhyeok Kim , Jaewoo Park , Junhee Park , Sangeyl Lee , Jiwan Chung , Jisung Kim , Ji Hoon Joung , Youngjae Yu

We present the PanAf20K dataset, the largest and most diverse open-access annotated video dataset of great apes in their natural environment. It comprises more than 7 million frames across ~20,000 camera trap videos of chimpanzees and…

Visual understanding of complex urban street scenes is an enabling factor for a wide range of applications. Object detection has benefited enormously from large-scale datasets, especially in the context of deep learning. For semantic urban…

计算机视觉与模式识别 · 计算机科学 2016-04-08 Marius Cordts , Mohamed Omran , Sebastian Ramos , Timo Rehfeld , Markus Enzweiler , Rodrigo Benenson , Uwe Franke , Stefan Roth , Bernt Schiele

3D human pose estimation captures the human joint points in three-dimensional space while keeping the depth information and physical structure. That is essential for applications that require precise pose information, such as human-computer…

计算机视觉与模式识别 · 计算机科学 2024-03-26 Jianbin Jiao , Xina Cheng , Weijie Chen , Xiaoting Yin , Hao Shi , Kailun Yang

Despite significant progress in the development of human action detection datasets and algorithms, no current dataset is representative of real-world aerial view scenarios. We present Okutama-Action, a new video dataset for aerial view…

计算机视觉与模式识别 · 计算机科学 2017-06-16 Mohammadamin Barekatain , Miquel Martí , Hsueh-Fu Shih , Samuel Murray , Kotaro Nakayama , Yutaka Matsuo , Helmut Prendinger

We present a novel large-scale dataset and comprehensive baselines for end-to-end pedestrian detection and person recognition in raw video frames. Our baselines address three issues: the performance of various combinations of detectors and…

计算机视觉与模式识别 · 计算机科学 2017-04-07 Liang Zheng , Hengheng Zhang , Shaoyan Sun , Manmohan Chandraker , Yi Yang , Qi Tian

Understanding animals' behaviors is significant for a wide range of applications. However, existing animal behavior datasets have limitations in multiple aspects, including limited numbers of animal classes, data samples and provided tasks,…

计算机视觉与模式识别 · 计算机科学 2022-06-06 Xun Long Ng , Kian Eng Ong , Qichen Zheng , Yun Ni , Si Yong Yeo , Jun Liu

Panoramic videos contain richer spatial information and have attracted tremendous amounts of attention due to their exceptional experience in some fields such as autonomous driving and virtual reality. However, existing datasets for video…

计算机视觉与模式识别 · 计算机科学 2024-07-30 Shilin Yan , Xiaohao Xu , Renrui Zhang , Lingyi Hong , Wenchao Chen , Wenqiang Zhang , Wei Zhang

Event cameras encode visual information with high temporal precision, low data-rate, and high-dynamic range. Thanks to these characteristics, event cameras are particularly suited for scenarios with high motion, challenging lighting…

计算机视觉与模式识别 · 计算机科学 2020-12-10 Etienne Perot , Pierre de Tournemire , Davide Nitti , Jonathan Masci , Amos Sironi

The field of autonomous driving increasingly demands high-quality annotated training data. In this paper, we propose Panacea, an innovative approach to generate panoramic and controllable videos in driving scenarios, capable of yielding an…

计算机视觉与模式识别 · 计算机科学 2023-11-29 Yuqing Wen , Yucheng Zhao , Yingfei Liu , Fan Jia , Yanhui Wang , Chong Luo , Chi Zhang , Tiancai Wang , Xiaoyan Sun , Xiangyu Zhang

Recognizing soft-biometric pedestrian attributes is essential in video surveillance and fashion retrieval. Recent works show promising results on single datasets. Nevertheless, the generalization ability of these methods under different…

计算机视觉与模式识别 · 计算机科学 2022-09-07 Andreas Specker , Mickael Cormier , Jürgen Beyerer