English
Related papers

Related papers: iQIYI-VID: A Large Dataset for Multi-modal Person …

200 papers

In the domain of video surveillance, describing the behavior of each individual within the video is becoming increasingly essential, especially in complex scenarios with multiple individuals present. This is because describing each…

Computer Vision and Pattern Recognition · Computer Science 2023-10-05 Lingru Zhou , Yiqi Gao , Manqing Zhang , Peng Wu , Peng Wang , Yanning Zhang

Face verification is a significant component of identity authentication in various applications including online banking and secure access to personal devices. The majority of the existing face image datasets often suffer from notable…

Computer Vision and Pattern Recognition · Computer Science 2025-11-24 Georgia Baltsou , Ioannis Sarridis , Christos Koutlis , Symeon Papadopoulos

The Real Face Dataset is a pedestrian face detection benchmark dataset in the wild, comprising over 11,000 images and over 55,000 detected faces in various ambient conditions. The dataset aims to provide a comprehensive and diverse…

Computer Vision and Pattern Recognition · Computer Science 2024-09-04 Leonardo Ramos Thomas

In real-world applications, e.g. law enforcement and video retrieval, one often needs to search a certain person in long videos with just one portrait. This is much more challenging than the conventional settings for person…

Computer Vision and Pattern Recognition · Computer Science 2018-07-30 Qingqiu Huang , Wentao Liu , Dahua Lin

The Codec Avatars Lab at Meta introduces Embody 3D, a multimodal dataset of 500 individual hours of 3D motion data from 439 participants collected in a multi-camera collection stage, amounting to over 54 million frames of tracked 3D motion.…

A well-designed interactive human-like dialogue system is expected to take actions (e.g. smiling) and respond in a pattern similar to humans. However, due to the limitation of single-modality (only speech) or small volume of currently…

Human-Computer Interaction · Computer Science 2022-12-13 Zhiling Luo , Qiankun Shi , Sha Zhao , Wei Zhou , Haiqing Chen , Yuankai Ma , Haitao Leng

Cross-modal person re-identification (Re-ID) is critical for modern video surveillance systems. The key challenge is to align cross-modality representations induced by the semantic information present for a person and ignore background…

Computer Vision and Pattern Recognition · Computer Science 2022-07-22 Ammarah Farooq , Muhammad Awais , Josef Kittler , Syed Safwan Khalid

The ability to identify the same person from multiple camera views without the explicit use of facial recognition is receiving commercial and academic interest. The current status-quo solutions are based on attention neural models. In this…

Computer Vision and Pattern Recognition · Computer Science 2019-12-18 Priyank Pathak , Amir Erfan Eshratifar , Michael Gormish

Cutting-edge robot learning techniques including foundation models and imitation learning from humans all pose huge demands on large-scale and high-quality datasets which constitute one of the bottleneck in the general intelligent robot…

Robotics · Computer Science 2026-04-27 Shuo Jiang , Haonan Li , Ruochen Ren , Yanmin Zhou , Zhipeng Wang , Bin He

Video-based person re-identification (re-id) is a central application in surveillance systems with significant concern in security. Matching persons across disjoint camera views in their video fragments is inherently challenging due to the…

Computer Vision and Pattern Recognition · Computer Science 2018-10-17 Lin Wu , Yang Wang , Junbin Gao , Xue Li

Recent work has established the ecological importance of developing algorithms for identifying animals individually from images. Typically, a separate algorithm is trained for each species, a natural step but one that creates significant…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Lasha Otarashvili , Tamilselvan Subramanian , Jason Holmberg , J. J. Levenson , Charles V. Stewart

Real-world surveillance often renders faces and license plates unrecognizable in individual low-resolution (LR) frames, hindering reliable identification. To advance temporal recognition models, we present FANVID, a novel video-based…

Computer Vision and Pattern Recognition · Computer Science 2026-02-12 Kavitha Viswanathan , Vrinda Goel , Shlesh Gholap , Devayan Ghosh , Madhav Gupta , Dhruvi Ganatra , Sanket Potdar , Amit Sethi

Multimodal person re-identification (Re-ID) aims to match pedestrian images across different modalities. However, most existing methods focus on limited cross-modal settings and fail to support arbitrary query-retrieval combinations,…

Computer Vision and Pattern Recognition · Computer Science 2025-10-20 Zhen Sun , Lei Tan , Yunhang Shen , Chengmao Cai , Xing Sun , Pingyang Dai , Liujuan Cao , Rongrong Ji

The volumetric representation of human interactions is one of the fundamental domains in the development of immersive media productions and telecommunication applications. Particularly in the context of the rapid advancement of Extended…

Computer Vision and Pattern Recognition · Computer Science 2024-02-15 Fatemeh Ghorbani Lohesara , Davi Rabbouni Freitas , Christine Guillemot , Karen Eguiazarian , Sebastian Knorr

Most existing video tasks related to "human" focus on the segmentation of salient humans, ignoring the unspecified others in the video. Few studies have focused on segmenting and tracking all humans in a complex video, including pedestrians…

Computer Vision and Pattern Recognition · Computer Science 2021-08-17 Ran Yu , Chenyu Tian , Weihao Xia , Xinyuan Zhao , Haoqian Wang , Yujiu Yang

Video-based person re-identification (ReID) has become increasingly important due to its applications in video surveillance applications. By employing events in video-based person ReID, more motion information can be provided between…

Computer Vision and Pattern Recognition · Computer Science 2025-01-14 Renkai Li , Xin Yuan , Wei Liu , Xin Xu

Driver distraction has become a significant cause of severe traffic accidents over the past decade. Despite the growing development of vision-driven driver monitoring systems, the lack of comprehensive perception datasets restricts road…

Computer Vision and Pattern Recognition · Computer Science 2023-08-02 Dingkang Yang , Shuai Huang , Zhi Xu , Zhenpeng Li , Shunli Wang , Mingcheng Li , Yuzheng Wang , Yang Liu , Kun Yang , Zhaoyu Chen , Yan Wang , Jing Liu , Peixuan Zhang , Peng Zhai , Lihua Zhang

When humans converse, what a speaker will say next significantly depends on what he sees. Unfortunately, existing dialogue models generate dialogue utterances only based on preceding textual contexts, and visual contexts are rarely…

Computation and Language · Computer Science 2021-06-01 Yuxian Meng , Shuhe Wang , Qinghong Han , Xiaofei Sun , Fei Wu , Rui Yan , Jiwei Li

Videos capture events that typically contain multiple sequential, and simultaneous, actions even in the span of only a few seconds. However, most large-scale datasets built to train models for action recognition in video only provide a…

Computer Vision and Pattern Recognition · Computer Science 2021-09-29 Mathew Monfort , Bowen Pan , Kandan Ramakrishnan , Alex Andonian , Barry A McNamara , Alex Lascelles , Quanfu Fan , Dan Gutfreund , Rogerio Feris , Aude Oliva

Multimodal counterfactual reasoning is a vital yet challenging ability for AI systems. It involves predicting the outcomes of hypothetical circumstances based on vision and language inputs, which enables AI models to learn from failures and…

Computer Vision and Pattern Recognition · Computer Science 2023-11-06 Te-Lin Wu , Zi-Yi Dou , Qingyuan Hu , Yu Hou , Nischal Reddy Chandra , Marjorie Freedman , Ralph M. Weischedel , Nanyun Peng