English
Related papers

Related papers: MIDV-2020: A Comprehensive Benchmark Dataset for I…

200 papers

Current movie captioning architectures are not capable of mentioning characters with their proper name, replacing them with a generic "someone" tag. The lack of movie description datasets with characters' visual annotations surely plays a…

Computer Vision and Pattern Recognition · Computer Science 2019-03-06 Stefano Pini , Marcella Cornia , Federico Bolelli , Lorenzo Baraldi , Rita Cucchiara

As AI-driven document understanding and processing tools become increasingly prevalent in real-world applications, the need for rigorous evaluation standards has grown increasingly urgent. Existing benchmarks and evaluations often focus on…

Given the central role of charts in scientific, business, and communication contexts, enhancing the chart understanding capabilities of vision-language models (VLMs) has become increasingly critical. A key limitation of existing VLMs lies…

Computer Vision and Pattern Recognition · Computer Science 2025-10-17 Jiangning Zhu , Yuxing Zhou , Zheng Wang , Juntao Yao , Yima Gu , Yuhui Yuan , Shixia Liu

Descriptive video service (DVS) provides linguistic descriptions of movies and allows visually impaired people to follow a movie along with their peers. Such descriptions are by design mainly visual and thus naturally form an interesting…

Computer Vision and Pattern Recognition · Computer Science 2015-01-13 Anna Rohrbach , Marcus Rohrbach , Niket Tandon , Bernt Schiele

Video identity customization seeks to synthesize realistic, temporally coherent videos of a specific subject, given a single reference image and a text prompt. This task presents two core challenges: (1) maintaining identity consistency…

Computer Vision and Pattern Recognition · Computer Science 2026-02-04 Guiyu Zhang , Chen Shi , Zijian Jiang , Xunzhi Xiang , Jingjing Qian , Shaoshuai Shi , Li Jiang

Recent advancements in deep learning have significantly enhanced content-based retrieval methods, notably through models like CLIP that map images and texts into a shared embedding space. However, these methods often struggle with…

Computer Vision and Pattern Recognition · Computer Science 2025-02-11 Nicola Messina , Lucia Vadicamo , Leo Maltese , Claudio Gennaro

In this data article, we introduce the Multi-Modal Event-based Vehicle Detection and Tracking (MEVDT) dataset. This dataset provides a synchronized stream of event data and grayscale images of traffic scenes, captured using the Dynamic and…

Computer Vision and Pattern Recognition · Computer Science 2024-07-31 Zaid A. El Shair , Samir A. Rawashdeh

The emergence and popularity of facial deepfake methods spur the vigorous development of deepfake datasets and facial forgery detection, which to some extent alleviates the security concerns about facial-related artificial intelligence…

Computer Vision and Pattern Recognition · Computer Science 2025-09-22 Wenkui Yang , Zhida Zhang , Xiaoqiang Zhou , Junxian Duan , Jie Cao

Deepfakes, leveraging advanced AIGC (Artificial Intelligence-Generated Content) techniques, create hyper-realistic synthetic images and videos of human faces, posing a significant threat to the authenticity of social media. While this…

Computer Vision and Pattern Recognition · Computer Science 2025-10-10 Junyu Shi , Minghui Li , Junguo Zuo , Zhifei Yu , Yipeng Lin , Shengshan Hu , Ziqi Zhou , Yechao Zhang , Wei Wan , Yinzhe Xu , Leo Yu Zhang

While the significant advancements have made in the generation of deepfakes using deep learning technologies, its misuse is a well-known issue now. Deepfakes can cause severe security and privacy issues as they can be used to impersonate a…

Computer Vision and Pattern Recognition · Computer Science 2022-03-02 Hasam Khalid , Shahroz Tariq , Minha Kim , Simon S. Woo

AI systems rely on extensive training on large datasets to address various tasks. However, image-based systems, particularly those used for demographic attribute prediction, face significant challenges. Many current face image datasets…

Computer Vision and Pattern Recognition · Computer Science 2024-04-30 Georgia Baltsou , Ioannis Sarridis , Christos Koutlis , Symeon Papadopoulos

3D multi-person motion prediction is a challenging task that involves modeling individual behaviors and interactions between people. Despite the emergence of approaches for this task, comparing them is difficult due to the lack of…

Computer Vision and Pattern Recognition · Computer Science 2023-06-27 Xiaogang Peng , Xiao Zhou , Yikai Luo , Hao Wen , Yu Ding , Zizhao Wu

Ensuring fairness and robustness in machine learning models remains a challenge, particularly under domain shifts. We present Face4FairShifts, a large-scale facial image benchmark designed to systematically evaluate fairness-aware learning…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Yumeng Lin , Dong Li , Xintao Wu , Minglai Shao , Xujiang Zhao , Zhong Chen , Chen Zhao

Spatio-temporal action detection is an important and challenging problem in video understanding. The existing action detection benchmarks are limited in aspects of small numbers of instances in a trimmed video or low-level atomic actions.…

Computer Vision and Pattern Recognition · Computer Science 2021-08-19 Yixuan Li , Lei Chen , Runyu He , Zhenzhi Wang , Gangshan Wu , Limin Wang

The rapid surge of text-to-speech and face-voice reenactment models makes video fabrication easier and highly realistic. To encounter this problem, we require datasets that rich in type of generation methods and perturbation strategy which…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Zhixi Cai , Kartik Kuckreja , Shreya Ghosh , Akanksha Chuchra , Muhammad Haris Khan , Usman Tariq , Tom Gedeon , Abhinav Dhall

Accurate document layout analysis is a key requirement for high-quality PDF document conversion. With the recent availability of public, large ground-truth datasets such as PubLayNet and DocBank, deep-learning models have proven to be very…

Computer Vision and Pattern Recognition · Computer Science 2022-08-18 Birgit Pfitzmann , Christoph Auer , Michele Dolfi , Ahmed S Nassar , Peter W J Staar

Recent face recognition experiments on a major benchmark LFW show stunning performance--a number of algorithms achieve near to perfect score, surpassing human recognition rates. In this paper, we advocate evaluations at the million scale…

Computer Vision and Pattern Recognition · Computer Science 2015-12-03 Ira Kemelmacher-Shlizerman , Steve Seitz , Daniel Miller , Evan Brossard

Deep learning-based face recognition continues to face challenges due to its reliance on huge datasets obtained from web crawling, which can be costly to gather and raise significant real-world privacy concerns. To address this issue, we…

Computer Vision and Pattern Recognition · Computer Science 2025-03-14 Minsoo Kim , Min-Cheol Sagong , Gi Pyo Nam , Junghyun Cho , Ig-Jae Kim

Face anti-spoofing is essential to prevent face recognition systems from a security breach. Much of the progresses have been made by the availability of face anti-spoofing benchmark datasets in recent years. However, existing face…

Computer Vision and Pattern Recognition · Computer Science 2020-02-05 Shifeng Zhang , Ajian Liu , Jun Wan , Yanyan Liang , Guogong Guo , Sergio Escalera , Hugo Jair Escalante , Stan Z. Li

Ensuring the security of transactions is currently one of the major challenges that banking systems deal with. The usage of face for biometric authentication of users is attracting large investments from banks worldwide due to its…

Computer Vision and Pattern Recognition · Computer Science 2020-04-13 Johnatan S. Oliveira , Gustavo B. Souza , Anderson R. Rocha , Flávio E. Deus , Aparecido N. Marana
‹ Prev 1 8 9 10 Next ›