中文
相关论文

相关论文: MIDV-2020: A Comprehensive Benchmark Dataset for I…

200 篇论文

Current movie captioning architectures are not capable of mentioning characters with their proper name, replacing them with a generic "someone" tag. The lack of movie description datasets with characters' visual annotations surely plays a…

计算机视觉与模式识别 · 计算机科学 2019-03-06 Stefano Pini , Marcella Cornia , Federico Bolelli , Lorenzo Baraldi , Rita Cucchiara

As AI-driven document understanding and processing tools become increasingly prevalent in real-world applications, the need for rigorous evaluation standards has grown increasingly urgent. Existing benchmarks and evaluations often focus on…

Given the central role of charts in scientific, business, and communication contexts, enhancing the chart understanding capabilities of vision-language models (VLMs) has become increasingly critical. A key limitation of existing VLMs lies…

计算机视觉与模式识别 · 计算机科学 2025-10-17 Jiangning Zhu , Yuxing Zhou , Zheng Wang , Juntao Yao , Yima Gu , Yuhui Yuan , Shixia Liu

Descriptive video service (DVS) provides linguistic descriptions of movies and allows visually impaired people to follow a movie along with their peers. Such descriptions are by design mainly visual and thus naturally form an interesting…

计算机视觉与模式识别 · 计算机科学 2015-01-13 Anna Rohrbach , Marcus Rohrbach , Niket Tandon , Bernt Schiele

Video identity customization seeks to synthesize realistic, temporally coherent videos of a specific subject, given a single reference image and a text prompt. This task presents two core challenges: (1) maintaining identity consistency…

计算机视觉与模式识别 · 计算机科学 2026-02-04 Guiyu Zhang , Chen Shi , Zijian Jiang , Xunzhi Xiang , Jingjing Qian , Shaoshuai Shi , Li Jiang

Recent advancements in deep learning have significantly enhanced content-based retrieval methods, notably through models like CLIP that map images and texts into a shared embedding space. However, these methods often struggle with…

计算机视觉与模式识别 · 计算机科学 2025-02-11 Nicola Messina , Lucia Vadicamo , Leo Maltese , Claudio Gennaro

In this data article, we introduce the Multi-Modal Event-based Vehicle Detection and Tracking (MEVDT) dataset. This dataset provides a synchronized stream of event data and grayscale images of traffic scenes, captured using the Dynamic and…

计算机视觉与模式识别 · 计算机科学 2024-07-31 Zaid A. El Shair , Samir A. Rawashdeh

The emergence and popularity of facial deepfake methods spur the vigorous development of deepfake datasets and facial forgery detection, which to some extent alleviates the security concerns about facial-related artificial intelligence…

计算机视觉与模式识别 · 计算机科学 2025-09-22 Wenkui Yang , Zhida Zhang , Xiaoqiang Zhou , Junxian Duan , Jie Cao

Deepfakes, leveraging advanced AIGC (Artificial Intelligence-Generated Content) techniques, create hyper-realistic synthetic images and videos of human faces, posing a significant threat to the authenticity of social media. While this…

计算机视觉与模式识别 · 计算机科学 2025-10-10 Junyu Shi , Minghui Li , Junguo Zuo , Zhifei Yu , Yipeng Lin , Shengshan Hu , Ziqi Zhou , Yechao Zhang , Wei Wan , Yinzhe Xu , Leo Yu Zhang

While the significant advancements have made in the generation of deepfakes using deep learning technologies, its misuse is a well-known issue now. Deepfakes can cause severe security and privacy issues as they can be used to impersonate a…

计算机视觉与模式识别 · 计算机科学 2022-03-02 Hasam Khalid , Shahroz Tariq , Minha Kim , Simon S. Woo

AI systems rely on extensive training on large datasets to address various tasks. However, image-based systems, particularly those used for demographic attribute prediction, face significant challenges. Many current face image datasets…

计算机视觉与模式识别 · 计算机科学 2024-04-30 Georgia Baltsou , Ioannis Sarridis , Christos Koutlis , Symeon Papadopoulos

3D multi-person motion prediction is a challenging task that involves modeling individual behaviors and interactions between people. Despite the emergence of approaches for this task, comparing them is difficult due to the lack of…

计算机视觉与模式识别 · 计算机科学 2023-06-27 Xiaogang Peng , Xiao Zhou , Yikai Luo , Hao Wen , Yu Ding , Zizhao Wu

Ensuring fairness and robustness in machine learning models remains a challenge, particularly under domain shifts. We present Face4FairShifts, a large-scale facial image benchmark designed to systematically evaluate fairness-aware learning…

计算机视觉与模式识别 · 计算机科学 2025-09-03 Yumeng Lin , Dong Li , Xintao Wu , Minglai Shao , Xujiang Zhao , Zhong Chen , Chen Zhao

Spatio-temporal action detection is an important and challenging problem in video understanding. The existing action detection benchmarks are limited in aspects of small numbers of instances in a trimmed video or low-level atomic actions.…

计算机视觉与模式识别 · 计算机科学 2021-08-19 Yixuan Li , Lei Chen , Runyu He , Zhenzhi Wang , Gangshan Wu , Limin Wang

The rapid surge of text-to-speech and face-voice reenactment models makes video fabrication easier and highly realistic. To encounter this problem, we require datasets that rich in type of generation methods and perturbation strategy which…

计算机视觉与模式识别 · 计算机科学 2025-07-29 Zhixi Cai , Kartik Kuckreja , Shreya Ghosh , Akanksha Chuchra , Muhammad Haris Khan , Usman Tariq , Tom Gedeon , Abhinav Dhall

Accurate document layout analysis is a key requirement for high-quality PDF document conversion. With the recent availability of public, large ground-truth datasets such as PubLayNet and DocBank, deep-learning models have proven to be very…

计算机视觉与模式识别 · 计算机科学 2022-08-18 Birgit Pfitzmann , Christoph Auer , Michele Dolfi , Ahmed S Nassar , Peter W J Staar

Recent face recognition experiments on a major benchmark LFW show stunning performance--a number of algorithms achieve near to perfect score, surpassing human recognition rates. In this paper, we advocate evaluations at the million scale…

计算机视觉与模式识别 · 计算机科学 2015-12-03 Ira Kemelmacher-Shlizerman , Steve Seitz , Daniel Miller , Evan Brossard

Deep learning-based face recognition continues to face challenges due to its reliance on huge datasets obtained from web crawling, which can be costly to gather and raise significant real-world privacy concerns. To address this issue, we…

计算机视觉与模式识别 · 计算机科学 2025-03-14 Minsoo Kim , Min-Cheol Sagong , Gi Pyo Nam , Junghyun Cho , Ig-Jae Kim

Face anti-spoofing is essential to prevent face recognition systems from a security breach. Much of the progresses have been made by the availability of face anti-spoofing benchmark datasets in recent years. However, existing face…

计算机视觉与模式识别 · 计算机科学 2020-02-05 Shifeng Zhang , Ajian Liu , Jun Wan , Yanyan Liang , Guogong Guo , Sergio Escalera , Hugo Jair Escalante , Stan Z. Li

Ensuring the security of transactions is currently one of the major challenges that banking systems deal with. The usage of face for biometric authentication of users is attracting large investments from banks worldwide due to its…

计算机视觉与模式识别 · 计算机科学 2020-04-13 Johnatan S. Oliveira , Gustavo B. Souza , Anderson R. Rocha , Flávio E. Deus , Aparecido N. Marana
‹ 上一页 1 8 9 10 下一页 ›