中文
相关论文

相关论文: MIDV-500: A Dataset for Identity Documents Analysi…

200 篇论文

Datasets drive vision progress, yet existing driving datasets are impoverished in terms of visual content and supported tasks to study multitask learning for autonomous driving. Researchers are usually constrained to study a small set of…

计算机视觉与模式识别 · 计算机科学 2020-04-09 Fisher Yu , Haofeng Chen , Xin Wang , Wenqi Xian , Yingying Chen , Fangchen Liu , Vashisht Madhavan , Trevor Darrell

Recent years have witnessed a broader range of applications of image processing technologies in multiple industrial processes, such as smoke detection, security monitoring, and workpiece inspection. Different kinds of distortion types and…

计算机视觉与模式识别 · 计算机科学 2024-02-19 Xuanchao Ma , Yanlin Jiang , Hongyan Liu , Chengxu Zhou , Ke Gu

Objects falling from buildings, a frequently occurring event in daily life, can cause severe injuries to pedestrians due to the high impact force they exert. Surveillance cameras are often installed around buildings to detect falling…

计算机视觉与模式识别 · 计算机科学 2025-09-08 Zhigang Tu , Zhengbo Zhang , Zitao Gao , Chunluan Zhou , Junsong Yuan , Bo Du

Video aesthetic assessment, a vital area in multimedia computing, integrates computer vision with human cognition. Its progress is limited by the lack of standardized datasets and robust models, as the temporal dynamics of video and…

计算机视觉与模式识别 · 计算机科学 2025-11-14 Qianqian Qiao , DanDan Zheng , Yihang Bo , Bao Peng , Heng Huang , Longteng Jiang , Huaye Wang , Jingdong Chen , Jun Zhou , Xin Jin

Recently, many video enhancement methods have been proposed to improve video quality from different aspects such as color, brightness, contrast, and stability. Therefore, how to evaluate the quality of the enhanced video in a way consistent…

图像与视频处理 · 电气工程与系统科学 2023-03-17 Yixuan Gao , Yuqin Cao , Tengchuan Kou , Wei Sun , Yunlong Dong , Xiaohong Liu , Xiongkuo Min , Guangtao Zhai

Deepfakes represent a growing concern across domains such as disinformation, fraud, and non-consensual media. In particular, the rise of video conference and identity-driven attacks in high-stakes scenarios--such as impostor hiring--demands…

计算机视觉与模式识别 · 计算机科学 2026-04-08 Sarah Barrington , Maty Bohacek , Hany Farid

FUNSD is one of the limited publicly available datasets for information extraction from document im-ages. The information in the FUNSD dataset is defined by text areas of four categories ("key", "value", "header", "other", and "background")…

计算机视觉与模式识别 · 计算机科学 2020-10-13 Hieu M. Vu , Diep Thi-Ngoc Nguyen

Research in face recognition has seen tremendous growth over the past couple of decades. Beginning from algorithms capable of performing recognition in constrained environments, the current face recognition systems achieve very high…

计算机视觉与模式识别 · 计算机科学 2018-11-22 Maneet Singh , Richa Singh , Mayank Vatsa , Nalini Ratha , Rama Chellappa

Video crime detection is a significant application of computer vision and artificial intelligence. However, existing datasets primarily focus on detecting severe crimes by analyzing entire video clips, often neglecting the precursor…

计算机视觉与模式识别 · 计算机科学 2024-12-06 Ryozo Masukawa , Sanggeon Yun , Yoshiki Yamaguchi , Mohsen Imani

Smart focal-plane and in-chip image processing has emerged as a crucial technology for vision-enabled embedded systems with energy efficiency and privacy. However, the lack of special datasets providing examples of the data that these…

计算机视觉与模式识别 · 计算机科学 2024-10-02 Riadul Islam , Sri Ranga Sai Krishna Tummala , Joey Mulé , Rohith Kankipati , Suraj Jalapally , Dhandeep Challagundla , Chad Howard , Ryan Robucci

RGB-D data is essential for solving many problems in computer vision. Hundreds of public RGB-D datasets containing various scenes, such as indoor, outdoor, aerial, driving, and medical, have been proposed. These datasets are useful for…

计算机视觉与模式识别 · 计算机科学 2022-08-10 Alexandre Lopes , Roberto Souza , Helio Pedrini

We propose a new semi-supervised learning method on face-related tasks based on Multi-Task Learning (MTL) and data distillation. The proposed method exploits multiple datasets with different labels for different-but-related tasks such as…

计算机视觉与模式识别 · 计算机科学 2019-07-10 Sepidehsadat Hosseini , Mohammad Amin Shabani , Nam Ik Cho

Sign language recognition is a challenging and often underestimated problem comprising multi-modal articulators (handshape, orientation, movement, upper body and face) that integrate asynchronously on multiple streams. Learning powerful…

计算机视觉与模式识别 · 计算机科学 2019-11-22 Hamid Reza Vaezi Joze , Oscar Koller

Retinal vessel segmentation is generally grounded in image-based datasets collected with bench-top devices. The static images naturally lose the dynamic characteristics of retina fluctuation, resulting in diminished dataset richness, and…

As tools for content editing mature, and artificial intelligence (AI) based algorithms for synthesizing media grow, the presence of manipulated content across online media is increasing. This phenomenon causes the spread of misinformation,…

计算机视觉与模式识别 · 计算机科学 2022-12-09 Trisha Mittal , Ritwik Sinha , Viswanathan Swaminathan , John Collomosse , Dinesh Manocha

While the significant advancements have made in the generation of deepfakes using deep learning technologies, its misuse is a well-known issue now. Deepfakes can cause severe security and privacy issues as they can be used to impersonate a…

计算机视觉与模式识别 · 计算机科学 2022-03-02 Hasam Khalid , Shahroz Tariq , Minha Kim , Simon S. Woo

We present MVIP, a novel dataset for multi-modal and multi-view application-oriented industrial part recognition. Here we are the first to combine a calibrated RGBD multi-view dataset with additional object context such as physical…

计算机视觉与模式识别 · 计算机科学 2025-02-24 Paul Koch , Marian Schlüter , Jörg Krüger

Document Visual Question Answering (VQA) aims to understand visually-rich documents to answer questions in natural language, which is an emerging research topic for both Natural Language Processing and Computer Vision. In this work, we…

计算机视觉与模式识别 · 计算机科学 2023-05-05 Fengbin Zhu , Wenqiang Lei , Fuli Feng , Chao Wang , Haozhou Zhang , Tat-Seng Chua

Most existing datasets for sound event recognition (SER) are relatively small and/or domain-specific, with the exception of AudioSet, based on over 2M tracks from YouTube videos and encompassing over 500 sound classes. However, AudioSet is…

声音 · 计算机科学 2022-04-26 Eduardo Fonseca , Xavier Favory , Jordi Pons , Frederic Font , Xavier Serra

In this paper, a novel dataset is introduced, designed to assess student attention within in-person classroom settings. This dataset encompasses RGB camera data, featuring multiple cameras per student to capture both posture and facial…

‹ 上一页 1 8 9 10 下一页 ›