中文
相关论文

相关论文: MFVLR: Multi-domain Fine-grained Vision-Language R…

200 篇论文

Three key challenges hinder the development of current deepfake video detection: (1) Temporal features can be complex and diverse: how can we identify general temporal artifacts to enhance model generalization? (2) Spatiotemporal models…

计算机视觉与模式识别 · 计算机科学 2024-12-03 Zhiyuan Yan , Yandan Zhao , Shen Chen , Mingyi Guo , Xinghe Fu , Taiping Yao , Shouhong Ding , Li Yuan

Domain Generalized person Re-identification (DG Re-ID) is a challenging task, where models are trained on source domains but tested on unseen target domains. Although previous pure vision-based models have achieved significant progress, the…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Jiachen Li , Xiaojin Gong , Dongping Zhang

Unveiling the real appearance of retouched faces to prevent malicious users from deceptive advertising and economic fraud has been an increasing concern in the era of digital economics. This article makes the first attempt to investigate…

计算机视觉与模式识别 · 计算机科学 2024-05-14 Fengchuang Xing , Xiaowen Shi , Yuan-Gen Wang , Chunsheng Yang

Existing deepfake detectors face several challenges in achieving robustness and generalization. One of the primary reasons is their limited ability to extract relevant information from forgery videos, especially in the presence of various…

计算机视觉与模式识别 · 计算机科学 2023-05-01 Zhiyuan Yan , Peng Sun , Yubo Lang , Shuo Du , Shanzhuo Zhang , Wei Wang , Lei Liu

Existing fine-grained image retrieval (FGIR) methods predominantly rely on supervision from predefined categories to learn discriminative representations for retrieving fine-grained objects. However, they inadvertently introduce…

计算机视觉与模式识别 · 计算机科学 2025-07-30 Shijie Wang , Jian Shi , Haojie Li

Fine-Grained Image Retrieval~(FGIR) faces challenges in learning discriminative visual representations to retrieve images with similar fine-grained features. Current leading FGIR solutions typically follow two regimes: enforce pairwise…

计算机视觉与模式识别 · 计算机科学 2025-12-02 Xin Jiang , Meiqi Cao , Hao Tang , Fei Shen , Zechao Li

While weakly supervised multi-view face reconstruction (MVR) is garnering increased attention, one critical issue still remains open: how to effectively interact and fuse multiple image information to reconstruct high-precision 3D models.…

计算机视觉与模式识别 · 计算机科学 2024-12-25 Weiguang Zhao , Chaolong Yang , Jianan Ye , Rui Zhang , Yuyao Yan , Xi Yang , Bin Dong , Amir Hussain , Kaizhu Huang

Fine-grained image classification, the task of distinguishing between visually similar subcategories within a broader category (e.g., bird species, car models, flower types), is a challenging computer vision problem. Traditional approaches…

计算机视觉与模式识别 · 计算机科学 2025-12-23 Dmitry Demidov , Zaigham Zaheer , Omkar Thawakar , Salman Khan , Fahad Shahbaz Khan

Recent advances in deep generative models have made it easier to manipulate face videos, raising significant concerns about their potential misuse for fraud and misinformation. Existing detectors often perform well in in-domain scenarios…

计算机视觉与模式识别 · 计算机科学 2026-01-26 Yinqi Cai , Jichang Li , Zhaolun Li , Weikai Chen , Rushi Lan , Xi Xie , Xiaonan Luo , Guanbin Li

The increasing realism and accessibility of deepfakes have raised critical concerns about media authenticity and information integrity. Despite recent advances, deepfake detection models often struggle to generalize beyond their training…

计算机视觉与模式识别 · 计算机科学 2025-11-11 Stelios Mylonas , Symeon Papadopoulos

Recent advances in multimodal large language models (MLLMs) have demonstrated strong capabilities in understanding general visual content. However, these general-domain MLLMs perform poorly in face perception tasks, often producing…

计算机视觉与模式识别 · 计算机科学 2025-04-29 Jingzhi Li , Changjiang Luo , Ruoyu Chen , Hua Zhang , Wenqi Ren , Jianhou Gan , Xiaochun Cao

Detecting diffusion-generated images has recently grown into an emerging research area. Existing diffusion-based datasets predominantly focus on general image generation. However, facial forgeries, which pose a more severe social risk, have…

计算机视觉与模式识别 · 计算机科学 2024-01-30 Harry Cheng , Yangyang Guo , Tianyi Wang , Liqiang Nie , Mohan Kankanhalli

The recovery of high-quality images from images corrupted by lens flare presents a significant challenge in low-level vision. Contemporary deep learning methods frequently entail training a lens flare removing model from scratch. However,…

计算机视觉与模式识别 · 计算机科学 2025-01-07 Tianwen Zhou , Qihao Duan , Zitong Yu

Vision-Language Models (VLMs) have demonstrated remarkable capabilities in cross-modal understanding and generation by integrating visual and textual information. While instruction tuning and parameter-efficient fine-tuning methods have…

机器学习 · 计算机科学 2025-06-12 Weiying Zheng , Ziyue Lin , Pengxin Guo , Yuyin Zhou , Feifei Wang , Liangqiong Qu

The rapid development of generative AI is a double-edged sword, which not only facilitates content creation but also makes image manipulation easier and more difficult to detect. Although current image forgery detection and localization…

计算机视觉与模式识别 · 计算机科学 2025-04-15 Zhipei Xu , Xuanyu Zhang , Runyi Li , Zecheng Tang , Qing Huang , Jian Zhang

High-dimensional images, known for their rich semantic information, are widely applied in remote sensing and other fields. The spatial information in these images reflects the object's texture features, while the spectral information…

计算机视觉与模式识别 · 计算机科学 2023-11-21 Daixun Li , Weiying Xie , Jiaqing Zhang , Yunsong Li

Existing Visual Speech Recognition (VSR) systems commonly rely on left-to-right autoregressive decoding, which can force premature decisions on visually ambiguous tokens before sufficient context is available. We propose DLLM-VSR, to the…

人工智能 · 计算机科学 2026-05-28 Jeong Hun Yeo , Chae Won Kim , Hyeongseop Rha , Yong Man Ro

Existing face super-resolution (FSR) methods have made significant advancements, but they primarily super-resolve face with limited visual information, original pixel-wise space in particular, commonly overlooking the pluralistic clues,…

计算机视觉与模式识别 · 计算机科学 2024-11-15 Chenyang Wang , Wenjie An , Kui Jiang , Xianming Liu , Junjun Jiang

Deepfake Generation Techniques are evolving at a rapid pace, making it possible to create realistic manipulated images and videos and endangering the serenity of modern society. The continual emergence of new and varied techniques brings…

计算机视觉与模式识别 · 计算机科学 2022-06-29 Davide Alessandro Coccomini , Roberto Caldelli , Fabrizio Falchi , Claudio Gennaro , Giuseppe Amato

Large Vision Language Models (LVLMs) have achieved significant progress in integrating visual and textual inputs for multimodal reasoning. However, a recurring challenge is ensuring these models utilize visual information as effectively as…

计算机视觉与模式识别 · 计算机科学 2025-03-20 Estelle Aflalo , Gabriela Ben Melech Stan , Tiep Le , Man Luo , Shachar Rosenman , Sayak Paul , Shao-Yen Tseng , Vasudev Lal