English
Related papers

Related papers: EvolveReason: Self-Evolving Reasoning Paradigm for…

200 papers

Fake News and especially deepfakes (generated, non-real image or video content) have become a serious topic over the last years. With the emergence of machine learning algorithms it is now easier than ever before to generate such fake…

Computer Vision and Pattern Recognition · Computer Science 2025-03-17 Lukas Kroiß , Johannes Reschke

Face Forgery Detection (FFD), or Deepfake detection, aims to determine whether a digital face is real or fake. Due to different face synthesis algorithms with diverse forgery patterns, FFD models often overfit specific patterns in training…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Zonghui Guo , Yingjie Liu , Jie Zhang , Haiyong Zheng , Shiguang Shan

Over the past few decades, numerous attempts have been made to address the problem of recovering a high-resolution (HR) facial image from its corresponding low-resolution (LR) counterpart, a task commonly referred to as face hallucination.…

Computer Vision and Pattern Recognition · Computer Science 2021-11-23 Ali Abbasi , Mohammad Rahmati

Over recent years, deep convolutional neural networks have significantly advanced the field of face recognition techniques for both verification and identification purposes. Despite the impressive accuracy, these neural networks are often…

Computer Vision and Pattern Recognition · Computer Science 2024-07-09 Yuhang Lu , Zewei Xu , Touradj Ebrahimi

Hallucination has been a major problem for large language models and remains a critical challenge when it comes to multimodality in which vision-language models (VLMs) have to deal with not just textual but also visual inputs. Despite rapid…

Computer Vision and Pattern Recognition · Computer Science 2024-07-23 Zhecan Wang , Garrett Bingham , Adams Yu , Quoc Le , Thang Luong , Golnaz Ghiasi

The emergence of LLMs, like ChatGPT and Gemini, has marked the modern era of artificial intelligence applications characterized by high-impact applications generating text, images, and videos. However, these models usually ensue with one…

Computation and Language · Computer Science 2025-07-08 Abdennour Boulesnane , Abdelhakim Souilah

Large Vision-Language Models (LVLMs) have shown remarkable capabilities, yet hallucinations remain a persistent challenge. This work presents a systematic analysis of the internal evolution of visual perception and token generation in…

Computer Vision and Pattern Recognition · Computer Science 2025-12-23 Guangtao Lyu , Xinyi Cheng , Chenghao Xu , Qi Liu , Muli Yang , Fen Fang , Huilin Chen , Jiexi Yan , Xu Yang , Cheng Deng

Vision-language models benefit from high-resolution images, but the increase in visual-token count incurs high compute overhead. Humans resolve this tension via foveation: a coarse view guides "where to look", while selectively acquired…

Computer Vision and Pattern Recognition · Computer Science 2026-04-24 Juhong Min , Lazar Valkov , Vitali Petsiuk , Hossein Souri , Deen Dayal Mohan

The discriminability of feature representation is the key to open-set face recognition. Previous methods rely on the learnable weights of the classification layer that represent the identities. However, the evaluation process learns no…

Computer Vision and Pattern Recognition · Computer Science 2023-04-25 Youzhe Song , Feng Wang

In the rapidly evolving landscape of digital security, biometric authentication systems, particularly facial recognition, have emerged as integral components of various security protocols. However, the reliability of these systems is…

Computer Vision and Pattern Recognition · Computer Science 2026-04-13 Oleksandr Kuznetsov , Emanuele Frontoni , Luca Romeo , Riccardo Rosati , Andrea Maranesi , Alessandro Muscatello

Vision-Language Models (VLMs) struggle with complex image annotation tasks, such as emotion classification and context-driven object detection, which demand sophisticated reasoning. Standard Supervised Fine-Tuning (SFT) focuses solely on…

Machine Learning · Computer Science 2025-09-16 Suhang Hu , Wei Hu , Yuhang Su , Fan Zhang

Visual Retrieval-Augmented Generation (VRAG) enhances Vision-Language Models (VLMs) by incorporating external visual documents to address a given query. Existing VRAG frameworks usually depend on rigid, pre-defined external tools to extend…

Artificial Intelligence · Computer Science 2026-04-10 Yuqi Xiong , Chunyi Peng , Zhipeng Xu , Zhenghao Liu , Zulong Chen , Yukun Yan , Shuo Wang , Yu Gu , Ge Yu

A fundamental challenge in artificial intelligence involves understanding the cognitive mechanisms underlying visual reasoning in sophisticated models like Vision-Language Models (VLMs). How do these models integrate visual perception with…

Computer Vision and Pattern Recognition · Computer Science 2025-05-07 Mohit Vaishnav , Tanel Tammet

Face recognition remains vulnerable to presentation attacks, calling for robust Face Anti-Spoofing (FAS) solutions. Recent MLLM-based FAS methods reformulate the binary classification task as the generation of brief textual descriptions to…

Computer Vision and Pattern Recognition · Computer Science 2026-03-23 Haoyuan Zhang , Keyao Wang , Guosheng Zhang , Haixiao Yue , Zhiwen Tan , Siran Peng , Tianshuo Zhang , Xiao Tan , Kunbin Chen , Wei He , Jingdong Wang , Ajian Liu , Xiangyu Zhu , Zhen Lei

Facial emotion perception in the vision large language model (VLLM) is crucial for achieving natural human-machine interaction. However, creating high-quality annotations for both coarse- and fine-grained facial emotion analysis demands…

Machine Learning · Computer Science 2025-05-27 Feifan Wang , Tengfei Song , Minggui He , Chang Su , Zhanglin Wu , Hao Yang , Wenming Zheng , Osamu Yoshie

While Retrieval-Augmented Generation (RAG) mitigates hallucination and knowledge staleness in Large Language Models (LLMs), existing frameworks often falter on complex, multi-hop queries that require synthesizing information from disparate…

Computation and Language · Computer Science 2025-10-28 Mohammad Aghajani Asl , Majid Asgari-Bidhendi , Behrooz Minaei-Bidgoli

Retrieval-Augmented Generation (RAG) significantly improves the factuality of Large Language Models (LLMs), yet standard pipelines often lack mechanisms to verify inter- mediate reasoning, leaving them vulnerable to hallucinations in…

Computation and Language · Computer Science 2026-03-12 Eeham Khan , Luis Rodriguez , Marc Queudot

Vision language models (VLMs) have achieved impressive performance across a variety of computer vision tasks. However, the multimodal reasoning capability has not been fully explored in existing models. In this paper, we propose a…

Computer Vision and Pattern Recognition · Computer Science 2025-12-08 Xintong Zhang , Zhi Gao , Bofei Zhang , Pengxiang Li , Xiaowen Zhang , Yang Liu , Tao Yuan , Yuwei Wu , Yunde Jia , Song-Chun Zhu , Qing Li

Existing methods for deepfake detection aim to develop generalizable detectors. Although "generalizable" is the ultimate target once and for all, with limited training forgeries and domains, it appears idealistic to expect generalization…

Computer Vision and Pattern Recognition · Computer Science 2025-12-05 Jikang Cheng , Renye Yan , Zhiyuan Yan , Yaozhong Gan , Xueyi Zhang , Zhongyuan Wang , Wei Peng , Ling Liang

Recent DeepFake detection methods have shown excellent performance on public datasets but are significantly degraded on new forgeries. Solving this problem is important, as new forgeries emerge daily with the continuously evolving…

Computer Vision and Pattern Recognition · Computer Science 2024-08-20 Qingxuan Lv , Yuezun Li , Junyu Dong , Sheng Chen , Hui Yu , Huiyu Zhou , Shu Zhang
‹ Prev 1 8 9 10 Next ›