English
Related papers

Related papers: HAMLET-FFD: Hierarchical Adaptive Multi-modal Lear…

200 papers

Appearance-based gaze estimation, aiming to predict accurate 3D gaze direction from a single facial image, has made promising progress in recent years. However, most methods suffer significant performance degradation in cross-domain…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Qida Tan , Hongyu Yang , Wenchao Du

With the rapid growth of video data, text-video retrieval technology has become increasingly important in numerous application scenarios such as recommendation and search. Early text-video retrieval methods suffer from two critical…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Jiaao Yu , Mingjie Han , Tao Gong , Jian Zhang , Man Lan

Generative models have enabled the creation of highly realistic facial-synthetic images, raising significant concerns due to their potential for misuse. Despite rapid advancements in the field of deepfake detection, developing efficient…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Yue-Hua Han , Tai-Ming Huang , Kai-Lung Hua , Jun-Cheng Chen

Scale variation is one of the most challenging problems in face detection. Modern face detectors employ feature pyramids to deal with scale variation. However, it might break the feature consistency across different scales of faces. In this…

Computer Vision and Pattern Recognition · Computer Science 2021-05-24 Leilei Cao , Yao Xiao , Lin Xu

Deception detection is an interdisciplinary field attracting researchers from psychology, criminology, computer science, and economics. We propose a multimodal approach combining deep learning and discriminative models for automated…

Computer Vision and Pattern Recognition · Computer Science 2024-11-07 Laslo Dinges , Marc-André Fiedler , Ayoub Al-Hamadi , Thorsten Hempel , Ahmed Abdelrahman , Joachim Weimann , Dmitri Bershadskyy

Feature matching is a cornerstone task in computer vision, essential for applications such as image retrieval, stereo matching, 3D reconstruction, and SLAM. This survey comprehensively reviews modality-based feature matching, exploring…

Computer Vision and Pattern Recognition · Computer Science 2025-07-31 Weide Liu , Wei Zhou , Jun Liu , Ping Hu , Jun Cheng , Jungong Han , Weisi Lin

Face presentation attack detection (PAD) is an essential measure to protect face recognition systems from being spoofed by malicious users and has attracted great attention from both academia and industry. Although most of the existing…

Computer Vision and Pattern Recognition · Computer Science 2021-10-22 Zhi Li , Haoliang Li , Xin Luo , Yongjian Hu , Kwok-Yan Lam , Alex C. Kot

Open-set face forgery detection poses significant security threats and presents substantial challenges for existing detection models. These detectors primarily have two limitations: they cannot generalize across unknown forgery domains and…

Computer Vision and Pattern Recognition · Computer Science 2026-02-27 Chenqi Kong , Anwei Luo , Peijun Bao , Haoliang Li , Renjie Wan , Zengwei Zheng , Anderson Rocha , Alex C. Kot

The rapid evolution of deepfake generation technologies poses critical challenges for detection systems, as non-continual learning methods demand frequent and expensive retraining. We reframe deepfake detection (DFD) as a Continual Learning…

Machine Learning · Computer Science 2025-09-11 Federico Fontana , Anxhelo Diko , Romeo Lanzino , Marco Raoul Marini , Bachir Kaddar , Gian Luca Foresti , Luigi Cinque

Hallucination remains a critical barrier for deploying large language models (LLMs) in reliability-sensitive applications. Existing detection methods largely fall into two categories: factuality checking, which is fundamentally constrained…

Computation and Language · Computer Science 2025-09-17 Jinxin Li , Gang Tu , ShengYu Cheng , Junjie Hu , Jinting Wang , Rui Chen , Zhilong Zhou , Dongbo Shan

Facial expression classification remains a challenging task due to the high dimensionality and inherent complexity of facial image data. This paper presents Hy-Facial, a hybrid feature extraction framework that integrates both deep learning…

Computer Vision and Pattern Recognition · Computer Science 2025-10-01 Xinjin Li , Yu Ma , Kaisen Ye , Jinghan Cao , Minghao Zhou , Yeyang Zhou

Hallucinations in Large Language Models (LLMs) represent a critical barrier to their reliable deployment, a vulnerability heavily exacerbated in non-English and resource-constrained contexts. Existing detection approaches that rely on…

Computation and Language · Computer Science 2026-05-26 Riasad Alvi , Nurul Labib Sayeedi , Md. Faiyaz Abdullah Sayeedi

Multimodal large language models have unlocked new possibilities for various multimodal tasks. However, their potential in image manipulation detection remains unexplored. When directly applied to the IMD task, M-LLMs often produce…

Computer Vision and Pattern Recognition · Computer Science 2025-12-30 Zhihao Sun , Haoran Jiang , Haoran Chen , Yixin Cao , Xipeng Qiu , Zuxuan Wu , Yu-Gang Jiang

The embodied intelligence bridges the physical world and information space. As its typical physical embodiment, humanoid robots have shown great promise through robot learning algorithms in recent years. In this study, a hardware platform,…

Robotics · Computer Science 2025-10-17 Jiaxin Huang , Hanyu Liu , Yunsheng Ma , Jian Shen , Yilin Zheng , Jiayi Wen , Baishu Wan , Pan Li , Zhigong Song

Surveillance and security scenarios usually require high efficient facial image compression scheme for face recognition and identification. While either traditional general image codecs or special facial image compression schemes only…

Multimedia · Computer Science 2019-03-08 Zhibo Chen , Tianyu He

Heterogeneous Face Recognition (HFR) focuses on matching faces from different domains, for instance, thermal to visible images, making Face Recognition (FR) systems more versatile for challenging scenarios. However, the domain gap between…

Computer Vision and Pattern Recognition · Computer Science 2024-04-23 Anjith George , Sebastien Marcel

Learning from Demonstration (LfD) offers a promising paradigm for robot skill acquisition. Recent approaches attempt to extract manipulation commands directly from video demonstrations, yet face two critical challenges: (1) general video…

Robotics · Computer Science 2026-02-24 Thanh Nguyen Canh , Thanh-Tuan Tran , Haolan Zhang , Ziyan Gao , Nak Young Chong , Xiem HoangVan

Low-light image enhancement techniques have significantly progressed, but unstable image quality recovery and unsatisfactory visual perception are still significant challenges. To solve these problems, we propose a novel and robust…

Computer Vision and Pattern Recognition · Computer Science 2024-04-18 Minglong Xue , Jinhong He , Wenhai Wang , Mingliang Zhou

We describe Forensics Adapter, an adapter network designed to transform CLIP into an effective and generalizable face forgery detector. Although CLIP is highly versatile, adapting it for face forgery detection is non-trivial as…

Computer Vision and Pattern Recognition · Computer Science 2025-05-26 Xinjie Cui , Yuezun Li , Delong Zhu , Jiaran Zhou , Junyu Dong , Siwei Lyu

As artificial intelligence systems increasingly operate in Real-world environments, the integration of multi-modal data sources such as vision, language, and audio presents both unprecedented opportunities and critical challenges for…

Machine Learning · Computer Science 2025-07-01 Sree Bhargavi Balija