English
Related papers

Related papers: HAMLET-FFD: Hierarchical Adaptive Multi-modal Lear…

200 papers

The goal of this work is to enhance balanced multimodal understanding in audio-visual large language models (AV-LLMs) by addressing modality bias without additional training. In current AV-LLMs, audio and video features are typically…

Computer Vision and Pattern Recognition · Computer Science 2025-10-01 Chaeyoung Jung , Youngjoon Jang , Jongmin Choi , Joon Son Chung

The rapid development of generative AI is a double-edged sword, which not only facilitates content creation but also makes image manipulation easier and more difficult to detect. Although current image forgery detection and localization…

Computer Vision and Pattern Recognition · Computer Science 2025-04-15 Zhipei Xu , Xuanyu Zhang , Runyi Li , Zecheng Tang , Qing Huang , Jian Zhang

Recent advances in tuning-free personalized image generation based on diffusion models are impressive. However, to improve subject fidelity, existing methods either retrain the diffusion model or infuse it with dense visual embeddings, both…

Computer Vision and Pattern Recognition · Computer Science 2024-03-25 Zhichao Wei , Qingkun Su , Long Qin , Weizhi Wang

Multimodal large language models (MLLMs) have demonstrated strong capabilities in vision-related tasks, capitalizing on their visual semantic comprehension and reasoning capabilities. However, their ability to detect subtle visual spoofing…

Computer Vision and Pattern Recognition · Computer Science 2025-04-22 Yichen Shi , Yuhao Gao , Yingxin Lai , Hongyang Wang , Jun Feng , Lei He , Jun Wan , Changsheng Chen , Zitong Yu , Xiaochun Cao

Face forgery detection is raising ever-increasing interest in computer vision since facial manipulation technologies cause serious worries. Though recent works have reached sound achievements, there are still unignorable problems: a)…

Computer Vision and Pattern Recognition · Computer Science 2021-03-17 Jiaming Li , Hongtao Xie , Jiahong Li , Zhongyuan Wang , Yongdong Zhang

The stunning progress in face manipulation methods has made it possible to synthesize realistic fake face images, which poses potential threats to our society. It is urgent to have face forensics techniques to distinguish those tampered…

Computer Vision and Pattern Recognition · Computer Science 2019-12-13 Jia Li , Tong Shen , Wei Zhang , Hui Ren , Dan Zeng , Tao Mei

Recent advances in representation learning have demonstrated an ability to represent information from different modalities such as video, text, and audio in a single high-level embedding vector. In this work we present a self-supervised…

Computer Vision and Pattern Recognition · Computer Science 2021-06-11 Alexander H. Liu , SouYoung Jin , Cheng-I Jeff Lai , Andrew Rouditchenko , Aude Oliva , James Glass

Face manipulation methods develop rapidly in recent years, whose potential risk to society accounts for the emerging of researches on detection methods. However, due to the diversity of manipulation methods and the high quality of fake…

Computer Vision and Pattern Recognition · Computer Science 2021-05-24 Zehao Chen , Hua Yang

Heterogeneous face matching is a challenge issue in face recognition due to large domain difference as well as insufficient pairwise images in different modalities during training. This paper proposes a coupled deep learning (CDL) approach…

Computer Vision and Pattern Recognition · Computer Science 2017-11-17 Xiang Wu , Lingxiao Song , Ran He , Tieniu Tan

Detecting maliciously falsified facial images and videos has attracted extensive attention from digital-forensics and computer-vision communities. An important topic in manipulation detection is the localization of the fake regions.…

Computer Vision and Pattern Recognition · Computer Science 2023-04-18 Weinan Guan , Wei Wang , Jing Dong , Bo Peng , Tieniu Tan

The rapid advancement of facial forgery techniques poses severe threats to public trust and information security, making facial DeepFake detection a critical research priority. Continual learning provides an effective approach to adapt…

Computer Vision and Pattern Recognition · Computer Science 2026-04-10 Yushuo Zhang , Yu Cheng , Yongkang Hu , Jiuan Zhou , Jiawei Chen , Yuan Xie , Zhaoxia Yin

Despite the impressive progress of general face detection, the tuning of hyper-parameters and architectures is still critical for the performance of a domain-specific face detector. Though existing AutoML works can speedup such process,…

Computer Vision and Pattern Recognition · Computer Science 2022-03-17 Chenqian Yan , Yuge Zhang , Quanlu Zhang , Yaming Yang , Xinyang Jiang , Yuqing Yang , Baoyuan Wang

Data augmentation is crucial for improving the robustness of face detection systems, especially under challenging conditions such as occlusion, illumination variation, and complex environments. Traditional copy paste augmentation often…

Computer Vision and Pattern Recognition · Computer Science 2025-12-15 Qiushi Guo

Current multi-modal models exhibit a notable misalignment with the human visual system when identifying objects that are visually assimilated into the background. Our observations reveal that these multi-modal models cannot distinguish…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Ruolin Shen , Xiaozhong Ji , Kai WU , Jiangning Zhang , Yijun He , HaiHua Yang , Xiaobin Hu , Xiaoyu Sun

Driven by the rapid progress in vision-language models (VLMs), the responsible behavior of large-scale multimodal models has become a prominent research area, particularly focusing on hallucination detection and factuality checking. In this…

Computer Vision and Pattern Recognition · Computer Science 2025-06-18 Zijian Zhang , Xuecheng Wu , Danlei Huang , Siyu Yan , Chong Peng , Xuezhi Cao

Multi-modal face anti-spoofing (FAS) aims to detect genuine human presence by extracting discriminative liveness cues from multiple modalities, such as RGB, infrared (IR), and depth images, to enhance the robustness of biometric…

Computer Vision and Pattern Recognition · Computer Science 2025-07-09 Jun-Xiong Chong , Fang-Yu Hsu , Ming-Tsung Hsu , Yi-Ting Lin , Kai-Heng Chien , Chiou-Ting Hsu , Pei-Kai Huang

We introduce a non-parametric hierarchical Bayesian approach for open-ended 3D object categorization, named the Local Hierarchical Dirichlet Process (Local-HDP). This method allows an agent to learn independent topics for each category…

Computer Vision and Pattern Recognition · Computer Science 2021-04-13 H. Ayoobi , H. Kasaei , M. Cao , R. Verbrugge , B. Verheij

High-resolution simulation models are essential for representing complex physical systems, yet their substantial computational cost severely limits the number of feasible high-fidelity (HF) evaluations. This problem is often addressed…

Methodology · Statistics 2026-04-21 Hossein Mohammadi

The rapid advancement of deepfake technologies raises significant concerns about the security of face recognition systems. While existing methods leverage the clues left by deepfake techniques for face forgery detection, malicious users may…

Computer Vision and Pattern Recognition · Computer Science 2023-10-10 Weihua Liu , Lin Li , Chaochao Lin , Said Boumaraf

With the advance in user-friendly and powerful video editing tools, anyone can easily manipulate videos without leaving prominent visual traces. Frame-rate up-conversion (FRUC), a representative temporal-domain operation, increases the…

Multimedia · Computer Science 2021-03-26 Minseok Yoon , Seung-Hun Nam , In-Jae Yu , Wonhyuk Ahn , Myung-Joon Kwon , Heung-Kyu Lee