中文
相关论文

相关论文: Unlocking the Capabilities of Large Vision-Languag…

200 篇论文

Previous studies on federated learning (FL) often encounter performance degradation due to data heterogeneity among different clients. In light of the recent advances in multimodal large language models (MLLMs), such as GPT-4v and LLaVA,…

人工智能 · 计算机科学 2024-12-03 Jianyi Zhang , Hao Frank Yang , Ang Li , Xin Guo , Pu Wang , Haiming Wang , Yiran Chen , Hai Li

Deepfake detection remains a critical challenge in the era of advanced generative models, particularly as synthetic media becomes more sophisticated. In this study, we explore the potential of state of the art multi-modal (reasoning) large…

计算机视觉与模式识别 · 计算机科学 2025-04-01 Simiao Ren , Yao Yao , Kidus Zewde , Zisheng Liang , Tsang , Ng , Ning-Yau Cheng , Xiaoou Zhan , Qinzhe Liu , Yifei Chen , Hengwei Xu

Advances in generative models have led to AI-generated images visually indistinguishable from authentic ones. Despite numerous studies on detecting AI-generated images with classifiers, a gap persists between such methods and human…

计算机视觉与模式识别 · 计算机科学 2025-08-05 Chuangchuang Tan , Jinglu Wang , Xiang Ming , Renshuai Tao , Yunchao Wei , Yao Zhao , Yan Lu

Vision-Language Models (VLMs) are increasingly used as perceptual modules for visual content reasoning, including through captioning and DeepFake detection. In this work, we expose a critical vulnerability of VLMs when exposed to subtle,…

计算机视觉与模式识别 · 计算机科学 2025-08-14 Jordan Vice , Naveed Akhtar , Yansong Gao , Richard Hartley , Ajmal Mian

In recent years, multimodal large language models (MLLMs) have made significant strides by training on vast high-quality image-text datasets, enabling them to generally understand images well. However, the inherent difficulty in explicitly…

计算机视觉与模式识别 · 计算机科学 2024-07-08 Yuanze Lin , Yunsheng Li , Dongdong Chen , Weijian Xu , Ronald Clark , Philip Torr , Lu Yuan

In recent years, the emergence of models capable of generating images from text has attracted considerable interest, offering the possibility of creating realistic images from text descriptions. Yet these advances have also raised concerns…

计算机视觉与模式识别 · 计算机科学 2024-04-04 Mamadou Keita , Wassim Hamidouche , Hassen Bougueffa , Abdenour Hadid , Abdelmalik Taleb-Ahmed

The increasing realism and accessibility of deepfakes have raised critical concerns about media authenticity and information integrity. Despite recent advances, deepfake detection models often struggle to generalize beyond their training…

计算机视觉与模式识别 · 计算机科学 2025-11-11 Stelios Mylonas , Symeon Papadopoulos

Multi-modal large language models (MLLMs) have shown remarkable abilities in various visual understanding tasks. However, MLLMs still struggle with fine-grained visual recognition (FGVR), which aims to identify subordinate-level categories…

计算机视觉与模式识别 · 计算机科学 2025-04-01 Hulingxiao He , Geng Li , Zijun Geng , Jinglin Xu , Yuxin Peng

Face forgery detection is essential in combating malicious digital face attacks. Previous methods mainly rely on prior expert knowledge to capture specific forgery clues, such as noise patterns, blending boundaries, and frequency artifacts.…

计算机视觉与模式识别 · 计算机科学 2023-04-26 Anwei Luo , Chenqi Kong , Jiwu Huang , Yongjian Hu , Xiangui Kang , Alex C. Kot

Vision-Language Models (VLMs) have demonstrated remarkable capabilities in cross-modal understanding and generation by integrating visual and textual information. While instruction tuning and parameter-efficient fine-tuning methods have…

机器学习 · 计算机科学 2025-06-12 Weiying Zheng , Ziyue Lin , Pengxin Guo , Yuyin Zhou , Feifei Wang , Liangqiong Qu

Large Visual Language Models (LVLMs) integrate visual and linguistic modalities, exhibiting exceptional performance across various multimodal tasks. Nevertheless, LVLMs remain vulnerable to the issue of object hallucinations. Previous…

计算机视觉与模式识别 · 计算机科学 2025-02-04 Chao Wang , Xuancheng Zhou , Weiwei Fu , Yang Zhou

Large Vision-Language Models (LVLMs) have demonstrated impressive multimodal reasoning capabilities, but they remain susceptible to hallucination, particularly object hallucination where non-existent objects or incorrect attributes are…

计算机视觉与模式识别 · 计算机科学 2025-08-06 Cong-Duy Nguyen , Xiaobao Wu , Duc Anh Vu , Shuai Zhao , Thong Nguyen , Anh Tuan Luu

Large language models (LLMs) have recently achieved significant success across various application domains, garnering substantial attention from different communities. Unfortunately, even for the best LLM, many \textit{faults} still exist…

软件工程 · 计算机科学 2024-11-06 Qiang Hu , Jin Wen , Maxime Cordy , Yuheng Huang , Wei Ma , Xiaofei Xie , Lei Ma

As ultra-realistic face forgery techniques emerge, deepfake detection has attracted increasing attention due to security concerns. Many detectors cannot achieve accurate results when detecting unseen manipulations despite excellent…

计算机视觉与模式识别 · 计算机科学 2022-11-08 Zihan Liu , Hanyi Wang , Shilin Wang

Speech deepfake detection (SDD) focuses on identifying whether a given speech signal is genuine or has been synthetically generated. Existing audio large language model (LLM)-based methods excel in content understanding; however, their…

声音 · 计算机科学 2026-02-02 Xiaoxuan Guo , Yuankun Xie , Haonan Cheng , Jiayi Zhou , Jian Liu , Hengyan Huang , Long Ye , Qin Zhang

Integrating multimodal knowledge into large language models (LLMs) represents a significant advancement in dialogue generation capabilities. However, the effective incorporation of such knowledge in zero-resource scenarios remains a…

计算与语言 · 计算机科学 2025-02-06 Bo Zhang , Hui Ma , Jian Ding , Jian Wang , Bo Xu , Hongfei Lin

Referring camouflaged object detection (Ref-COD) is a recently-proposed problem aiming to segment out specified camouflaged objects matched with a textual or visual reference. This task involves two major challenges: the COD domain-specific…

计算机视觉与模式识别 · 计算机科学 2023-11-30 Shupeng Cheng , Ge-Peng Ji , Pengda Qin , Deng-Ping Fan , Bowen Zhou , Peng Xu

Audio-visual deepfake detection (AVD) is increasingly important as modern generators can fabricate convincing speech and video. Most current multimodal detectors are small, task-specific models: they work well on curated tests but scale…

声音 · 计算机科学 2026-03-02 Songjun Cao , Yuqi Li , Yunpeng Luo , Jianjun Yin , Long Ma

Current video understanding models excel at recognizing "what" is happening but fall short in high-level cognitive tasks like causal reasoning and future prediction, a limitation rooted in their lack of commonsense world knowledge. To…

计算机视觉与模式识别 · 计算机科学 2025-12-30 L'ea Dubois , Klaus Schmidt , Chengyu Wang , Ji-Hoon Park , Lin Wang , Santiago Munoz

Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities in multimodal task reasoning. However, they often generate responses that appear plausible yet do not accurately reflect the visual content, a phenomenon known…

计算机视觉与模式识别 · 计算机科学 2025-03-18 Jiaqi Wang , Yifei Gao , Jitao Sang