中文
相关论文

相关论文: Beyond Artificial Misalignment: Detecting and Grou…

200 篇论文

Visual grounding focuses on detecting objects from images based on language expressions. Recent Large Vision-Language Models (LVLMs) have significantly advanced visual grounding performance by training large models with large-scale…

计算机视觉与模式识别 · 计算机科学 2025-09-25 Yangxiao Lu , Ruosen Li , Liqiang Jing , Jikai Wang , Xinya Du , Yunhui Guo , Nicholas Ruozzi , Yu Xiang

Image manipulation detection and localization have received considerable attention from the research community given the blooming of Generative Models (GMs). Detection methods that follow a passive approach may overfit to specific GMs,…

计算机视觉与模式识别 · 计算机科学 2024-09-27 Filippo Bartolucci , Iacopo Masi , Giuseppe Lisanti

As a new communication paradigm, semantic communication has received widespread attention in communication fields. However, since the decoding of semantic signals relies on contextual knowledge, misalignment between the starting position of…

信号处理 · 电气工程与系统科学 2023-12-19 Xiaoyi Liu , Haotai Liang , Chen Dong , Xiaodong Xu

Bimanual mobile manipulation requires a seamless integration between high-level semantic reasoning and safe, compliant physical interaction - a challenge that end-to-end models approach opaquely and classical controllers lack the context to…

Text-based person anomaly search retrieves specific behavioral events from surveillance archives using natural-language queries. Although recent pose-aware methods align geometric structures well, they face a fundamental Pose-Semantic Gap:…

计算机视觉与模式识别 · 计算机科学 2026-05-28 Zequn Xie , Guijin Luo , Chuxin Wang , Sihang Cai , Tao Jin , Zhou Zhao , Yixuan Tang

Content generation and manipulation approaches based on deep learning methods have seen significant advancements, leading to an increased need for techniques to detect whether an image has been generated or edited. Another area of research…

计算机视觉与模式识别 · 计算机科学 2025-01-20 Philip Wootaek Shin , Jack Sampson , Vijaykrishnan Narayanan , Andres Marquez , Mahantesh Halappanavar

This survey provides a comprehensive overview of recent advances in multimodal alignment and fusion within the field of machine learning, driven by the increasing availability and diversity of data modalities such as text, images, audio,…

计算机视觉与模式识别 · 计算机科学 2025-10-14 Songtao Li , Hao Tang

Multimodal deepfakes can exhibit subtle visual artifacts and cross-modal inconsistencies, which remain challenging to detect, especially when detectors are trained primarily on curated synthetic forgeries. Such synthetic dependence can…

计算机视觉与模式识别 · 计算机科学 2026-03-27 Sahibzada Adil Shahzad , Ammarah Hashmi , Junichi Yamagishi , Yusuke Yasuda , Yu Tsao , Chia-Wen Lin , Yan-Tsung Peng , Hsin-Min Wang

Geometric information in the normalized digital surface models (nDSM) is highly correlated with the semantic class of the land cover. Exploiting two modalities (RGB and nDSM (height)) jointly has great potential to improve the segmentation…

计算机视觉与模式识别 · 计算机科学 2023-05-25 Zhitong Xiong , Sining Chen , Yi Wang , Lichao Mou , Xiao Xiang Zhu

The rapid advancement of Generative Artificial Intelligence has fueled deepfake proliferation-synthetic media encompassing fully generated content and subtly edited authentic material-posing challenges to digital security, misinformation…

密码学与安全 · 计算机科学 2025-07-30 Naseem Khan , Tuan Nguyen , Amine Bermak , Issa Khalil

Semi-supervised Domain Generalization (SSDG) addresses the challenge of generalizing to unseen target domains with limited labeled data. Existing SSDG methods highlight the importance of achieving high pseudo-labeling (PL) accuracy and…

计算机视觉与模式识别 · 计算机科学 2026-01-21 Muditha Fernando , Kajhanan Kailainathan , Krishnakanth Nagaratnam , Isuranga Udaravi Bandara Senavirathne , Ranga Rodrigo

Despite the strong abilities, large language models (LLMs) still suffer from hallucinations and reliance on outdated knowledge, raising concerns in knowledge-intensive tasks. Graph-based retrieval-augmented generation (GRAG) enriches LLMs…

In recent years, detecting fake multimodal content on social media has drawn increasing attention. Two major forms of deception dominate: human-crafted misinformation (e.g., rumors and misleading posts) and AI-generated content produced by…

人工智能 · 计算机科学 2025-10-17 Haiyang Li , Yaxiong Wang , Shengeng Tang , Lianwei Wu , Lechao Cheng , Zhun Zhong

Large Vision-Language Models (LVLMs) have made remarkable strides in multimodal tasks such as visual question answering, visual grounding, and complex reasoning. However, they remain limited by static training data, susceptibility to…

人工智能 · 计算机科学 2025-08-27 Chan-Wei Hu , Yueqi Wang , Shuo Xing , Chia-Ju Chen , Suofei Feng , Ryan Rossi , Zhengzhong Tu

Multi-modality image fusion, particularly infrared and visible, plays a crucial role in integrating diverse modalities to enhance scene understanding. Although early research prioritized visual quality, preserving fine details and adapting…

计算机视觉与模式识别 · 计算机科学 2025-03-26 Guanyao Wu , Haoyu Liu , Hongming Fu , Yichuan Peng , Jinyuan Liu , Xin Fan , Risheng Liu

Multimodal deepfakes involving audiovisual manipulations are a growing threat because they are difficult to detect with the naked eye or using unimodal deep learningbased forgery detection methods. Audiovisual forensic models, while more…

计算机视觉与模式识别 · 计算机科学 2024-11-15 Sahibzada Adil Shahzad , Ammarah Hashmi , Yan-Tsung Peng , Yu Tsao , Hsin-Min Wang

The rapid development of generative AI facilitates content creation and makes image manipulation easier and more difficult to detect. While multimodal Large Language Models (LLMs) have encoded rich world knowledge, they are not inherently…

计算机视觉与模式识别 · 计算机科学 2025-04-17 Yiran He , Yun Cao , Bowen Yang , Zeyu Zhang

Fine-grained robotic manipulation requires grounding natural language into appropriate affordance targets. However, most existing methods driven by foundation models often compress rich semantics into oversimplified affordances, preventing…

机器人学 · 计算机科学 2025-12-30 Chenyu Su , Weiwei Shang , Chen Qian , Fei Zhang , Shuang Cong

Video Temporal Grounding (VTG) aims to localize temporal segments in long, untrimmed videos that align with a given natural language query. This task typically comprises two subtasks: Moment Retrieval (MR) and Highlight Detection (HD).…

计算机视觉与模式识别 · 计算机科学 2025-10-24 Minseok Kang , Minhyeok Lee , Minjung Kim , Donghyeong Kim , Sangyoun Lee

Multimodal fake news detection often involves modelling heterogeneous data sources, such as vision and language. Existing detection methods typically rely on fusion effectiveness and cross-modal consistency to model the content,…

机器学习 · 计算机科学 2025-03-04 Lingzhi Shen , Yunfei Long , Xiaohao Cai , Imran Razzak , Guanming Chen , Kang Liu , Shoaib Jameel