中文
相关论文

相关论文: LAVID: An Agentic LVLM Framework for Diffusion-Gen…

200 篇论文

Vision-Language Models (VLMs) are powerful open-set reasoners, yet their direct use as anomaly detectors in video surveillance is fragile: without calibrated anomaly priors, they alternate between missed detections and hallucinated false…

计算机视觉与模式识别 · 计算机科学 2026-04-03 Mohamed Eltahir , Ahmed O. Ibrahim , Obada Siralkhatim , Tabarak Abdallah , Sondos Mohamed

Out-of-distribution (OOD) detection is essential for reliable and trustworthy machine learning. Recent multi-modal OOD detection leverages textual information from in-distribution (ID) class names for visual OOD detection, yet it currently…

计算与语言 · 计算机科学 2023-10-13 Yi Dai , Hao Lang , Kaisheng Zeng , Fei Huang , Yongbin Li

Large Vision-Language Models (LVLMs) generate contextually relevant responses by jointly interpreting visual and textual inputs. However, our finding reveals they often mistakenly perceive text inputs lacking visual evidence as being part…

计算机视觉与模式识别 · 计算机科学 2025-09-08 Sohee Kim , Soohyun Ryu , Joonhyung Park , Eunho Yang

The development of Large Vision-Language Models (LVLMs) is striving to catch up with the success of Large Language Models (LLMs), yet it faces more challenges to be resolved. Very recent works enable LVLMs to localize object-level visual…

计算机视觉与模式识别 · 计算机科学 2024-03-20 Zhipeng Huang , Zhizheng Zhang , Zheng-Jun Zha , Yan Lu , Baining Guo

The collection and detection of video anomaly data has long been a challenging problem due to its rare occurrence and spatio-temporal scarcity. Existing video anomaly detection (VAD) methods under perform in open-world scenarios. Key…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Zunkai Dai , Ke Li , Jiajia Liu , Jie Yang , Yuanyuan Qiao

Large Language Models (LLMs), with remarkable conversational capability, have emerged as AI assistants that can handle both visual and textual modalities. However, their effectiveness in joint video and language understanding has not been…

计算机视觉与模式识别 · 计算机科学 2025-03-18 Ruipu Luo , Ziwang Zhao , Min Yang , Zheming Yang , Minghui Qiu , Tao Wang , Zhongyu Wei , Yanhao Wang , Cen Chen

This work introduces Robots Imitating Generated Videos (RIGVid), a system that enables robots to perform complex manipulation tasks--such as pouring, wiping, and mixing--purely by imitating AI-generated videos, without requiring any…

机器人学 · 计算机科学 2026-05-14 Shivansh Patel , Shraddhaa Mohan , Hanlin Mai , Unnat Jain , Svetlana Lazebnik , Yunzhu Li

Assessing whether AI-generated images are substantially similar to source works is a crucial step in resolving copyright disputes. In this paper, we propose CopyJudge, a novel automated infringement identification framework that leverages…

计算机视觉与模式识别 · 计算机科学 2025-07-29 Shunchang Liu , Zhuan Shi , Lingjuan Lyu , Yaochu Jin , Boi Faltings

The rapid development of deep learning and generative AI technologies has profoundly transformed the digital contact landscape, creating realistic Deepfake that poses substantial challenges to public trust and digital media integrity. This…

计算机视觉与模式识别 · 计算机科学 2025-03-11 Ying Xu , Marius Pedersen , Kiran Raja

The advent of Large Language Models (LLMs) has significantly reshaped the trajectory of the AI revolution. Nevertheless, these LLMs exhibit a notable limitation, as they are primarily adept at processing textual information. To address this…

计算机视觉与模式识别 · 计算机科学 2025-10-15 Akash Ghosh , Arkadeep Acharya , Sriparna Saha , Vinija Jain , Aman Chadha

The growing capability of video generation poses escalating security risks, making reliable detection increasingly essential. In this paper, we introduce VideoVeritas, a framework that integrates fine-grained perception and fact-based…

计算机视觉与模式识别 · 计算机科学 2026-02-10 Hao Tan , Jun Lan , Senyuan Shi , Zichang Tan , Zijian Yu , Huijia Zhu , Weiqiang Wang , Jun Wan , Zhen Lei

In this paper, we introduce PruneVid, a visual token pruning method designed to enhance the efficiency of multi-modal video understanding. Large Language Models (LLMs) have shown promising performance in video tasks due to their extended…

计算机视觉与模式识别 · 计算机科学 2024-12-23 Xiaohu Huang , Hao Zhou , Kai Han

Current multimodal latent reasoning often relies on external supervision (e.g., auxiliary images), ignoring intrinsic visual attention dynamics. In this work, we identify a critical Perception Gap in distillation: student models frequently…

计算机视觉与模式识别 · 计算机科学 2026-01-16 Linquan Wu , Tianxiang Jiang , Yifei Dong , Haoyu Yang , Fengji Zhang , Shichaang Meng , Ai Xuan , Linqi Song , Jacky Keung

The development of large language models (LLMs) has successfully transformed knowledge-based systems such as open domain question nswering, which can automatically produce vast amounts of seemingly coherent information. Yet, those models…

人工智能 · 计算机科学 2026-01-28 Eduardo C. Garrido-Merchán , Cristina Puente

With the surge of large language models (LLMs), Large Vision-Language Models (VLMs)--which integrate vision encoders with LLMs for accurate visual grounding--have shown great potential in tasks like generalist agents and robotic control.…

计算机视觉与模式识别 · 计算机科学 2025-04-28 Hongyu Zhu , Sichu Liang , Wenwen Wang , Boheng Li , Tongxin Yuan , Fangqi Li , ShiLin Wang , Zhuosheng Zhang

Large vision-language models (VLMs) have shown promising capabilities in scene understanding, enhancing the explainability of driving behaviors and interactivity with users. Existing methods primarily fine-tune VLMs on on-board multi-view…

计算机视觉与模式识别 · 计算机科学 2025-08-19 Nan Song , Bozhou Zhang , Xiatian Zhu , Jiankang Deng , Li Zhang

A well-known dilemma in large vision-language models (e.g., GPT-4, LLaVA) is that while increasing the number of vision tokens generally enhances visual understanding, it also significantly raises memory and computational costs, especially…

计算机视觉与模式识别 · 计算机科学 2024-08-30 Shiwei Wu , Joya Chen , Kevin Qinghong Lin , Qimeng Wang , Yan Gao , Qianli Xu , Tong Xu , Yao Hu , Enhong Chen , Mike Zheng Shou

In recent years, online lecture videos have become an increasingly popular resource for acquiring new knowledge. Systems capable of effectively understanding/indexing lecture videos are thus highly desirable, enabling downstream tasks like…

计算机视觉与模式识别 · 计算机科学 2025-03-04 Kangda Wei , Zhengyu Zhou , Bingqing Wang , Jun Araki , Lukas Lange , Ruihong Huang , Zhe Feng

The ease of access to large language models (LLMs) has enabled a widespread of machine-generated texts, and now it is often hard to tell whether a piece of text was human-written or machine-generated. This raises concerns about potential…

The rapid development of Artificial Intelligence Generated Content (AIGC) techniques has enabled the creation of high-quality synthetic content, but it also raises significant security concerns. Current detection methods face two major…

计算机视觉与模式识别 · 计算机科学 2026-05-06 Changjiang Jiang , Wenhui Dong , Zhonghao Zhang , Fengchang Yu , Wei Peng , Xinbin Yuan , Yifei Bi , Ming Zhao , Zian Zhou , Chenyang Si , Caifeng Shan