English
Related papers

Related papers: DVAR: Adversarial Multi-Agent Debate for Video Aut…

200 papers

With rich visual data, such as images, becoming readily associated with items, visually-aware recommendation systems (VARS) have been widely used in different applications. Recent studies have shown that VARS are vulnerable to item-image…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Minglei Yin , Bin Liu , Neil Zhenqiang Gong , Xin Li

Video reasoning, which requires multi-step deduction across frames, remains a major challenge for multimodal large language models (MLLMs). While reinforcement learning (RL)-based methods enhance reasoning capabilities, they often rely on…

Computer Vision and Pattern Recognition · Computer Science 2025-11-21 Kun Ouyang , Yuanxin Liu , Linli Yao , Yishuo Cai , Hao Zhou , Jie Zhou , Fandong Meng , Xu Sun

Long-form video understanding remains challenging due to the extended temporal structure and dense multimodal cues. Despite recent progress, many existing approaches still rely on hand-crafted reasoning pipelines or employ token-consuming…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Yufei Yin , Qianke Meng , Minghao Chen , Jiajun Ding , Zhenwei Shao , Zhou Yu

Recent advances in video manipulation techniques have made the generation of fake videos more accessible than ever before. Manipulated videos can fuel disinformation and reduce trust in media. Therefore detection of fake videos has garnered…

Computer Vision and Pattern Recognition · Computer Science 2020-11-10 Shehzeen Hussain , Paarth Neekhara , Malhar Jere , Farinaz Koushanfar , Julian McAuley

Diffusion models have achieved impressive performance in video generation, but their iterative denoising process remains computationally expensive due to the large number of tokens processed at each timestep. Recently, progressive…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Shikang Zheng , Jingkai Huang , Jiacheng Liu , Guantao Chen , Lixuan , Yuqi Lin , Peiliang Cai , Linfeng Zhang

Retrieval-Augmented Generation (RAG) grounds Large Language Models (LLMs) in external knowledge but often suffers from flat context representations and stateless retrieval, leading to unstable performance. We propose Stateful…

Computation and Language · Computer Science 2026-04-17 Qi Dong , Ziheng Lin , Ning Ding

Visual reasoning (VR), which is crucial in many fields for enabling human-like visual understanding, remains highly challenging. Recently, compositional visual reasoning approaches, which leverage the reasoning abilities of large language…

Computer Vision and Pattern Recognition · Computer Science 2026-02-04 Fucai Ke , Vijay Kumar B G , Xingjian Leng , Zhixi Cai , Zaid Khan , Weiqing Wang , Pari Delir Haghighi , Hamid Rezatofighi , Manmohan Chandraker

Text-video retrieval, a prominent sub-field within the domain of multimodal information retrieval, has witnessed remarkable growth in recent years. However, existing methods assume video scenes are consistent with unbiased descriptions.…

Computer Vision and Pattern Recognition · Computer Science 2024-06-04 Huy Le , Tung Kieu , Anh Nguyen , Ngan Le

Ensuring traffic safety and preventing accidents is a critical goal in daily driving, where the advancement of computer vision technologies can be leveraged to achieve this goal. In this paper, we present a multi-view, multi-scale framework…

Computer Vision and Pattern Recognition · Computer Science 2023-05-17 Yunsheng Ma , Liangqi Yuan , Amr Abdelraouf , Kyungtae Han , Rohit Gupta , Zihao Li , Ziran Wang

Real-time detection of irregularities in visual data is very invaluable and useful in many prospective applications including surveillance, patient monitoring systems, etc. With the surge of deep learning methods in the recent years,…

Computer Vision and Pattern Recognition · Computer Science 2018-07-19 Mohammad Sabokrou , Masoud Pourreza , Mohsen Fayyaz , Rahim Entezari , Mahmood Fathy , Jürgen Gall , Ehsan Adeli

The reasoning abilities of large language models (LLMs) have been substantially improved by reinforcement learning with verifiable rewards (RLVR). At test time, collaborative reasoning through Multi-Agent Debate (MAD) has emerged as a…

Computation and Language · Computer Science 2026-05-19 Chenxi Liu , Yanshuo Chen , Ruibo Chen , Tianyi Xiong , Tong Zheng , Heng Huang

In open-domain dialogue systems, generative approaches have attracted much attention for response generation. However, existing methods are heavily plagued by generating safe responses and unnatural responses. To alleviate these two…

Computation and Language · Computer Science 2019-06-25 Shaobo Cui , Rongzhong Lian , Di Jiang , Yuanfeng Song , Siqi Bao , Yong Jiang

Generative models have emerged as an essential building block for many image synthesis and editing tasks. Recent advances in this field have also enabled high-quality 3D or video content to be generated that exhibits either multi-view or…

Computer Vision and Pattern Recognition · Computer Science 2023-08-10 Sherwin Bahmani , Jeong Joon Park , Despoina Paschalidou , Hao Tang , Gordon Wetzstein , Leonidas Guibas , Luc Van Gool , Radu Timofte

Audio-Visual Question Answering (AVQA) is a challenging multimodal reasoning task requiring intelligent systems to answer natural language queries based on paired audio-video inputs accurately. However, existing AVQA approaches often suffer…

Multimedia · Computer Science 2025-04-03 Jie Ma , Zhitao Gao , Qi Chai , Jun Liu , Pinghui Wang , Jing Tao , Zhou Su

Retrieval-augmented generation (RAG) is key to enhancing large language models (LLMs) to systematically access richer factual knowledge. Yet, using RAG brings intrinsic challenges, as LLMs must deal with potentially conflicting knowledge,…

Computation and Language · Computer Science 2025-04-08 Leonardo Ranaldi , Federico Ranaldi , Fabio Massimo Zanzotto , Barry Haddow , Alexandra Birch

Anomaly recognition plays a vital role in surveillance, transportation, healthcare, and public safety. However, most existing approaches rely solely on visual data, making them unreliable under challenging conditions such as occlusion, low…

Computer Vision and Pattern Recognition · Computer Science 2025-11-12 Amjid Ali , Zulfiqar Ahmad Khan , Altaf Hussain , Muhammad Munsif , Adnan Hussain , Sung Wook Baik

We present RAVEN an adaptive AI agent framework designed for multimodal entity discovery and retrieval in large-scale video collections. Synthesizing information across visual, audio, and textual modalities, RAVEN autonomously processes…

Information Retrieval · Computer Science 2025-04-10 Kevin Dela Rosa

Cross-Video Reasoning (CVR) has emerged as a critical frontier in multimodal intelligence, requiring models to retrieve, align, and aggregate evidence distributed across multiple videos. Current Multimodal Large Language Models (MLLMs)…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Yilun Qiu , Jiahe Wang , Cilin Yan , Jiayin Cai , Xiaolong Jiang , Yao Hu , Chun Yuan

Visual abductive reasoning (VAR) is a challenging task that requires AI systems to infer the most likely explanation for incomplete visual observations. While recent MLLMs develop strong general-purpose multimodal reasoning capabilities,…

Computer Vision and Pattern Recognition · Computer Science 2026-01-07 Boyu Chang , Qi Wang , Xi Guo , Zhixiong Nan , Yazhou Yao , Tianfei Zhou

Recently, image manipulation has achieved rapid growth due to the advancement of sophisticated image editing tools. A recent surge of generated fake imagery and videos using neural networks is DeepFake. DeepFake algorithms can create fake…

Computer Vision and Pattern Recognition · Computer Science 2022-06-01 Naciye Celebi , Qingzhong Liu , Muhammed Karatoprak