English
Related papers

Related papers: VistaDPO: Video Hierarchical Spatial-Temporal Dire…

200 papers

Large Visual Language Models (LVLMs) have demonstrated impressive capabilities across multiple tasks. However, their trustworthiness is often challenged by hallucinations, which can be attributed to the modality misalignment and the…

Computer Vision and Pattern Recognition · Computer Science 2025-09-23 Jiulong Wu , Zhengliang Shi , Shuaiqiang Wang , Jizhou Huang , Dawei Yin , Lingyong Yan , Min Cao , Min Zhang

Large Vision-Language Models (LVLMs) or multimodal large language models represent a significant advancement in artificial intelligence, enabling systems to understand and generate content across both visual and textual modalities. While…

Machine Learning · Computer Science 2025-09-09 Thanh Thi Nguyen , Campbell Wilson , Janis Dalins

Video diffusion models (VDMs) have demonstrated remarkable capabilities in text-to-video (T2V) generation. Despite their success, VDMs still suffer from degraded image quality and flickering artifacts. To address these issues, some…

Computer Vision and Pattern Recognition · Computer Science 2026-01-01 Jiacheng Zhang , Jie Wu , Weifeng Chen , Yatai Ji , Xuefeng Xiao , Weilin Huang , Kai Han

Hallucination remains a fundamental challenge in vision-language models (VLMs), where autoregressive generation may produce linguistically plausible yet physically inconsistent or visually ungrounded responses due to likelihood maximization…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Qinwu Xu

The rapid growth of autonomous driving datasets has enabled the scaling of powerful motion forecasting models. While large-scale pretraining provides strong performance, the standard imitation objective may not fully capture the complex…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Zhefan Xu , Ghassen Jerfel , Marina Haliem , Qi Zhao , Jeonhyung Kang , Khaled S. Refaat

Visual preference alignment involves training Large Vision-Language Models (LVLMs) to predict human preferences between visual inputs. This is typically achieved by using labeled datasets of chosen/rejected pairs and employing optimization…

Computer Vision and Pattern Recognition · Computer Science 2024-10-24 Ziyu Liu , Yuhang Zang , Xiaoyi Dong , Pan Zhang , Yuhang Cao , Haodong Duan , Conghui He , Yuanjun Xiong , Dahua Lin , Jiaqi Wang

Multimodal Large Language Models (MLLMs) still struggle with hallucinations despite their impressive capabilities. Recent studies have attempted to mitigate this by applying Direct Preference Optimization (DPO) to multimodal scenarios using…

Computation and Language · Computer Science 2025-01-29 Jinlan Fu , Shenzhen Huangfu , Hao Fei , Xiaoyu Shen , Bryan Hooi , Xipeng Qiu , See-Kiong Ng

Recent progress in generative diffusion models has greatly advanced text-to-video generation. While text-to-video models trained on large-scale, diverse datasets can produce varied outputs, these generations often deviate from user…

Computer Vision and Pattern Recognition · Computer Science 2024-12-19 Runtao Liu , Haoyu Wu , Zheng Ziqiang , Chen Wei , Yingqing He , Renjie Pi , Qifeng Chen

Vision generation remains a challenging frontier in artificial intelligence, requiring seamless integration of visual understanding and generative capabilities. In this paper, we propose a novel framework, Vision-Driven Prompt Optimization…

Computer Vision and Pattern Recognition · Computer Science 2025-01-07 Leo Franklin , Apiradee Boonmee , Kritsada Wongsuwan

Large Vision-Language Models (LVLMs) frequently suffer from hallucinations. Existing preference learning-based approaches largely rely on proprietary models to construct preference datasets. We identify that this reliance introduces a…

Artificial Intelligence · Computer Science 2026-04-28 Byeonggeuk Lim , JungMin Yun , Junehyoung Kwon , Kyeonghyun Kim , YoungBin Kim

In this paper, we propose VidLA, an approach for video-language alignment at scale. There are two major limitations of previous video-language alignment approaches. First, they do not capture both short-range and long-range temporal…

Computer Vision and Pattern Recognition · Computer Science 2024-03-25 Mamshad Nayeem Rizve , Fan Fei , Jayakrishnan Unnikrishnan , Son Tran , Benjamin Z. Yao , Belinda Zeng , Mubarak Shah , Trishul Chilimbi

Video-language models (VLMs) achieve strong multimodal understanding but remain prone to hallucinations, especially when reasoning about actions and temporal order. Existing mitigation strategies, such as textual filtering or random video…

Computer Vision and Pattern Recognition · Computer Science 2026-01-09 Tobia Poppi , Burak Uzkent , Amanmeet Garg , Lucas Porto , Garin Kessler , Yezhou Yang , Marcella Cornia , Lorenzo Baraldi , Rita Cucchiara , Florian Schiffers

Multimodal Large Language Models (MLLMs) emerge as a unified interface to address a multitude of tasks, ranging from NLP to computer vision. Despite showcasing state-of-the-art results in many benchmarks, a long-standing issue is the…

Computer Vision and Pattern Recognition · Computer Science 2025-08-29 Alberto Compagnoni , Davide Caffagni , Nicholas Moratelli , Lorenzo Baraldi , Marcella Cornia , Rita Cucchiara

Preference alignment has become a crucial component in enhancing the performance of Large Language Models (LLMs), yet its impact in Multimodal Large Language Models (MLLMs) remains comparatively underexplored. Similar to language models,…

Computer Vision and Pattern Recognition · Computer Science 2024-07-03 Elmira Amirloo , Jean-Philippe Fauconnier , Christoph Roesmann , Christian Kerl , Rinu Boney , Yusu Qian , Zirui Wang , Afshin Dehghan , Yinfei Yang , Zhe Gan , Peter Grasch

Multi-modal large language models (MLLMs) are expected to support multi-turn queries of interchanging image and text modalities in production. However, the current MLLMs trained with visual-question-answering (VQA) datasets could suffer…

Computation and Language · Computer Science 2024-11-06 Shengzhi Li , Rongyu Lin , Shichao Pei

Vision Large Language Models (VLLMs) are widely acknowledged to be prone to hallucinations. Existing research addressing this problem has primarily been confined to image inputs, with limited exploration of video-based hallucinations.…

Computer Vision and Pattern Recognition · Computer Science 2026-04-24 Wey Yeh Choong , Yangyang Guo , Mohan Kankanhalli

Multimodal large language models (MLLMs) have recently shown significant advancements in video understanding, excelling in content reasoning and instruction-following tasks. However, hallucination, where models generate inaccurate or…

Computer Vision and Pattern Recognition · Computer Science 2025-04-02 Chaoyu Li , Eun Woo Im , Pooyan Fazli

As language models continue to scale, Large Language Models (LLMs) have exhibited emerging capabilities in In-Context Learning (ICL), enabling them to solve language tasks by prefixing a few in-context demonstrations (ICDs) as context.…

Computer Vision and Pattern Recognition · Computer Science 2024-11-26 Hongrui Jia , Chaoya Jiang , Haiyang Xu , Wei Ye , Mengfan Dong , Ming Yan , Ji Zhang , Fei Huang , Shikun Zhang

Fine-grained video captioning aims to generate detailed, temporally coherent descriptions of video content. However, existing methods struggle to capture subtle video dynamics and rich detailed information. In this paper, we leverage…

Artificial Intelligence · Computer Science 2026-03-24 Jisheng Dang , Yizhou Zhang , Hao Ye , Teng Wang , Siming Chen , Huicheng Zheng , Yulan Guo , Jianhuang Lai , Bin Hu

Multimodal Large Language Models (MLLMs) have significantly improved the performance of various tasks, but continue to suffer from visual hallucinations, a critical issue where generated responses contradict visual evidence. While Direct…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Yuanshuai Li , Yuping Yan , Junfeng Tang , Yunxuan Li , Zeqi Zheng , Yaochu Jin