中文
相关论文

相关论文: LucidDream: Controlled Temporally-Consistent DeepD…

200 篇论文

Video Diffusion Models have been developed for video generation, usually integrating text and image conditioning to enhance control over the generated content. Despite the progress, ensuring consistency across frames remains a challenge,…

计算机视觉与模式识别 · 计算机科学 2024-11-12 Tian Xia , Xuweiyi Chen , Sihan Xu

Large Vision-Language Models often generate hallucinated content that is not grounded in its visual inputs. While prior work focuses on mitigating hallucinations, we instead explore leveraging hallucination correction as a training…

计算机视觉与模式识别 · 计算机科学 2025-02-24 Lingjun Zhao , Mingyang Xie , Paola Cascante-Bonilla , Hal Daumé , Kwonjoon Lee

We present a simple and effective deep convolutional neural network (CNN) model for video deblurring. The proposed algorithm mainly consists of optical flow estimation from intermediate latent frames and latent frame restoration steps. It…

计算机视觉与模式识别 · 计算机科学 2020-04-07 Jinshan Pan , Haoran Bai , Jinhui Tang

Despite significant progress in video-language modeling, hallucinations remain a persistent challenge in Video Large Language Models (Vid-LLMs), referring to outputs that appear plausible yet contradict the content of the input video. This…

计算机视觉与模式识别 · 计算机科学 2026-04-15 Yiyang Huang , Yitian Zhang , Yizhou Wang , Mingyuan Zhang , Liang Shi , Huimin Zeng , Yun Fu

Large Vision-Language Models (LVLMs) are susceptible to object hallucinations, an issue in which their generated text contains non-existent objects, greatly limiting their reliability and practicality. Current approaches often rely on the…

计算机视觉与模式识别 · 计算机科学 2024-04-24 Ailin Deng , Zhirui Chen , Bryan Hooi

Large vision-language models (LVLMs) tend to hallucinate, especially when visual inputs are corrupted at test time. We show that such corruptions act as additional distribution shifts, significantly amplifying hallucination rates in…

计算机视觉与模式识别 · 计算机科学 2026-03-30 Mriganka Nath , Anurag Das , Jiahao Xie , Bernt Schiele

Recent Video Large Language Models (Video-LLMs) have demonstrated strong capability in video understanding, yet they still suffer from hallucinations. Existing mitigation methods typically rely on training, input modification, auxiliary…

计算机视觉与模式识别 · 计算机科学 2026-04-15 Zijian Liu , Sihan Cao , Pengcheng Zheng , Kuien Liu , Caiyan Qin , Xiaolin Qin , Jiwei Wei , Chaoning Zhang

Existing deep learning-based low-light enhancement methods are typically trained on limited datasets with single enhancement targets, which restricts their generalization ability and controllability in real-world applications. To overcome…

计算机视觉与模式识别 · 计算机科学 2026-05-27 Yufeng Yang , Jianzhuang Liu , Jisheng Chu , Yuqi Peng , Xianfang Zeng , Jiancheng Huang , Shifeng Chen

Deep image denoisers achieve state-of-the-art results but with a hidden cost. As witnessed in recent literature, these deep networks are capable of overfitting their training distributions, causing inaccurate hallucinations to be added to…

图像与视频处理 · 电气工程与系统科学 2022-01-04 Qiyuan Liang , Florian Cassayre , Haley Owsianko , Majed El Helou , Sabine Süsstrunk

Large Vision-Language Models (LVLMs) with discrete image tokenizers unify multimodal representations by encoding visual inputs into a finite set of tokens. Despite their effectiveness, we find that these models still hallucinate…

计算机视觉与模式识别 · 计算机科学 2025-05-29 Weixing Wang , Zifeng Ding , Jindong Gu , Rui Cao , Christoph Meinel , Gerard de Melo , Haojin Yang

Multimodal large language models (MLLMs) frequently suffer from object hallucinations, yet the visual perceptual mechanism underlying this failure remains poorly understood. In this work, we reveal that hallucinations are strongly…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Quanjiang Li , Zhiming Liu , Wei Luo , Tingjin Luo , Chenping Hou

Recent diffusion models have achieved remarkable success in image relighting, and this success has quickly been extended to video relighting. However, existing methods offer limited explicit control over illumination in the relighted…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Yizuo Peng , Xuelin Chen , Kai Zhang , Xiaodong Cun

Unsupervised object-centric learning from videos is a promising approach to extract structured representations from large, unlabeled collections of videos. To support downstream tasks like autonomous control, these representations must be…

计算机视觉与模式识别 · 计算机科学 2025-03-19 Anna Manasyan , Maximilian Seitzer , Filip Radovic , Georg Martius , Andrii Zadaianchuk

The goal of our work is to generate high-quality novel views from monocular videos of complex and dynamic scenes. Prior methods, such as DynamicNeRF, have shown impressive performance by leveraging time-varying dynamic radiation fields.…

计算机视觉与模式识别 · 计算机科学 2024-07-03 Xingyu Miao , Yang Bai , Haoran Duan , Yawen Huang , Fan Wan , Yang Long , Yefeng Zheng

We introduce a method to generate temporally coherent human animation from a single image, a video, or a random noise. This problem has been formulated as modeling of an auto-regressive generation, i.e., to regress past frames to decode…

计算机视觉与模式识别 · 计算机科学 2024-03-25 Tserendorj Adiya , Jae Shin Yoon , Jungeun Lee , Sanghun Kim , Hwasup Lim

Videos captured by consumer cameras often exhibit temporal variations in color and tone that are caused by camera auto-adjustments like white-balance and exposure. When such videos are sub-sampled to play fast-forward, as in the…

图形学 · 计算机科学 2017-10-02 Xuaner Cecilia Zhang , Joon-Young Lee , Kalyan Sunkavalli , Zhaowen Wang

Recent work has shown that diffusion models can serve as powerful neural rendering engines that can be leveraged for inserting virtual objects into images. However, unlike typical physics-based renderers, these neural rendering engines are…

计算机视觉与模式识别 · 计算机科学 2025-12-12 Frédéric Fortier-Chouinard , Zitian Zhang , Louis-Etienne Messier , Mathieu Garon , Anand Bhattad , Jean-François Lalonde

Tomographic image reconstruction is generally an ill-posed linear inverse problem. Such ill-posed inverse problems are typically regularized using prior knowledge of the sought-after object property. Recently, deep neural networks have been…

图像与视频处理 · 电气工程与系统科学 2021-09-28 Sayantan Bhadra , Varun A. Kelkar , Frank J. Brooks , Mark A. Anastasio

Multimodal large language models (MLLMs) have achieved remarkable progress in video understanding.However, hallucination, where the model generates plausible yet incorrect outputs, persists as a significant and under-addressed challenge in…

计算机视觉与模式识别 · 计算机科学 2025-05-20 Jianfeng Cai , Wengang Zhou , Zongmeng Zhang , Jiale Hong , Nianji Zhan , Houqiang Li

Multimodal Large Language Models (MLLMs) have made significant progress in bridging the gap between visual and language modalities. However, hallucinations in MLLMs, where the generated text does not align with image content, continue to be…

人工智能 · 计算机科学 2024-08-05 Kohou Wang , Xiang Liu , Zhaoxiang Liu , Kai Wang , Shiguo Lian