English
Related papers

Related papers: Efficient High-Resolution Image Editing with Hallu…

200 papers

Ptychography is an enabling coherent diffraction imaging technique for both fundamental and applied sciences. Its applications in optical microscopy, however, fall short for its low imaging throughput and limited resolution. Here, we report…

Contrastive pretraining of image-text foundation models, such as CLIP, demonstrated excellent zero-shot performance and improved robustness on a wide range of downstream tasks. However, these models utilize large transformer-based encoders…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Pavan Kumar Anasosalu Vasu , Hadi Pouransari , Fartash Faghri , Raviteja Vemulapalli , Oncel Tuzel

Multimodal Large Language Models (MLLMs) have experienced significant advancements recently. Nevertheless, challenges persist in the accurate recognition and comprehension of intricate details within high-resolution images. Despite being…

Computer Vision and Pattern Recognition · Computer Science 2024-03-05 Haogeng Liu , Quanzeng You , Xiaotian Han , Yiqi Wang , Bohan Zhai , Yongfei Liu , Yunzhe Tao , Huaibo Huang , Ran He , Hongxia Yang

We present Pippo, a generative model capable of producing 1K resolution dense turnaround videos of a person from a single casually clicked photo. Pippo is a multi-view diffusion transformer and does not require any additional inputs - e.g.,…

Computer Vision and Pattern Recognition · Computer Science 2025-02-12 Yash Kant , Ethan Weber , Jin Kyu Kim , Rawal Khirodkar , Su Zhaoen , Julieta Martinez , Igor Gilitschenski , Shunsuke Saito , Timur Bagautdinov

Visual hallucinations in Large Language Models (LLMs), where the model generates responses that are inconsistent with the visual input, pose a significant challenge to their reliability, particularly in contexts where precise and…

Computer Vision and Pattern Recognition · Computer Science 2025-06-30 Nokimul Hasan Arif , Shadman Rabby , Md Hefzul Hossain Papon , Sabbir Ahmed

As recent advances in mobile camera technology have enabled the capability to capture high-resolution images, such as 4K images, the demand for an efficient deblurring model handling large motion has increased. In this paper, we discover…

Computer Vision and Pattern Recognition · Computer Science 2024-04-19 Insoo Kim , Jae Seok Choi , Geonseok Seo , Kinam Kwon , Jinwoo Shin , Hyong-Euk Lee

Image Phase Alignment Super-Sampling (ImPASS) is a computational imaging algorithm for converting a sequence of displaced low-resolution images into a single high-resolution image. The method consists of a unique combination of Phase…

Optics · Physics 2025-09-01 James N. Caron

Spike cameras, as innovative neuromorphic devices, generate continuous spike streams to capture high-speed scenes with lower bandwidth and higher dynamic range than traditional RGB cameras. However, reconstructing high-quality images from…

Computer Vision and Pattern Recognition · Computer Science 2025-03-05 Kang Chen , Yajing Zheng , Tiejun Huang , Zhaofei Yu

While image generation with diffusion models has achieved a great success, generating images of higher resolution than the training size remains a challenging task due to the high computational cost. Current methods typically perform the…

Computer Vision and Pattern Recognition · Computer Science 2025-02-28 Zhengqiang Zhang , Ruihuang Li , Lei Zhang

Recent advances in camera designs and imaging pipelines allow us to capture high-quality images using smartphones. However, due to the small size and lens limitations of the smartphone cameras, we commonly find artifacts or degradation in…

Computer Vision and Pattern Recognition · Computer Science 2023-11-27 Marcos V. Conde , Florin Vasluianu , Javier Vazquez-Corral , Radu Timofte

In this work we apply commonly known methods of non-adaptive interpolation (nearest pixel, bilinear, B-spline, bicubic, Hermite spline) and sampling (point sampling, supersampling, mip-map pre-filtering, rip-map pre-filtering and FAST) to…

Computer Vision and Pattern Recognition · Computer Science 2020-09-15 Anton Trusov , Elena Limonova

In recent years, diffusion models have emerged as the most powerful approach in image synthesis. However, applying these models directly to video synthesis presents challenges, as it often leads to noticeable flickering contents. Although…

Computer Vision and Pattern Recognition · Computer Science 2023-08-11 Zhongjie Duan , Lizhou You , Chengyu Wang , Cen Chen , Ziheng Wu , Weining Qian , Jun Huang

Recently, multimodal large language models have made significant advancements in video understanding tasks. However, their ability to understand unprocessed long videos is very limited, primarily due to the difficulty in supporting the…

Computer Vision and Pattern Recognition · Computer Science 2024-06-18 Yiwei Sun , Zhihang Liu , Chuanbin Liu , Bowei Pu , Zhihan Zhang , Hongtao Xie

This work focuses on model-free zero-shot 6D object pose estimation for robotics applications. While existing methods can estimate the precise 6D pose of objects, they heavily rely on curated CAD models or reference images, the preparation…

Computer Vision and Pattern Recognition · Computer Science 2025-02-18 Yibo Liu , Zhaodong Jiang , Binbin Xu , Guile Wu , Yuan Ren , Tongtong Cao , Bingbing Liu , Rui Heng Yang , Amir Rasouli , Jinjun Shan

Object hallucination in Large Vision-Language Models (LVLMs) significantly hinders their reliable deployment. Existing methods struggle to balance efficiency and accuracy: they often require expensive reference models and multiple forward…

Computer Vision and Pattern Recognition · Computer Science 2026-02-27 Yangguang Lin , Quan Fang , Yufei Li , Jiachen Sun , Junyu Gao , Jitao Sang

High-resolution satellite imagery has proven useful for a broad range of tasks, including measurement of global human population, local economic livelihoods, and biodiversity, among many others. Unfortunately, high-resolution imagery is…

Computer Vision and Pattern Recognition · Computer Science 2022-04-05 Yutong He , Dingjie Wang , Nicholas Lai , William Zhang , Chenlin Meng , Marshall Burke , David B. Lobell , Stefano Ermon

Large Vision-Language Models (LVLMs) have demonstrated impressive multimodal understanding capabilities, yet they remain prone to object hallucination, where models describe non-existent objects or attribute incorrect factual information,…

Computer Vision and Pattern Recognition · Computer Science 2026-02-24 Ahmed Akl , Abdelwahed Khamis , Ali Cheraghian , Zhe Wang , Sara Khalifa , Kewen Wang

Hallucinations in large vision-language models (LVLMs) often stem from the model's sensitivity to image tokens during decoding, as evidenced by attention peaks observed when generating both real and hallucinated entities. To address this,…

Computer Vision and Pattern Recognition · Computer Science 2025-06-24 Shuaiye Lu , Linjiang Zhou , Xiaochuan Shi

Text-to-image diffusion models have recently received increasing interest for their astonishing ability to produce high-fidelity images from solely text inputs. Subsequent research efforts aim to exploit and apply their capabilities to real…

Computer Vision and Pattern Recognition · Computer Science 2024-06-27 Manuel Brack , Felix Friedrich , Katharina Kornmeier , Linoy Tsaban , Patrick Schramowski , Kristian Kersting , Apolinário Passos

Image denoising is one of the most critical problems in mobile photo processing. While many solutions have been proposed for this task, they are usually working with synthetic data and are too computationally expensive to run on mobile…

‹ Prev 1 4 5 6 7 8 10 Next ›