English
Related papers

Related papers: From Pixels to Views: Learning Angular-Aware and P…

200 papers

Reconstructing 3D objects from a single image is an intriguing but challenging problem. One promising solution is to utilize multi-view (MV) 3D reconstruction to fuse generated MV images into consistent 3D objects. However, the generated…

Computer Vision and Pattern Recognition · Computer Science 2024-01-30 Yizheng Chen , Rengan Xie , Qi Ye , Sen Yang , Zixuan Xie , Tianxiao Chen , Rong Li , Yuchi Huo

Light field cameras have been proved to be powerful tools for 3D reconstruction and virtual reality applications. However, the limited resolution of light field images brings a lot of difficulties for further information display and…

Image and Video Processing · Electrical Eng. & Systems 2020-08-27 Qingyan Sun , Shuo Zhang , Song Chang , Lixi Zhu , Youfang Lin

Deep learning and Convolutional Neural Networks (CNNs) have driven major transformations in diverse research areas. However, their limitations in handling low-frequency information present obstacles in certain tasks like interpreting global…

Computer Vision and Pattern Recognition · Computer Science 2024-03-14 Fuzhi Wu , Jiasong Wu , Youyong Kong , Chunfeng Yang , Guanyu Yang , Huazhong Shu , Guy Carrault , Lotfi Senhadji

Visual Language Models (VLMs) are now increasingly being merged with Large Language Models (LLMs) to enable new capabilities, particularly in terms of improved interactivity and open-ended responsiveness. While these are remarkable…

Ultra-low-field (ULF) MRI promises broader accessibility but suffers from low signal-to-noise ratio (SNR), reduced spatial resolution, and contrasts that deviate from high-field standards. Image-to-image translation can map ULF images to a…

Image and Video Processing · Electrical Eng. & Systems 2025-11-13 Felix F Zimmermann

Large-scale pre-trained Vision-Language Models (VLMs) have demonstrated strong few-shot learning capabilities. However, these methods typically learn holistic representations where an image's domain-invariant structure is implicitly…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Hieu Dinh Trung Pham , Huy Minh Nhat Nguyen , Cuong Tuan Nguyen

Following the successes in the fields of vision and language, self-supervised pretraining via masked autoencoding of 3D point set data, or Masked Point Modeling (MPM), has achieved state-of-the-art accuracy in various downstream tasks.…

Computer Vision and Pattern Recognition · Computer Science 2024-05-24 Takahiko Furuya

Deep learning has transformed computational imaging, but traditional pixel-based representations limit their ability to capture continuous, multiscale details of objects. Here we introduce a novel Local Conditional Neural Fields (LCNF)…

Image and Video Processing · Electrical Eng. & Systems 2023-07-25 Hao Wang , Jiabei Zhu , Yunzhe Li , QianWan Yang , Lei Tian

Light field (LF) cameras can record scenes from multiple perspectives, and thus introduce beneficial angular information for image super-resolution (SR). However, it is challenging to incorporate angular information due to disparities among…

Image and Video Processing · Electrical Eng. & Systems 2020-12-30 Yingqian Wang , Jungang Yang , Longguang Wang , Xinyi Ying , Tianhao Wu , Wei An , Yulan Guo

Neural Radiance Fields (NeRF) has demonstrated remarkable 3D reconstruction capabilities with dense view images. However, its performance significantly deteriorates under sparse view settings. We observe that learning the 3D consistency of…

Computer Vision and Pattern Recognition · Computer Science 2023-05-19 Shoukang Hu , Kaichen Zhou , Kaiyu Li , Longhui Yu , Lanqing Hong , Tianyang Hu , Zhenguo Li , Gim Hee Lee , Ziwei Liu

Light field (LF) imaging, which captures both spatial and angular information of a scene, is undoubtedly beneficial to numerous applications. Although various techniques have been proposed for LF acquisition, achieving both angularly and…

Computer Vision and Pattern Recognition · Computer Science 2022-01-05 Trung-Hieu Tran , Jan Berberich , Sven Simon

Vision-Language Models (VLMs) integrate visual knowledge with the analytical capabilities of Large Language Models (LLMs) through supervised visual instruction tuning, using image-question-answer triplets. However, the potential of VLMs…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Yunlong Deng , Guangyi Chen , Tianpei Gu , Lingjing Kong , Yan Li , Zeyu Tang , Kun Zhang

Lithography is fundamental to integrated circuit fabrication, necessitating large computation overhead. The advancement of machine learning (ML)-based lithography models alleviates the trade-offs between manufacturing process expense and…

Computer Vision and Pattern Recognition · Computer Science 2023-04-11 Guojin Chen , Zehua Pei , Haoyu Yang , Yuzhe Ma , Bei Yu , Martin D. F. Wong

We propose a novel framework for training neural networks which is capable of learning 3D information of non-rigid objects when only 2D annotations are available as ground truths. Recently, there have been some approaches that incorporate…

Computer Vision and Pattern Recognition · Computer Science 2020-07-22 Sungheon Park , Minsik Lee , Nojun Kwak

Lensless cameras replace bulky optics with thin modulation masks, enabling compact imaging systems. However, existing methods rely on an idealized model that assumes a globally shift-invariant point spread function (PSF) and sufficiently…

Image and Video Processing · Electrical Eng. & Systems 2025-12-02 Yu Ren , Xiaoling Zhang , Xu Zhan , Xiangdong Ma , Yunqi Wang , Edmund Y. Lam , Tianjiao Zeng

The impressive performance of Large Language Model (LLM) has prompted researchers to develop Multi-modal LLM (MLLM), which has shown great potential for various multi-modal tasks. However, current MLLM often struggles to effectively address…

Computer Vision and Pattern Recognition · Computer Science 2024-12-24 Yeyuan Wang , Dehong Gao , Bin Li , Rujiao Long , Lei Yi , Xiaoyan Cai , Libin Yang , Jinxia Zhang , Shanqing Yu , Qi Xuan

Image restoration of biological structures in microscopy poses unique challenges for preserving fine textures and sharp edges. While recent GAN-based image restoration formulations have introduced frequency-domain losses for natural images,…

Quantitative Methods · Quantitative Biology 2026-01-30 Xingjian Zhang , Claire Leclech , Louison Blivet-Bailly , Abdul I. Barakat , Elsa D. Angelini

We introduce a novel framework for solving inverse problems using NeRF-style generative models. We are interested in the problem of 3-D scene reconstruction given a single 2-D image and known camera parameters. We show that naively…

Computer Vision and Pattern Recognition · Computer Science 2021-12-17 Giannis Daras , Wen-Sheng Chu , Abhishek Kumar , Dmitry Lagun , Alexandros G. Dimakis

The hardware challenges associated with light-field(LF) imaging has made it difficult for consumers to access its benefits like applications in post-capture focus and aperture control. Learning-based techniques which solve the ill-posed…

Image and Video Processing · Electrical Eng. & Systems 2022-07-22 Shrisudhan Govindarajan , Prasan Shedligeri , Sarah , Kaushik Mitra

Vision-Language Models (VLMs) excel at many multimodal tasks, yet they frequently struggle with tasks requiring precise understanding and handling of fine-grained visual elements. This is mainly due to information loss during image encoding…

Computer Vision and Pattern Recognition · Computer Science 2025-10-03 Xuchen Li , Xuzhao Li , Jiahui Gao , Renjie Pi , Shiyu Hu , Wentao Zhang