English
Related papers

Related papers: RASLF: Representation-Aware State Space Model for …

200 papers

Geometry problem solving (GPS) is a challenging mathematical reasoning task requiring multi-modal understanding, fusion, and reasoning. Existing neural solvers take GPS as a vision-language task but are short in the representation of…

Computer Vision and Pattern Recognition · Computer Science 2024-02-21 Zhong-Zhi Li , Ming-Liang Zhang , Fei Yin , Cheng-Lin Liu

We present a method for transferring the artistic features of an arbitrary style image to a 3D scene. Previous methods that perform 3D stylization on point clouds or meshes are sensitive to geometric reconstruction errors for complex…

Computer Vision and Pattern Recognition · Computer Science 2022-06-14 Kai Zhang , Nick Kolkin , Sai Bi , Fujun Luan , Zexiang Xu , Eli Shechtman , Noah Snavely

Recent years have seen significant developments in the field of License Plate Recognition (LPR) through the integration of deep learning techniques and the increasing availability of training data. Nevertheless, reconstructing license…

Computer Vision and Pattern Recognition · Computer Science 2023-05-30 Valfride Nascimento , Rayson Laroca , Jorge de A. Lambert , William Robson Schwartz , David Menotti

We introduce NeRF-GS, a novel framework that jointly optimizes Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS). This framework leverages the inherent continuous spatial representation of NeRF to mitigate several limitations…

Computer Vision and Pattern Recognition · Computer Science 2025-08-01 Shuangkang Fang , I-Chao Shen , Takeo Igarashi , Yufeng Wang , ZeSheng Wang , Yi Yang , Wenrui Ding , Shuchang Zhou

Current hyperspectral anomaly detection (HAD) benchmark datasets suffer from low resolution, simple background, and small size of the detection data. These factors also limit the performance of the well-known low-rank representation (LRR)…

Image and Video Processing · Electrical Eng. & Systems 2024-02-26 Chenyu Li , Bing Zhang , Danfeng Hong , Jing Yao , Jocelyn Chanussot

Scalable Vector Graphics (SVG) offer a powerful format for representing visual designs as interpretable code. Recent advances in vision-language models (VLMs) have enabled high-quality SVG generation by framing the problem as a code…

Shadow removal under diverse lighting conditions requires disentangling illumination from intrinsic reflectance, a challenge compounded when physical priors are not properly aligned. We propose PhaSR (Physically Aligned Shadow Removal),…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Chia-Ming Lee , Yu-Fan Lin , Yu-Jou Hsiao , Jin-Hui Jiang , Yu-Lun Liu , Chih-Chung Hsu

Recently, most of state-of-the-art single image super-resolution (SISR) methods have attained impressive performance by using deep convolutional neural networks (DCNNs). The existing SR methods have limited performance due to a fixed…

Image and Video Processing · Electrical Eng. & Systems 2021-07-08 Rao Muhammad Umer , Asad Munir , Christian Micheloni

Neural Radiance Fields (NeRFs) have proven to be powerful 3D representations, capable of high quality novel view synthesis of complex scenes. While NeRFs have been applied to graphics, vision, and robotics, problems with slow rendering…

Computer Vision and Pattern Recognition · Computer Science 2023-10-30 Tristan Aumentado-Armstrong , Ashkan Mirzaei , Marcus A. Brubaker , Jonathan Kelly , Alex Levinshtein , Konstantinos G. Derpanis , Igor Gilitschenski

In the last few years, the fusion of multi-modal data has been widely studied for various applications such as robotics, gesture recognition, and autonomous navigation. Indeed, high-quality visual sensors are expensive, and consumer-grade…

Image and Video Processing · Electrical Eng. & Systems 2024-11-13 Aditya Kasliwal , Ishaan Gakhar , Aryan Kamani , Pratinav Seth , Ujjwal Verma

Medical image segmentation remains challenging due to intensity inhomogeneity, noise, blurred boundaries, and irregular structures. Traditional level set methods, while effective in certain cases, often depend on approximate bias field…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Wenqi Zhao , Jiacheng Sang , Fenghua Cheng , Yonglu Shu , Dong Li , Xiaofeng Yang

Latent 3D reconstruction has shown great promise in empowering 3D semantic understanding and 3D generation by distilling 2D features into the 3D space. However, existing approaches struggle with the domain gap between 2D feature space and…

Computer Vision and Pattern Recognition · Computer Science 2025-02-14 Chaoyi Zhou , Xi Liu , Feng Luo , Siyu Huang

We propose a novel approach for 3D mesh reconstruction from multi-view images. Our method takes inspiration from large reconstruction models like LRM that use a transformer-based triplane generator and a Neural Radiance Field (NeRF) model…

Computer Vision and Pattern Recognition · Computer Science 2025-08-08 Peiye Zhuang , Songfang Han , Chaoyang Wang , Aliaksandr Siarohin , Jiaxu Zou , Michael Vasilkovsky , Vladislav Shakhrai , Sergey Korolev , Sergey Tulyakov , Hsin-Ying Lee

Various SDF-based neural implicit surface reconstruction methods have been proposed recently, and have demonstrated remarkable modeling capabilities. However, due to the global nature and limited representation ability of a single network,…

Computer Vision and Pattern Recognition · Computer Science 2025-01-17 Leyuan Yang , Bailin Deng , Juyong Zhang

Complete reconstruction of surgical scenes is crucial for robot-assisted surgery (RAS). Deep depth estimation is promising but existing works struggle with depth discontinuities, resulting in noisy predictions at object boundaries and do…

Computer Vision and Pattern Recognition · Computer Science 2025-03-31 Xu Wang , Shuai Zhang , Baoru Huang , Danail Stoyanov , Evangelos B. Mazomenos

Deep learning based approaches has achieved great performance in single image super-resolution (SISR). However, recent advances in efficient super-resolution focus on reducing the number of parameters and FLOPs, and they aggregate more…

Computer Vision and Pattern Recognition · Computer Science 2022-05-17 Fangyuan Kong , Mingxi Li , Songwei Liu , Ding Liu , Jingwen He , Yang Bai , Fangmin Chen , Lean Fu

Hyperspectral image fusion aims to reconstruct high-spatial-resolution hyperspectral images (HR-HSI) by integrating complementary information from multi-source inputs. Despite recent progress, existing methods still face two critical…

Computer Vision and Pattern Recognition · Computer Science 2026-04-08 Qiya Song , Hongzhi Zhou , Lishan Tan , Renwei Dian , Shutao Li

Latent scene representation plays a significant role in training reinforcement learning (RL) agents. To obtain good latent vectors describing the scenes, recent works incorporate the 3D-aware latent-conditioned NeRF pipeline into scene…

Robotics · Computer Science 2024-09-30 Jiaxu Wang , Ziyi Zhang , Qiang Zhang , Jia Li , Jingkai Sun , Mingyuan Sun , Junhao He , Renjing Xu

Retrieval-augmented generation (RAG) aims to mitigate the hallucination of Large Language Models (LLMs) by retrieving and incorporating relevant external knowledge into the generation process. However, the external knowledge may contain…

Computation and Language · Computer Science 2026-01-12 Yi Sui , Chaozhuo Li , Chen Zhang , Dawei song , Qiuchi Li

Recent years have witnessed significant advancements in light field image super-resolution (LFSR) owing to the progress of modern neural networks. However, these methods often face challenges in capturing long-range dependencies (CNN-based)…

Computer Vision and Pattern Recognition · Computer Science 2024-06-19 Wang xia , Yao Lu , Shunzhou Wang , Ziqi Wang , Peiqi Xia , Tianfei Zhou