English
Related papers

Related papers: A Multi-Grid Implicit Neural Representation for Mu…

200 papers

Reconstructing general dynamic scenes is important for many computer vision and graphics applications. Recent works represent the dynamic scene with neural radiance fields for photorealistic view synthesis, while their surface geometry is…

Computer Vision and Pattern Recognition · Computer Science 2023-03-02 Decai Chen , Haofei Lu , Ingo Feldmann , Oliver Schreer , Peter Eisert

With the emergence of neural radiance fields (NeRFs), view synthesis quality has reached an unprecedented level. Compared to traditional mesh-based assets, this volumetric representation is more powerful in expressing scene geometry but…

Computer Vision and Pattern Recognition · Computer Science 2023-03-29 Yuan-Chen Guo , Yan-Pei Cao , Chen Wang , Yu He , Ying Shan , Xiaohu Qie , Song-Hai Zhang

Generating consecutive descriptions for videos, i.e., Video Captioning, requires taking full advantage of visual representation along with the generation process. Existing video captioning methods focus on making an exploration of…

Computer Vision and Pattern Recognition · Computer Science 2022-11-18 Pengpeng Zeng , Haonan Zhang , Lianli Gao , Xiangpeng Li , Jin Qian , Heng Tao Shen

Human Mesh Recovery (HMR) from an image is a challenging problem because of the inherent ambiguity of the task. Existing HMR methods utilized either temporal information or kinematic relationships to achieve higher accuracy, but there is no…

Computer Vision and Pattern Recognition · Computer Science 2025-07-15 Hanbyel Cho , Jaesung Ahn , Yooshin Cho , Junmo Kim

Implicit Neural Representations (INR) or neural fields have emerged as a popular framework to encode multimedia signals such as images and radiance fields while retaining high-quality. Recently, learnable feature grids proposed by…

Computer Vision and Pattern Recognition · Computer Science 2023-09-28 Sharath Girish , Abhinav Shrivastava , Kamal Gupta

Videos are inherently multimodal. This paper studies the problem of how to fully exploit the abundant multimodal clues for improved video categorization. We introduce a hybrid deep learning framework that integrates useful clues from…

Multimedia · Computer Science 2017-06-15 Yu-Gang Jiang , Zuxuan Wu , Jinhui Tang , Zechao Li , Xiangyang Xue , Shih-Fu Chang

We present MVSNeRF, a novel neural rendering approach that can efficiently reconstruct neural radiance fields for view synthesis. Unlike prior works on neural radiance fields that consider per-scene optimization on densely captured images,…

Computer Vision and Pattern Recognition · Computer Science 2021-08-25 Anpei Chen , Zexiang Xu , Fuqiang Zhao , Xiaoshuai Zhang , Fanbo Xiang , Jingyi Yu , Hao Su

Recent sparse multi-view scene reconstruction advances like DUSt3R and MASt3R no longer require camera calibration and camera pose estimation. However, they only process a pair of views at a time to infer pixel-aligned pointmaps. When…

Computer Vision and Pattern Recognition · Computer Science 2024-12-11 Zhenggang Tang , Yuchen Fan , Dilin Wang , Hongyu Xu , Rakesh Ranjan , Alexander Schwing , Zhicheng Yan

Recent advancements in generative video codec (GVC) typically encode video into a 2D latent grid and employ high-capacity generative decoders for reconstruction. However, this paradigm still leaves two key challenges in fully exploiting…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Zihan Zheng , Zhaoyang Jia , Naifu Xue , Jiahao Li , Bin Li , Zongyu Guo , Xiaoyi Zhang , Zhenghao Chen , Houqiang Li , Yan Lu

With the representation learning capability of the deep learning models, deep embedded multi-view clustering (MVC) achieves impressive performance in many scenarios and has become increasingly popular in recent years. Although great…

Machine Learning · Computer Science 2022-05-10 Zongmo Huang , Yazhou Ren , Xiaorong Pu , Lifang He

This paper addresses the challenge of novel view synthesis for a human performer from a very sparse set of camera views. Some recent works have shown that learning implicit neural representations of 3D scenes achieves remarkable view…

Computer Vision and Pattern Recognition · Computer Science 2021-03-30 Sida Peng , Yuanqing Zhang , Yinghao Xu , Qianqian Wang , Qing Shuai , Hujun Bao , Xiaowei Zhou

Neural implicit 3D representations have emerged as a powerful paradigm for reconstructing surfaces from multi-view images and synthesizing novel views. Unfortunately, existing methods such as DVR or IDR require accurate per-pixel object…

Computer Vision and Pattern Recognition · Computer Science 2021-10-11 Michael Oechsle , Songyou Peng , Andreas Geiger

Clinical routine and retrospective cohorts commonly include multi-parametric Magnetic Resonance Imaging; however, they are mostly acquired in different anisotropic 2D views due to signal-to-noise-ratio and scan-time constraints. Thus…

We propose flexgrid2vec, a novel approach for image representation learning. Existing visual representation methods suffer from several issues, including the need for highly intensive computation, the risk of losing in-depth structural…

Computer Vision and Pattern Recognition · Computer Science 2021-09-30 Ali Hamdi , Du Yong Kim , Flora D. Salim

Recently, the image-wise implicit neural representation of videos, NeRV, has gained popularity for its promising results and swift speed compared to regular pixel-wise implicit representations. However, the redundant parameters within the…

Computer Vision and Pattern Recognition · Computer Science 2022-07-19 Zizhang Li , Mengmeng Wang , Huaijin Pi , Kechun Xu , Jianbiao Mei , Yong Liu

This work, termed MH-LVC, presents a multi-hypothesis temporal prediction scheme that employs long- and short-term reference frames in a conditional residual video coding framework. Recent temporal context mining approaches to conditional…

Image and Video Processing · Electrical Eng. & Systems 2025-10-15 Huu-Tai Phung , Zong-Lin Gao , Yi-Chen Yao , Kuan-Wei Ho , Yi-Hsin Chen , Yu-Hsiang Lin , Alessandro Gnutti , Wen-Hsiao Peng

In this paper, we propose a scene-level inverse rendering framework that uses multi-view images to decompose the scene into geometry, SVBRDF, and 3D spatially-varying lighting. While multi-view images have been widely used for object-level…

Computer Vision and Pattern Recognition · Computer Science 2025-03-07 JunYong Choi , SeokYeong Lee , Haesol Park , Seung-Won Jung , Ig-Jae Kim , Junghyun Cho

Understanding the structure of complex activities in untrimmed videos is a challenging task in the area of action recognition. One problem here is that this task usually requires a large amount of hand-annotated minute- or even hour-long…

Computer Vision and Pattern Recognition · Computer Science 2020-10-01 Rosaura G. VidalMata , Walter J. Scheirer , Anna Kukleva , David Cox , Hilde Kuehne

In recent years, the introduction of Multi-modal Large Language Models (MLLMs) into video understanding tasks has become increasingly prevalent. However, how to effectively integrate temporal information remains a critical research focus.…

Computer Vision and Pattern Recognition · Computer Science 2025-07-22 Xiaoyi Bao , Chenwei Xie , Hao Tang , Tingyu Weng , Xiaofeng Wang , Yun Zheng , Xingang Wang

Implicit neural representations (INRs) have recently emerged as a powerful tool that provides an accurate and resolution-independent encoding of data. Their robustness as general approximators has been shown in a wide variety of data…

Machine Learning · Computer Science 2022-08-12 Elizabeth Fons , Alejandro Sztrajman , Yousef El-laham , Alexandros Iosifidis , Svitlana Vyetrenko