English
Related papers

Related papers: EasyVolcap: Accelerating Neural Volumetric Video R…

200 papers

Generating 4D scenes from a single-view video is inherently ill-posed: a single viewpoint lacks the information needed to recover a complete, dynamic scene with full coverage. Existing methods are typically limited to monocular videos,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Tingxi Chen , Ke Hao , Yabo Chen , Zhengxue Cheng , Rong Xie , Li Song , Haibin Huang , Chi Zhang , Xuelong Li

Traditional volume visualization (VolVis) methods, like direct volume rendering, suffer from rigid transfer function designs and high computational costs. Although novel view synthesis approaches enhance rendering efficiency, they require…

Human-Computer Interaction · Computer Science 2025-07-21 Kuangshi Ai , Kaiyuan Tang , Chaoli Wang

We present a real-time deep learning framework for video-based facial performance capture -- the dense 3D tracking of an actor's face given a monocular video. Our pipeline begins with accurately capturing a subject using a high-end…

Computer Vision and Pattern Recognition · Computer Science 2017-06-05 Samuli Laine , Tero Karras , Timo Aila , Antti Herva , Shunsuke Saito , Ronald Yu , Hao Li , Jaakko Lehtinen

Virtual reality (VR) is increasingly used to enhance the ecological validity of motor control and learning studies by providing immersive, interactive environments with precise motion tracking. However, designing realistic VR-based motor…

Quantitative Methods · Quantitative Biology 2025-05-01 Cristina Rossi , Rini Varghese , Amy J Bastian

Large language models (LLMs)-based image captioning has the capability of describing objects not explicitly observed in training data; yet novel objects occur frequently, necessitating the requirement of sustaining up-to-date object…

Computer Vision and Pattern Recognition · Computer Science 2024-04-09 Jiaxuan Li , Duc Minh Vo , Akihiro Sugimoto , Hideki Nakayama

In this work, we pioneer Semantic Flow, a neural semantic representation of dynamic scenes from monocular videos. In contrast to previous NeRF methods that reconstruct dynamic scenes from the colors and volume densities of individual…

Computer Vision and Pattern Recognition · Computer Science 2024-04-09 Fengrui Tian , Yueqi Duan , Angtian Wang , Jianfei Guo , Shaoyi Du

We introduce VividDream, a method for generating explorable 4D scenes with ambient dynamics from a single input image or text prompt. VividDream first expands an input image into a static 3D point cloud through iterative inpainting and…

Computer Vision and Pattern Recognition · Computer Science 2024-05-31 Yao-Chih Lee , Yi-Ting Chen , Andrew Wang , Ting-Hsuan Liao , Brandon Y. Feng , Jia-Bin Huang

We present an algorithm for generating novel views at arbitrary viewpoints and any input time step given a monocular video of a dynamic scene. Our work builds upon recent advances in neural implicit representation and uses continuous and…

Computer Vision and Pattern Recognition · Computer Science 2021-05-14 Chen Gao , Ayush Saraf , Johannes Kopf , Jia-Bin Huang

As virtual reality gains popularity, the demand for controllable creation of immersive and dynamic omnidirectional videos (ODVs) is increasing. While previous text-to-ODV generation methods achieve impressive results, they struggle with…

Computer Vision and Pattern Recognition · Computer Science 2024-12-13 Weiqi Li , Shijie Zhao , Chong Mou , Xuhan Sheng , Zhenyu Zhang , Qian Wang , Junlin Li , Li Zhang , Jian Zhang

Event camera sensors are bio-inspired sensors which asynchronously capture per-pixel brightness changes and output a stream of events encoding the polarity, location and time of these changes. These systems are witnessing rapid advancements…

Computer Vision and Pattern Recognition · Computer Science 2025-09-15 Aupendu Kar , Vishnu Raj , Guan-Ming Su

Visual data comes in various forms, ranging from small icons of just a few pixels to long videos spanning hours. Existing multi-modal LLMs usually standardize these diverse visual inputs to a fixed resolution for visual encoders and yield…

Computer Vision and Pattern Recognition · Computer Science 2025-02-28 Zuyan Liu , Yuhao Dong , Ziwei Liu , Winston Hu , Jiwen Lu , Yongming Rao

EasyVis2 is a system designed to provide hands-free, real-time 3D visualization for laparoscopic surgery. It incorporates a surgical trocar equipped with an array of micro-cameras, which can be inserted into the body cavity to offer an…

Computer Vision and Pattern Recognition · Computer Science 2025-04-10 Yung-Hong Sun , Gefei Shen , Jiangang Chen , Jayer Fernandes , Amber L. Shada , Charles P. Heise , Hongrui Jiang , Yu Hen Hu

Humans possess a remarkable ability to mentally explore and replay 3D environments they have previously experienced. Inspired by this mental process, we present EvoWorld: a world model that bridges panoramic video generation with evolving…

Computer Vision and Pattern Recognition · Computer Science 2025-10-02 Jiahao Wang , Luoxin Ye , TaiMing Lu , Junfei Xiao , Jiahan Zhang , Yuxiang Guo , Xijun Liu , Rama Chellappa , Cheng Peng , Alan Yuille , Jieneng Chen

Modern visual agents require representations that are general, causal, and physically structured to operate in real-time streaming environments. However, current vision foundation models remain fragmented, specializing narrowly in image…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Yibin Yan , Jilan Xu , Shangzhe Di , Haoning Wu , Weidi Xie

Video editing is a creative and complex endeavor and we believe that there is potential for reimagining a new video editing interface to better support the creative and exploratory nature of video editing. We take inspiration from latent…

Human-Computer Interaction · Computer Science 2024-06-26 David Chuan-En Lin , Fabian Caba Heilbron , Joon-Young Lee , Oliver Wang , Nikolas Martelaro

We propose a novel neural rendering pipeline, Hybrid Volumetric-Textural Rendering (HVTR), which synthesizes virtual human avatars from arbitrary poses efficiently and at high quality. First, we learn to encode articulated human motions on…

Computer Vision and Pattern Recognition · Computer Science 2022-09-02 Tao Hu , Tao Yu , Zerong Zheng , He Zhang , Yebin Liu , Matthias Zwicker

Videos typically record the streaming and continuous visual data as discrete consecutive frames. Since the storage cost is expensive for videos of high fidelity, most of them are stored in a relatively low resolution and frame rate. Recent…

Image and Video Processing · Electrical Eng. & Systems 2022-06-10 Zeyuan Chen , Yinbo Chen , Jingwen Liu , Xingqian Xu , Vidit Goel , Zhangyang Wang , Humphrey Shi , Xiaolong Wang

Video-text retrieval (VTR) aims to locate relevant videos using natural language queries. Current methods, often based on pre-trained models like CLIP, are hindered by video's inherent redundancy and their reliance on coarse, final-layer…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Zequn Xie , Boyun Zhang , Yuxiao Lin , Tao Jin

Reconstructing dynamic 3D garment surfaces with open boundaries from monocular videos is an important problem as it provides a practical and low-cost solution for clothes digitization. Recent neural rendering methods achieve high-quality…

Computer Vision and Pattern Recognition · Computer Science 2023-05-30 Lingteng Qiu , Guanying Chen , Jiapeng Zhou , Mutian Xu , Junle Wang , Xiaoguang Han

Volumetric video has emerged as a key medium for immersive telepresence and augmented/virtual reality, enabling six-degrees-of-freedom (6DoF) navigation and realistic spatial interactions. However, delivering high-quality dynamic volumetric…

Computer Vision and Pattern Recognition · Computer Science 2025-09-23 Houqiang Zhong , Zihan Zheng , Qiang Hu , Yuan Tian , Ning Cao , Lan Xu , Xiaoyun Zhang , Zhengxue Cheng , Li Song , Wenjun Zhang