中文
相关论文

相关论文: ShowMak3r: Compositional TV Show Reconstruction

200 篇论文

Multi-view transformers such as DUSt3R are revolutionizing 3D vision by solving 3D tasks in a feed-forward manner. However, contrary to previous optimization-based pipelines, the inner mechanisms of multi-view transformers are unclear.…

计算机视觉与模式识别 · 计算机科学 2025-10-30 Michal Stary , Julien Gaubil , Ayush Tewari , Vincent Sitzmann

We present a 3D-aware one-shot head reenactment method based on a fully volumetric neural disentanglement framework for source appearance and driver expressions. Our method is real-time and produces high-fidelity and view-consistent output,…

计算机视觉与模式识别 · 计算机科学 2023-12-11 Phong Tran , Egor Zakharov , Long-Nhat Ho , Anh Tuan Tran , Liwen Hu , Hao Li

Reconstructing dynamic articulated objects from a singular monocular video is challenging, requiring joint estimation of shape, motion, and camera parameters from limited views. Current methods typically demand extensive computational…

计算机视觉与模式识别 · 计算机科学 2024-05-22 Hao Zhang , Fang Li , Samyak Rawlekar , Narendra Ahuja

The recent paradigm shift in 3D vision led to the rise of foundation models with remarkable capabilities in 3D perception from uncalibrated images. However, extending these models to large-scale RGB stream 3D reconstruction remains…

计算机视觉与模式识别 · 计算机科学 2026-02-05 Leonid Antsfeld , Boris Chidlovskii , Yohann Cabon , Vincent Leroy , Jerome Revaud

Animating realistic character interactions with the surrounding environment is important for autonomous agents in gaming, AR/VR, and robotics. However, current methods for human motion reconstruction struggle with accurately placing humans…

计算机视觉与模式识别 · 计算机科学 2025-10-20 Joshua Li , Brendan Chharawala , Chang Shu , Xue Bin Peng , Pengcheng Xi

In this paper, we rethink the problem of scene reconstruction from an embodied agent's perspective: While the classic view focuses on the reconstruction accuracy, our new perspective emphasizes the underlying functions and constraints such…

机器人学 · 计算机科学 2021-03-31 Muzhi Han , Zeyu Zhang , Ziyuan Jiao , Xu Xie , Yixin Zhu , Song-Chun Zhu , Hangxin Liu

Much progress has been made in reconstructing garments from an image or a video. However, none of existing works meet the expectations of digitizing high-quality animatable dynamic garments that can be adjusted to various unseen poses. In…

计算机视觉与模式识别 · 计算机科学 2023-11-03 Xiongzheng Li , Jinsong Zhang , Yu-Kun Lai , Jingyu Yang , Kun Li

Active-stereo-based 3D shape measurement is crucial for various purposes, such as industrial inspection, reverse engineering, and medical systems, due to its strong ability to accurately acquire the shape of textureless objects. Active…

计算机视觉与模式识别 · 计算机科学 2025-06-03 Ryo Furukawa , Kota Nishihara , Hiroshi Kawasaki

Humans exhibit an innate capacity to rapidly perceive and segment objects from video observations, and even mentally assemble them into structured 3D scenes. Replicating such capability, termed compositional 3D reconstruction, is pivotal…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Mingyu Dong , Chong Xia , Mingyuan Jia , Weichen Lyu , Long Xu , Zheng Zhu , Yueqi Duan

Despite the impressive progress of telepresence systems for room-scale scenes with static and dynamic scene entities, expanding their capabilities to scenarios with larger dynamic environments beyond a fixed size of a few square-meters…

计算机视觉与模式识别 · 计算机科学 2024-02-14 Leif Van Holland , Patrick Stotko , Stefan Krumpen , Reinhard Klein , Michael Weinmann

Simultaneous localization and mapping (SLAM) are crucial for autonomous robots (e.g., self-driving cars, autonomous drones), 3D mapping systems, and AR/VR applications. This work proposed a novel LiDAR-inertial-visual fusion framework…

计算机视觉与模式识别 · 计算机科学 2022-09-09 Jiarong Lin , Fu Zhang

Tracking surgical modifications based on endoscopic videos is technically feasible and of great clinical advantages; however, it still remains challenging. This report presents a modular pipeline to divide and conquer the clinical…

图像与视频处理 · 电气工程与系统科学 2025-01-06 Yicheng Zhu

Reconstructing general dynamic scenes is important for many computer vision and graphics applications. Recent works represent the dynamic scene with neural radiance fields for photorealistic view synthesis, while their surface geometry is…

计算机视觉与模式识别 · 计算机科学 2023-03-02 Decai Chen , Haofei Lu , Ingo Feldmann , Oliver Schreer , Peter Eisert

We explore the task of embodied view synthesis from monocular videos of deformable scenes. Given a minute-long RGBD video of people interacting with their pets, we render the scene from novel camera trajectories derived from the in-scene…

计算机视觉与模式识别 · 计算机科学 2023-10-03 Chonghyuk Song , Gengshan Yang , Kangle Deng , Jun-Yan Zhu , Deva Ramanan

In this paper, we present frame reconstruction model: FrameRS. It consists self-supervised video frame reconstructor and key frame selector. The frame reconstructor, FrameMAE, is developed by adapting the principles of the Masked…

计算机视觉与模式识别 · 计算机科学 2023-09-19 Qiqian Fu , Guanhong Wang , Gaoang Wang

In this paper, we present a framework to represent mock 3D objects and scenes, which are not 3D but appear 3D. In our framework, each mock-3D object is represented using 2D non-conservative vector fields and thickness information that are…

图形学 · 计算机科学 2024-01-02 Ergun Akleman , Youyou Wang , Ozgur Gonen

This paper primarily investigates the task of expression-only portrait video performance editing based on a driving video, which plays a crucial role in animation and film industries. Most existing research mainly focuses on portrait…

计算机视觉与模式识别 · 计算机科学 2026-03-23 Jiadong Liang , Bojun Xiong , Jie Tian , Hua Li , Xiao Long , Yong Zheng , Huan Fu

Lip-syncing videos with given audio is the foundation for various applications including the creation of virtual presenters or performers. While recent studies explore high-fidelity lip-sync with different techniques, their task-orientated…

Light detection and ranging (Lidar) data can be used to capture the depth and intensity profile of a 3D scene. This modality relies on constructing, for each pixel, a histogram of time delays between emitted light pulses and detected photon…

图像与视频处理 · 电气工程与系统科学 2019-09-13 Julián Tachella , Yoann Altmann , Ximing Ren , Aongus McCarthy , Gerald S. Buller , Jean-Yves Tourneret , Steve McLaughlin

Reconstructing people, objects, and their interactions in 3D is a long-standing goal for intelligent systems. Often the input is RGB video from a moving camera, making the task ill-posed; depth is ambiguous, humans and objects occlude each…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Lixin Xue , Chengwei Zheng , Georgios Paschalidis , Chen Guo , Manuel Kaufmann , Juan Zarate , Dimitrios Tzionas