中文
相关论文

相关论文: Anything in Any Scene: Photorealistic Video Object…

200 篇论文

We investigate how to enhance the physical fidelity of video generation models by leveraging synthetic videos derived from computer graphics pipelines. These rendered videos respect real-world physics, such as maintaining 3D consistency,…

图像与视频处理 · 电气工程与系统科学 2025-03-28 Qi Zhao , Xingyu Ni , Ziyu Wang , Feng Cheng , Ziyan Yang , Lu Jiang , Bohan Wang

Effective spatio-temporal representation is fundamental to modeling, understanding, and predicting dynamics in videos. The atomic unit of a video, the pixel, traces a continuous 3D trajectory over time, serving as the primitive element of…

计算机视觉与模式识别 · 计算机科学 2025-10-16 Xinhang Liu , Yuxi Xiao , Donny Y. Chen , Jiashi Feng , Yu-Wing Tai , Chi-Keung Tang , Bingyi Kang

We propose Text2Scene, a method to automatically create realistic textures for virtual scenes composed of multiple objects. Guided by a reference image and text descriptions, our pipeline adds detailed texture on labeled 3D geometries in…

计算机视觉与模式识别 · 计算机科学 2023-09-01 Inwoo Hwang , Hyeonwoo Kim , Young Min Kim

We present UrbanIR (Urban Scene Inverse Rendering), a new inverse graphics model that enables realistic, free-viewpoint renderings of scenes under various lighting conditions with a single video. It accurately infers shape, albedo,…

计算机视觉与模式识别 · 计算机科学 2025-01-16 Chih-Hao Lin , Bohan Liu , Yi-Ting Chen , Kuan-Sheng Chen , David Forsyth , Jia-Bin Huang , Anand Bhattad , Shenlong Wang

In this work, we focus on a challenging task: synthesizing multiple imaginary videos given a single image. Major problems come from high dimensionality of pixel space and the ambiguity of potential motions. To overcome those problems, we…

计算机视觉与模式识别 · 计算机科学 2017-06-16 Baoyang Chen , Wenmin Wang , Jinzhuo Wang , Xiongtao Chen

Representing scenes from multi-view images is a crucial task in computer vision with extensive applications. However, inherent photometric distortions in the camera imaging can significantly degrade image quality. Without accounting for…

计算机视觉与模式识别 · 计算机科学 2025-06-27 Weichen Dai , Kangcheng Ma , Jiaxin Wang , Kecen Pan , Yuhang Ming , Hua Zhang , Wanzeng Kong

This paper presents a generalized framework for the simulation of multiple robots and drones in highly realistic models of natural environments. The proposed simulation architecture uses the Unreal Engine4 for generating both optical and…

机器人学 · 计算机科学 2017-08-08 Ori Ganoni , Ramakrishnan Mukundan

We propose a deep inverse rendering framework for indoor scenes. From a single RGB image of an arbitrary indoor scene, we create a complete scene reconstruction, estimating shape, spatially-varying lighting, and spatially-varying,…

计算机视觉与模式识别 · 计算机科学 2019-05-09 Zhengqin Li , Mohammad Shafiei , Ravi Ramamoorthi , Kalyan Sunkavalli , Manmohan Chandraker

Current video retrieval systems, especially those used in competitions, primarily focus on querying individual keyframes or images rather than encoding an entire clip or video segment. However, queries often describe an action or event over…

计算机视觉与模式识别 · 计算机科学 2026-04-29 Quoc-Bao Nguyen-Le , Thanh-Huy Le-Nguyen

Modern video editing techniques have achieved high visual fidelity when inserting video objects. However, they focus on optimizing visual fidelity rather than physical causality, leading to edits that are physically inconsistent with their…

计算机视觉与模式识别 · 计算机科学 2026-03-09 Bohai Gu , Taiyi Wu , Dazhao Du , Jian Liu , Shuai Yang , Xiaotong Zhao , Alan Zhao , Song Guo

Humans exhibit an innate capacity to rapidly perceive and segment objects from video observations, and even mentally assemble them into structured 3D scenes. Replicating such capability, termed compositional 3D reconstruction, is pivotal…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Mingyu Dong , Chong Xia , Mingyuan Jia , Weichen Lyu , Long Xu , Zheng Zhu , Yueqi Duan

We propose a system that uses video as the input to track the position of objects relative to their surrounding environment in real-time. The neural network employed is trained on a 100% synthetic dataset coming from our own automated…

计算机视觉与模式识别 · 计算机科学 2020-10-30 David Albarracín , Jesús Hormigo , José David Fernández

We present CAT-V (Caption AnyThing in Video), a training-free framework for fine-grained object-centric video captioning that enables detailed descriptions of user-selected objects through time. CAT-V integrates three key components: a…

Inserting 3D objects into videos is a longstanding challenge in computer graphics with applications in augmented reality, virtual try-on, and video composition. Achieving both temporal consistency, or realistic lighting remains difficult,…

计算机视觉与模式识别 · 计算机科学 2025-07-30 Chenjian Gao , Lihe Ding , Rui Han , Zhanpeng Huang , Zibin Wang , Tianfan Xue

Visual object tracking is a fundamental video task in computer vision. Recently, the notably increasing power of perception algorithms allows the unification of single/multiobject and box/mask-based tracking. Among them, the Segment…

计算机视觉与模式识别 · 计算机科学 2023-07-27 Jiawen Zhu , Zhenyu Chen , Zeqi Hao , Shijie Chang , Lu Zhang , Dong Wang , Huchuan Lu , Bin Luo , Jun-Yan He , Jin-Peng Lan , Hanyuan Chen , Chenyang Li

User-generated cinematic creations are gaining popularity as our daily entertainment, yet it is a challenge to master cinematography for producing immersive contents. Many existing automatic methods focus on roughly controlling predefined…

多媒体 · 计算机科学 2024-05-24 Xinyi Wu , Haohong Wang , Aggelos K. Katsaggelos

Large text-to-image diffusion models have exhibited impressive proficiency in generating high-quality images. However, when applying these models to video domain, ensuring temporal consistency across video frames remains a formidable…

计算机视觉与模式识别 · 计算机科学 2023-09-19 Shuai Yang , Yifan Zhou , Ziwei Liu , Chen Change Loy

In this paper we present a framework for the rendering of dynamic 3D virtual environments which can be integrated in the development of videogames. It includes methods to manage sounds and particle effects, paged static geometries, the…

图形学 · 计算机科学 2013-06-06 Salvatore Catanese , Emilio Ferrara , Giacomo Fiumara , Francesco Pagano

Mask-free video object insertion has emerged as a challenging task, requiring harmonious integration of reference objects into source videos. However, existing methods struggle when references exhibit severe stylistic domain gaps with the…

计算机视觉与模式识别 · 计算机科学 2026-05-25 Xiao Cao , Yansong Qu , Xiangzhen , Chang , Wen Xiao , Jiakui Hu , Heyuan Li , Jialun Liu , Zhiyong Huang , Xuelong Li

While the satellite-based Global Positioning System (GPS) is adequate for some outdoor applications, many other applications are held back by its multi-meter positioning errors and poor indoor coverage. In this paper, we study the…

计算机视觉与模式识别 · 计算机科学 2020-02-20 Abm Musa , Jakob Eriksson