中文
相关论文

相关论文: SceneFactory: A Workflow-centric and Unified Frame…

200 篇论文

Research on 3D Vision-Language Models (3D-VLMs) is gaining increasing attention, which is crucial for developing embodied AI within 3D scenes, such as visual navigation and embodied question answering. Due to the high density of visual…

计算机视觉与模式识别 · 计算机科学 2025-02-04 Hongyan Zhi , Peihao Chen , Junyan Li , Shuailei Ma , Xinyu Sun , Tianhang Xiang , Yinjie Lei , Mingkui Tan , Chuang Gan

We present SceneFactor, a diffusion-based approach for large-scale 3D scene generation that enables controllable generation and effortless editing. SceneFactor enables text-guided 3D scene synthesis through our factored diffusion…

计算机视觉与模式识别 · 计算机科学 2024-12-04 Alexey Bokhovkin , Quan Meng , Shubham Tulsiani , Angela Dai

In the field of sketch generation, raster-format trained models often produce non-stroke artifacts, while vector-format trained models typically lack a holistic understanding of sketches, leading to compromised recognizability. Moreover,…

图形学 · 计算机科学 2025-11-19 Jin Zhou , Yi Zhou , Hongliang Yang , Pengfei Xu , Hui Huang

We present NextFlow, a unified decoder-only autoregressive transformer trained on 6 trillion interleaved text-image discrete tokens. By leveraging a unified vision representation within a unified autoregressive architecture, NextFlow…

Many multimodal learning tasks require supervision that remains consistent across edits, viewpoints, and scene-level interventions. However, such supervision is difficult to obtain from observation-level datasets, which do not expose the…

计算机视觉与模式识别 · 计算机科学 2026-05-15 Jizhizi Li , Jiayang Ao , Danny Wicks , Petru-Daniel Tudosiu

Dense 3D facial motion capture from only monocular in-the-wild pairs of RGB images is a highly challenging problem with numerous applications, ranging from facial expression recognition to facial reenactment. In this work, we propose…

计算机视觉与模式识别 · 计算机科学 2020-05-18 Mohammad Rami Koujan , Anastasios Roussos , Stefanos Zafeiriou

We present TraceFlow, a novel framework for high-fidelity rendering of dynamic specular scenes by addressing two key challenges: precise reflection direction estimation and physically accurate reflection modeling. To achieve this, we…

计算机视觉与模式识别 · 计算机科学 2025-12-12 Jiachen Tao , Junyi Wu , Haoxuan Wang , Zongxin Yang , Dawen Cai , Yan Yan

Multi-frame depth estimation improves over single-frame approaches by also leveraging geometric relationships between images via feature matching, in addition to learning appearance-based features. In this paper we revisit feature matching…

计算机视觉与模式识别 · 计算机科学 2022-06-14 Vitor Guizilini , Rares Ambrus , Dian Chen , Sergey Zakharov , Adrien Gaidon

The whole is greater than the sum of its parts-even in 3D-text contrastive learning. We introduce SceneForge, a novel framework that enhances contrastive alignment between 3D point clouds and text through structured multi-object scene…

计算机视觉与模式识别 · 计算机科学 2025-10-17 Cristian Sbrolli , Matteo Matteucci

World building with 3D scene representations is increasingly important for content creation, simulation, and interactive experiences, yet real workflows are inherently iterative: creators must repeatedly extend an existing scene under user…

计算机视觉与模式识别 · 计算机科学 2026-04-01 Zijian He , Renjie Liu , Yihao Wang , Weizhi Zhong , Huan Yuan , Kun Gai , Guangrun Wang , Guanbin Li

3D content generation has recently attracted significant research interest, driven by its critical applications in VR/AR and embodied AI. In this work, we tackle the challenging task of synthesizing multiple 3D assets within a single scene…

计算机视觉与模式识别 · 计算机科学 2025-12-10 Yanxu Meng , Haoning Wu , Ya Zhang , Weidi Xie

When working with 3D facial data, improving fidelity and avoiding the uncanny valley effect is critically dependent on accurate 3D facial performance capture. Because such methods are expensive and due to the widespread availability of 2D…

计算机视觉与模式识别 · 计算机科学 2024-06-26 Felix Taubner , Prashant Raina , Mathieu Tuli , Eu Wern Teh , Chul Lee , Jinmiao Huang

Generating high-dimensional visual modalities is a computationally intensive task. A common solution is progressive generation, where the outputs are synthesized in a coarse-to-fine spectral autoregressive manner. While diffusion models…

计算机视觉与模式识别 · 计算机科学 2025-06-25 Moayed Haji-Ali , Willi Menapace , Ivan Skorokhodov , Arpit Sahni , Sergey Tulyakov , Vicente Ordonez , Aliaksandr Siarohin

The assumption of scene rigidity is typical in SLAM algorithms. Such a strong assumption limits the use of most visual SLAM systems in populated real-world environments, which are the target of several relevant applications like service…

计算机视觉与模式识别 · 计算机科学 2018-08-16 Berta Bescos , José M. Fácil , Javier Civera , José Neira

Unsupervised methods have showed promising results on monocular depth estimation. However, the training data must be captured in scenes without moving objects. To push the envelope of accuracy, recent methods tend to increase their model…

计算机视觉与模式识别 · 计算机科学 2023-03-09 Tak-Wai Hui

Indoor scene synthesis has become increasingly important with the rise of Embodied AI, which requires 3D environments that are not only visually realistic but also physically plausible and functionally diverse. While recent approaches have…

图形学 · 计算机科学 2025-10-28 Yandan Yang , Baoxiong Jia , Shujie Zhang , Siyuan Huang

Multi-camera systems have been shown to improve the accuracy and robustness of SLAM estimates, yet state-of-the-art SLAM systems predominantly support monocular or stereo setups. This paper presents a generic sparse visual SLAM framework…

机器人学 · 计算机科学 2024-05-10 Pushyami Kaveti , Shankara Narayanan Vaidyanathan , Arvind Thamilchelvan , Hanumant Singh

Scene flow estimation predicts the 3D motion at each point in successive LiDAR scans. This detailed, point-level, information can help autonomous vehicles to accurately predict and understand dynamic changes in their surroundings. Current…

计算机视觉与模式识别 · 计算机科学 2024-09-18 Qingwen Zhang , Yi Yang , Peizheng Li , Olov Andersson , Patric Jensfelt

3D Multi-Object Tracking (MOT) is an important part of the unmanned vehicle perception module. Most methods optimize object detection and data association independently. These methods make the network structure complicated and limit the…

计算机视觉与模式识别 · 计算机科学 2022-03-07 Yueling Shen , Guangming Wang , Hesheng Wang

Collaborative perception plays a crucial role in enhancing environmental understanding by expanding the perceptual range and improving robustness against sensor failures, which primarily involves collaborative 3D detection and tracking…

计算机视觉与模式识别 · 计算机科学 2025-06-10 Xunjie He , Christina Dao Wen Lee , Meiling Wang , Chengran Yuan , Zefan Huang , Yufeng Yue , Marcelo H. Ang