中文
相关论文

相关论文: AREA3D: Active Reconstruction Agent with Unified F…

200 篇论文

Point-Vision-Language Models promise to empower embodied agents with executable spatial reasoning, yet they frequently succumb to geometric hallucination where predicted 3D structures contradict the observed 2D reality. We identify a key…

计算机视觉与模式识别 · 计算机科学 2026-04-24 Jingkun Chen , Ruoshi Xu , Mingqi Gao , Shengda Luo , Jungong Han

Generating immersive 3D scenes from texts is a core task in computer vision, crucial for applications in virtual reality and game development. Despite the promise of leveraging 2D diffusion priors, existing methods suffer from spatial…

计算机视觉与模式识别 · 计算机科学 2026-01-28 Jisheng Chu , Wenrui Li , Rui Zhao , Wangmeng Zuo , Shifeng Chen , Xiaopeng Fan

Remote sensing image fusion aims to create a high-resolution multi/hyper-spectral image from a high-resolution image with limited spectral information and a low-resolution image with abundant spectral data. Recently, deep learning (DL)…

计算机视觉与模式识别 · 计算机科学 2026-05-12 Siran Peng , Xiangyu Zhu , Shang-Qi Deng , Liang-Jian Deng , Zhen Lei

Remote sighted assistance (RSA) has emerged as a conversational assistive technology, where remote sighted workers, i.e., agents, provide real-time assistance to users with vision impairments via video-chat-like communication. Researchers…

人机交互 · 计算机科学 2022-02-04 Jingyi Xie , Rui Yu , Sooyeon Lee , Yao Lyu , Syed Masum Billah , John M. Carroll

With the proliferation of small aerial vehicles, acquiring close up aerial imagery for high quality reconstruction of complex scenes is gaining importance. We present an adaptive view planning method to collect such images in an automated…

计算机视觉与模式识别 · 计算机科学 2019-11-20 Cheng Peng , Volkan Isler

Novel view synthesis from monocular videos of dynamic scenes with unknown camera poses remains a fundamental challenge in computer vision and graphics. While recent advances in 3D representations such as Neural Radiance Fields (NeRF) and 3D…

计算机视觉与模式识别 · 计算机科学 2025-11-10 Mengqi Guo , Bo Xu , Yanyan Li , Gim Hee Lee

We introduce Structured 3D Features, a model based on a novel implicit 3D representation that pools pixel-aligned image features onto dense 3D points sampled from a parametric, statistical human mesh surface. The 3D points have associated…

计算机视觉与模式识别 · 计算机科学 2023-04-18 Enric Corona , Mihai Zanfir , Thiemo Alldieck , Eduard Gabriel Bazavan , Andrei Zanfir , Cristian Sminchisescu

We present ARCH++, an image-based method to reconstruct 3D avatars with arbitrary clothing styles. Our reconstructed avatars are animation-ready and highly realistic, in both the visible regions from input views and the unseen regions.…

计算机视觉与模式识别 · 计算机科学 2022-03-01 Tong He , Yuanlu Xu , Shunsuke Saito , Stefano Soatto , Tony Tung

This work concerns itself with the task of reconstructing all edges of an arbitrary 3D wire-frame model projected to an image plane. We explore a bottom-up part-wise procedure undertaken by an RL agent to segment and reconstruct these 2D…

机器学习 · 计算机科学 2025-03-24 Julian Ziegler , Patrick Frenzel , Mirco Fuchs

Recent AI-based 3D content creation has largely evolved along two paths: feed-forward image-to-3D reconstruction approaches and 3D generative models trained with 2D or 3D supervision. In this work, we show that existing feed-forward…

计算机视觉与模式识别 · 计算机科学 2025-01-07 Suttisak Wizadwongsa , Jinfan Zhou , Edward Li , Jeong Joon Park

Despite recent advancements in surface reconstruction, Level of Detail (LoD) 3 building reconstruction remains an unresolved challenge. The main issue pertains to the object-oriented modelling paradigm, which requires georeferencing,…

计算机视觉与模式识别 · 计算机科学 2025-04-08 Wenzhao Tang , Weihang Li , Xiucheng Liang , Olaf Wysocki , Filip Biljecki , Christoph Holst , Boris Jutzi

Holistic 3D scene understanding entails estimation of both layout configuration and object geometry in a 3D environment. Recent works have shown advances in 3D scene estimation from various input modalities (e.g., images, 3D scans), by…

计算机视觉与模式识别 · 计算机科学 2022-11-28 Yinyu Nie , Angela Dai , Xiaoguang Han , Matthias Nießner

Accurate, reproducible burn assessment is critical for treatment planning, healing monitoring, and medico-legal documentation, yet conventional visual inspection and 2D photography are subjective and limited for longitudinal comparison.…

计算机视觉与模式识别 · 计算机科学 2026-02-03 S. Kalaycioglu , C. Hong , K. Zhai , H. Xie , J. N. Wong

Vision-Language-Action (VLA) models have recently achieved remarkable progress in robotic perception and control, yet most existing approaches primarily rely on VLM trained using 2D images, which limits their spatial understanding and…

计算机视觉与模式识别 · 计算机科学 2026-04-27 Zhifeng Rao , Wenlong Chen , Lei Xie , Xia Hua , Dongfu Yin , Zhen Tian , F. Richard Yu

In this paper, we propose an adaptive keyframe selection method for improved 3D scene reconstruction in dynamic environments. The proposed method integrates two complementary modules: an error-based selection module utilizing photometric…

机器人学 · 计算机科学 2025-12-30 Raman Jha , Yang Zhou , Giuseppe Loianno

Editing a 3D indoor scene from natural language is conceptually straightforward but technically challenging. Existing open-vocabulary systems often regenerate large portions of a scene or rely on image-space edits that disrupt spatial…

计算机视觉与模式识别 · 计算机科学 2026-03-19 Seongrae Noh , SeungWon Seo , Gyeong-Moon Park , HyeongYeop Kang

Localizing objects and parts from natural language in 3D space is essential for robotics, AR, and embodied AI, yet existing methods face a trade-off between the accuracy and geometric consistency of per-scene optimization and the efficiency…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Bryce Grant , Aryeh Rothenberg , Atri Banerjee , Peng Wang

Although Multimodal Large Language Models have achieved remarkable progress, they still struggle with complex 3D spatial reasoning due to the reliance on 2D visual priors. Existing approaches typically mitigate this limitation either…

计算机视觉与模式识别 · 计算机科学 2026-04-09 Jiahua Chen , Qihong Tang , Weinong Wang , Qi Fan

3D visual grounding (3DVG), which aims to correlate a natural language description with the target object within a 3D scene, is a significant yet challenging task. Despite recent advancements in this domain, existing approaches commonly…

计算机视觉与模式识别 · 计算机科学 2025-01-17 Xinyi Wang , Na Zhao , Zhiyuan Han , Dan Guo , Xun Yang

Existing image-to-video generation methods often produce physically implausible motions and lack precise control over object dynamics. While prior approaches have incorporated physics simulators, they remain confined to 2D planar motions…

计算机视觉与模式识别 · 计算机科学 2026-04-28 Tianyidan Xie , Zhentao Huang , Mingjie Wang , Xin Huang , Jun Zhou , Minglun Gong , Zili Yi
‹ 上一页 1 8 9 10 下一页 ›