中文
相关论文

相关论文: Z3D: Zero-Shot 3D Visual Grounding from Images

200 篇论文

3D object grounding aims to locate the most relevant target object in a raw point cloud scene based on a free-form language description. Understanding complex and diverse descriptions, and lifting them directly to a point cloud is a new and…

计算机视觉与模式识别 · 计算机科学 2021-03-31 Mingtao Feng , Zhen Li , Qi Li , Liang Zhang , XiangDong Zhang , Guangming Zhu , Hui Zhang , Yaonan Wang , Ajmal Mian

We introduce ZeroVO, a novel visual odometry (VO) algorithm that achieves zero-shot generalization across diverse cameras and environments, overcoming limitations in existing methods that depend on predefined or static camera calibration…

计算机视觉与模式识别 · 计算机科学 2025-06-10 Lei Lai , Zekai Yin , Eshed Ohn-Bar

We introduce Zero-1-to-3, a framework for changing the camera viewpoint of an object given just a single RGB image. To perform novel view synthesis in this under-constrained setting, we capitalize on the geometric priors that large-scale…

计算机视觉与模式识别 · 计算机科学 2023-03-21 Ruoshi Liu , Rundi Wu , Basile Van Hoorick , Pavel Tokmakov , Sergey Zakharov , Carl Vondrick

We introduce LiveSVG, a zero-shot approach for generating Scalable Vector Graphics (SVG) animations using video diffusion models. Current SVG animation methods struggle with complex motions: LLM-based code synthesis fails to express fine,…

计算机视觉与模式识别 · 计算机科学 2026-05-29 Matan Levy , Ran Margolin , Bar Cavia , Dvir Samuel , Yael Pritch , Shmuel Peleg , Alex Rav Acha , Ariel Shamir , Dani Lischinski

General 3D foundation models have started to lead the trend of unifying diverse vision tasks, yet most assume RGB-only inputs and ignore readily available geometric cues (e.g., camera intrinsics, poses, and depth maps). To address this…

计算机视觉与模式识别 · 计算机科学 2025-11-17 Haosong Peng , Hao Li , Yalun Dai , Yushi Lan , Yihang Luo , Tianyu Qi , Zhengshen Zhang , Yufeng Zhan , Junfei Zhang , Wenchao Xu , Ziwei Liu

Multi-view Detection (MVD) is highly effective for occlusion reasoning in a crowded environment. While recent works using deep learning have made significant advances in the field, they have overlooked the generalization aspect, which makes…

计算机视觉与模式识别 · 计算机科学 2022-03-15 Jeet Vora , Swetanjal Dutta , Kanishk Jain , Shyamgopal Karthik , Vineet Gandhi

Previous research has investigated the application of Multimodal Large Language Models (MLLMs) in understanding 3D scenes by interpreting them as videos. These approaches generally depend on comprehensive 3D data inputs, such as point…

计算机视觉与模式识别 · 计算机科学 2025-10-23 Duo Zheng , Shijia Huang , Yanyang Li , Liwei Wang

We propose a novel zero-shot approach for keypoint detection on 3D shapes. Point-level reasoning on visual data is challenging as it requires precise localization capability, posing problems even for powerful models like DINO or CLIP.…

计算机视觉与模式识别 · 计算机科学 2024-12-10 Bingchen Gong , Diego Gomez , Abdullah Hamdi , Abdelrahman Eldesokey , Ahmed Abdelreheem , Peter Wonka , Maks Ovsjanikov

Estimating the 3D position and orientation of objects in the environment with a single RGB camera is a critical and challenging task for low-cost urban autonomous driving and mobile robots. Most of the existing algorithms are based on the…

计算机视觉与模式识别 · 计算机科学 2021-02-02 Yuxuan Liu , Yuan Yixuan , Ming Liu

Video stabilization aims to mitigate camera shake but faces a fundamental trade-off between geometric robustness and full-frame consistency. While 2D methods suffer from aggressive cropping, 3D techniques are often undermined by fragile…

计算机视觉与模式识别 · 计算机科学 2026-03-09 Muhua Zhu , Xinhao Jin , Yu Zhang , Yifei Xue , Tie Ji , Yizhen Lao

Video temporal grounding (VTG) aims to locate specific temporal segments from an untrimmed video based on a linguistic query. Most existing VTG models are trained on extensive annotated video-text pairs, a process that not only introduces…

计算机视觉与模式识别 · 计算机科学 2024-03-05 Yifang Xu , Yunzhuo Sun , Zien Xie , Benxiang Zhai , Sidan Du

Scaling up representations for images or text has been extensively investigated in the past few years and has led to revolutions in learning vision and language. However, scalable representation for 3D objects and scenes is relatively…

计算机视觉与模式识别 · 计算机科学 2023-10-11 Junsheng Zhou , Jinsheng Wang , Baorui Ma , Yu-Shen Liu , Tiejun Huang , Xinlong Wang

Embodied navigation is a fundamental capability for robotic agents operating. Real-world deployment requires open vocabulary generalization and low training overhead, motivating zero-shot methods rather than task-specific RL training.…

计算机视觉与模式识别 · 计算机科学 2026-03-18 Xun Huang , Shijia Zhao , Yunxiang Wang , Xin Lu , Wanfa Zhang , Rongsheng Qu , Weixin Li , Yunhong Wang , Chenglu Wen

The ability to interpret and comprehend a 3D scene is essential for many vision and robotics systems. In numerous applications, this involves 3D object detection, i.e.~identifying the location and dimensions of objects belonging to a…

计算机视觉与模式识别 · 计算机科学 2025-09-22 Olivier Moliner , Viktor Larsson , Kalle Åström

A phrase grounding system localizes a particular object in an image referred to by a natural language query. In previous work, the phrases were restricted to have nouns that were encountered in training, we extend the task to Zero-Shot…

计算机视觉与模式识别 · 计算机科学 2019-08-21 Arka Sadhu , Kan Chen , Ram Nevatia

3D vision-language grounding, which focuses on aligning language with the 3D physical environment, stands as a cornerstone in the development of embodied agents. In comparison to recent advancements in the 2D domain, grounding language in…

计算机视觉与模式识别 · 计算机科学 2024-09-25 Baoxiong Jia , Yixin Chen , Huangyue Yu , Yan Wang , Xuesong Niu , Tengyu Liu , Qing Li , Siyuan Huang

As the development of deep neural networks, 3D object recognition is becoming increasingly popular in computer vision community. Many multi-view based methods are proposed to improve the category recognition accuracy. These approaches…

计算机视觉与模式识别 · 计算机科学 2019-06-18 Qi Xuan , Fuxian Li , Yi Liu , Yun Xiang

Visual grounding (VG) aims to localize target objects in an image based on natural language descriptions. In this paper, we propose AerialVG, a new task focusing on visual grounding from aerial views. Compared to traditional VG, AerialVG…

计算机视觉与模式识别 · 计算机科学 2025-10-09 Junli Liu , Qizhi Chen , Zhigang Wang , Yiwen Tang , Yiting Zhang , Chi Yan , Dong Wang , Xuelong Li , Bin Zhao

Though recent advances in vision-language models (VLMs) have achieved remarkable progress across a wide range of multimodal tasks, understanding 3D spatial relationships from limited views remains a significant challenge. Previous reasoning…

计算机视觉与模式识别 · 计算机科学 2026-03-16 Zhangquan Chen , Manyuan Zhang , Xinlei Yu , Xufang Luo , Mingze Sun , Zihao Pan , Xiang An , Yan Feng , Peng Pei , Xunliang Cai , Ruqi Huang

This paper presents VGGT-360, a novel training-free framework for zero-shot, geometry-consistent panoramic depth estimation. Unlike prior view-independent training-free approaches, VGGT-360 reformulates the task as panoramic reprojection…

计算机视觉与模式识别 · 计算机科学 2026-05-15 Jiayi Yuan , Haobo Jiang , De Wen Soh , Na Zhao