中文
相关论文

相关论文: OASIS: Real-Time Opti-Acoustic Sensing for Interve…

200 篇论文

We propose DeepASA, a multi-purpose model for auditory scene analysis that performs multi-input multi-output (MIMO) source separation, dereverberation, sound event detection (SED), audio classification, and direction-of-arrival estimation…

音频与语音处理 · 电气工程与系统科学 2026-04-16 Dongheon Lee , Younghoo Kwon , Jung-Woo Choi

Open-vocabulary semantic mapping enables robots to spatially ground previously unseen concepts without requiring predefined class sets. Current training-free methods commonly rely on multi-view fusion of semantic embeddings into a 3D map,…

Training large language models (LLMs) is constrained by memory requirements, with activations accounting for a substantial fraction of the total footprint. Existing approaches reduce memory using low-rank weight parameterizations or…

机器学习 · 计算机科学 2026-04-13 Sakshi Choudhary , Utkarsh Saxena , Kaushik Roy

The efficient fusion of depth maps is a key part of most state-of-the-art 3D reconstruction methods. Besides requiring high accuracy, these depth fusion methods need to be scalable and real-time capable. To this end, we present a novel…

计算机视觉与模式识别 · 计算机科学 2020-04-06 Silvan Weder , Johannes L. Schönberger , Marc Pollefeys , Martin R. Oswald

While 3D Gaussian representations (3DGS) have proven effective for modeling the geometry and appearance of objects, their potential for capturing other physical attributes-such as sound-remains largely unexplored. In this paper, we present…

声音 · 计算机科学 2025-07-29 Chunshi Wang , Hongxing Li , Yawei Luo

3D Gaussian Splatting (3DGS) provides an explicit and efficient scene representation, but its primitives lack inherent object-level identity, hindering downstream tasks such as open-vocabulary scene understanding. Existing methods typically…

计算机视觉与模式识别 · 计算机科学 2026-05-20 Guiyu Liu , Niklas Vaara , Janne Mustaniemi , Juho Kannala , Janne Heikkilä

Iterative image reconstruction algorithms for optoacoustic tomography (OAT), also known as photoacoustic tomography, have the ability to improve image quality over analytic algorithms due to their ability to incorporate accurate models of…

偏微分方程分析 · 数学 2015-06-05 Kun wang , Richard Su , Alexander A. Oraevsky , Mark A. Anastasio

This paper presents an adaptive visual servoing framework for robotic on-orbit servicing (OOS), specifically designed for capturing tumbling satellites. The vision-guided robotic system is capable of selecting optimal control actions in the…

机器人学 · 计算机科学 2024-09-10 Farhad Aghili

Recent advances in leveraging large-scale Internet photo collections for 3D reconstruction have enabled immersive virtual exploration of landmarks and historic sites worldwide. However, little attention has been given to the immersive…

图形学 · 计算机科学 2025-08-06 Yuze Wang , Yue Qi

This paper introduces 3DFIRES, a novel system for scene-level 3D reconstruction from posed images. Designed to work with as few as one view, 3DFIRES reconstructs the complete geometry of unseen scenes, including hidden surfaces. With…

计算机视觉与模式识别 · 计算机科学 2024-03-14 Linyi Jin , Nilesh Kulkarni , David Fouhey

Current 3D instance segmentation models generally use multi-stage methods to extract instance objects, including clustering, feature extraction, and post-processing processes. However, these multi-stage approaches rely on hyperparameter…

计算机视觉与模式识别 · 计算机科学 2023-03-14 Chuan Tang , Xi Yang

Robust 3D human pose estimation is crucial to ensure safe and effective human-robot collaboration. Accurate human perception,however, is particularly challenging in these scenarios due to strong occlusions and limited camera viewpoints.…

计算机视觉与模式识别 · 计算机科学 2024-08-29 Laura Bragagnolo , Matteo Terreran , Davide Allegro , Stefano Ghidoni

Unmanned Aerial Vehicles (UAVs) have gained significant popularity in scene reconstruction. This paper presents SOAR, a LiDAR-Visual heterogeneous multi-UAV system specifically designed for fast autonomous reconstruction of complex…

机器人学 · 计算机科学 2024-09-05 Mingjie Zhang , Chen Feng , Zengzhi Li , Guiyong Zheng , Yiming Luo , Zhu Wang , Jinni Zhou , Shaojie Shen , Boyu Zhou

We present a novel 3D mapping method leveraging the recent progress in neural implicit representation for 3D reconstruction. Most existing state-of-the-art neural implicit representation methods are limited to object-level reconstructions…

计算机视觉与模式识别 · 计算机科学 2023-08-08 Stefan Lionar , Lukas Schmid , Cesar Cadena , Roland Siegwart , Andrei Cramariuc

We introduce 3D-SIS, a novel neural network architecture for 3D semantic instance segmentation in commodity RGB-D scans. The core idea of our method is to jointly learn from both geometric and color signal, thus enabling accurate instance…

计算机视觉与模式识别 · 计算机科学 2019-04-30 Ji Hou , Angela Dai , Matthias Nießner

The research on neural radiance fields for new view synthesis has experienced explosive growth with the development of new models and extensions. The NERF algorithm, suitable for underwater scenes or scattering media, is also evolving.…

计算机视觉与模式识别 · 计算机科学 2023-12-15 Zhuoyifan Zhang , Lu Zhang , Liang Wang , Haoming Wu

Coded aperture snapshot spectral imaging (CASSI) retrieves a 3D hyperspectral image (HSI) from a single 2D compressed measurement, which is a highly challenging reconstruction task. Recent deep unfolding networks (DUNs), empowered by…

计算机视觉与模式识别 · 计算机科学 2025-09-16 Xiaodong Wang , Ping Wang , Zijun He , Mengjie Qin , Xin Yuan

3D scene understanding is crucial for facilitating seamless interaction between digital devices and the physical world. Real-time capturing and processing of the 3D scene are essential for achieving this seamless integration. While existing…

计算机视觉与模式识别 · 计算机科学 2024-10-04 Remco Royen , Kostas Pataridis , Ward van der Tempel , Adrian Munteanu

Optical coherence tomography (OCT) is a non-invasive imaging technique widely used for ophthalmology. It can be extended to OCT angiography (OCT-A), which reveals the retinal vasculature with improved contrast. Recent deep learning…

图像与视频处理 · 电气工程与系统科学 2021-07-12 Dewei Hu , Can Cui , Hao Li , Kathleen E. Larson , Yuankai K. Tao , Ipek Oguz

Open-vocabulary 3D scene understanding is indispensable for embodied agents. Recent works leverage pretrained vision-language models (VLMs) for object segmentation and project them to point clouds to build 3D maps. Despite progress, a point…

计算机视觉与模式识别 · 计算机科学 2026-03-18 Zhigang Wang , Yifei Su , Chenhui Li , Dong Wang , Yan Huang , Bin Zhao , Xuelong Li