中文
相关论文

相关论文: Explore, Listen, Inspect: Supporting Multimodal In…

200 篇论文

Monocular 3D object detection poses a significant challenge in 3D scene understanding due to its inherently ill-posed nature in monocular depth estimation. Existing methods heavily rely on supervised learning using abundant 3D labels,…

计算机视觉与模式识别 · 计算机科学 2024-04-02 Zihua Liu , Hiroki Sakuma , Masatoshi Okutomi

Humans naturally interact with their 3D surroundings using language, and modeling 3D language fields for scene understanding and interaction has gained growing interest. This paper introduces ChatSplat, a system that constructs a 3D…

计算机视觉与模式识别 · 计算机科学 2024-12-03 Hanlin Chen , Fangyin Wei , Gim Hee Lee

We investigate the perception of visual variables on wall-sized tiled displays within an immersive environment. We designed and conducted two formal user studies focusing on elementary visualization reading tasks in VR. The first study…

人机交互 · 计算机科学 2025-01-20 Dongyun Han , Anastasia Bezerianos , Petra Isenberg , Isaac Cho

Blind and Low Vision (BLV) people have adopted AI-powered visual interpretation applications to address their daily needs. While these applications have been helpful, prior work has found that users remain unsatisfied by their frequent…

人机交互 · 计算机科学 2025-03-11 Ricardo E. Gonzalez Penuela , Ruiying Hu , Sharon Lin , Tanisha Shende , Shiri Azenkot

Interactive robots navigating photo-realistic environments need to be trained to effectively leverage and handle the dynamic nature of dialogue in addition to the challenges underlying vision-and-language navigation (VLN). In this paper, we…

计算机视觉与模式识别 · 计算机科学 2022-03-17 Ayush Shrivastava , Karthik Gopalakrishnan , Yang Liu , Robinson Piramuthu , Gokhan Tür , Devi Parikh , Dilek Hakkani-Tür

Vision-based Interfaces (VIs) are pivotal in advancing Human-Computer Interaction (HCI), particularly in enhancing context awareness. However, there are significant opportunities for these interfaces due to rapid advancements in multimodal…

人机交互 · 计算机科学 2024-08-15 Yongquan Hu , Wen Hu , Aaron Quigley

Despite widespread use, charts remain largely inaccessible for Low-Vision Individuals (LVI). Reading charts requires viewing data points within a global context, which is difficult for LVI who may rely on magnification or experience a…

人机交互 · 计算机科学 2026-03-04 Yotam Sechayk , Hennes Rave , Max Rädler , Mark Colley , Zhongyi Zhou , Ariel Shamir , Takeo Igarashi

Open-vocabulary 3D scene understanding enables users to segment novel objects in complex 3D environments through natural language. However, existing approaches remain slow, memory-intensive, and overly complex due to iterative optimization…

计算机视觉与模式识别 · 计算机科学 2026-03-26 Jaehun Bang , Jinhyeok Kim , Minji Kim , Seungheon Jeong , Kyungdon Joo

Vision-and-Language Navigation (VLN) has gained significant research interest in recent years due to its potential applications in real-world scenarios. However, existing VLN methods struggle with the issue of spurious associations,…

计算机视觉与模式识别 · 计算机科学 2024-03-07 Liuyi Wang , Zongtao He , Ronghao Dang , Huiyi Chen , Chengju Liu , Qijun Chen

Augmented data storytelling enhances narrative delivery by integrating visualizations with physical environments and presenter actions. Existing systems predominantly rely on body gestures or speech to control visualizations, leaving…

人机交互 · 计算机科学 2025-08-08 Kentaro Takahira , Yue Yu , Takanori Fujiwara , Ryo Suzuki , Huamin Qu

Data is everywhere but may not be accessible to everyone. Conventional data visualization tools and guidelines often do not actively consider the specific needs and abilities of people with Intellectual and Developmental Disabilities (IDD),…

人机交互 · 计算机科学 2023-03-13 Keke Wu , Michelle H Tran , Emma Petersen , Varsha Koushik , Danielle Albers Szafir

In recent years, cross-modal domain adaptation has been studied on the paired 2D image and 3D LiDAR data to ease the labeling costs for 3D LiDAR semantic segmentation (3DLSS) in the target domain. However, in such a setting the paired 2D…

计算机视觉与模式识别 · 计算机科学 2023-08-08 Yiyang Chen , Shanshan Zhao , Changxing Ding , Liyao Tang , Chaoyue Wang , Dacheng Tao

Vision-language reward modeling faces a dilemma: generative approaches are interpretable but slow, while discriminative ones are efficient but act as opaque "black boxes." To bridge this gap, we propose VL-MDR (Vision-Language…

计算与语言 · 计算机科学 2026-04-08 Qiyuan Chen , Hongsen Huang , Jiahe Chen , Qian Shao , Jintai Chen , Hongxia Xu , Renjie Hua , Chuan Ren , Jian Wu

Understanding sensor data can be difficult for non-experts because of the complexity and different semantic meanings of sensor modalities. This leads to a need for intuitive and effective methods to present sensor information. However,…

人机交互 · 计算机科学 2025-03-26 Yunqi Guo , Kaiyuan Hou , Heming Fu , Hongkai Chen , Zhenyu Yan , Guoliang Xing , Xiaofan Jiang

The quest for deeper understanding of biological systems has driven the acquisition of increasingly larger multidimensional image datasets. Inspecting and manipulating data of this complexity is very challenging in traditional visualization…

图形学 · 计算机科学 2018-08-23 Stanislav Pidhorskyi , Michael Morehead , Quinn Jones , George Spirou , Gianfranco Doretto

New era has unlocked exciting possibilities for extending Large Language Models (LLMs) to tackle 3D vision-language tasks. However, most existing 3D multimodal LLMs (MLLMs) rely on compressing holistic 3D scene information or segmenting…

计算机视觉与模式识别 · 计算机科学 2025-07-23 Xiaoyan Wang , Zeju Li , Yifan Xu , Jiaxing Qi , Zhifei Yang , Ruifei Ma , Xiangde Liu , Chao Zhang

Accurate multi-view 3D object detection is essential for applications such as autonomous driving. Researchers have consistently aimed to leverage LiDAR's precise spatial information to enhance camera-based detectors through methods like…

计算机视觉与模式识别 · 计算机科学 2024-12-31 Shaoqing Xu , Fang Li , Peixiang Huang , Ziying Song , Zhi-Xin Yang

Accurate 3D lane estimation is crucial for ensuring safety in autonomous driving. However, prevailing monocular techniques suffer from depth loss and lighting variations, hampering accurate 3D lane detection. In contrast, LiDAR points offer…

计算机视觉与模式识别 · 计算机科学 2024-06-25 Yueru Luo , Shuguang Cui , Zhen Li

LiDAR has become one of the primary 3D object detection sensors in autonomous driving. However, LiDAR's diverging point pattern with increasing distance results in a non-uniform sampled point cloud ill-suited to discretized volumetric…

计算机视觉与模式识别 · 计算机科学 2022-03-23 Jordan S. K. Hu , Tianshu Kuai , Steven L. Waslander

Multimodal data has become a crucial element in the realm of big data analytics, driving advancements in data exploration, data mining, and empowering artificial intelligence applications. To support high-quality retrieval for these…

数据库 · 计算机科学 2025-02-11 Ming Sheng , Shuliang Wang , Yong Zhang , Kaige Wang , Jingyi Wang , Yi Luo , Rui Hao