中文
相关论文

相关论文: ClickAIXR: On-Device Multimodal Vision-Language In…

200 篇论文

Scalable interactive visual data exploration is crucial in many domains due to increasingly large datasets generated at rapid rates. Details-on-demand provides a useful interaction paradigm for exploring large datasets, where users start at…

人机交互 · 计算机科学 2019-05-14 Wenbo Tao , Xiaoyu Liu , Çağatay Demiralp , Remco Chang , Michael Stonebraker

Simulations constitute a fundamental component of medical and nursing education and traditionally employ standardized patients (SP) and high-fidelity manikins to develop clinical reasoning and communication skills. However, these methods…

人机交互 · 计算机科学 2025-10-23 Akilan Amithasagaran , Sagnik Dakshit , Bhavani Suryadevara , Lindsey Stockton

We introduce InteractVLM, a novel method to estimate 3D contact points on human bodies and objects from single in-the-wild images, enabling accurate human-object joint reconstruction in 3D. This is challenging due to occlusions, depth…

计算机视觉与模式识别 · 计算机科学 2025-04-08 Sai Kumar Dwivedi , Dimitrije Antić , Shashank Tripathi , Omid Taheri , Cordelia Schmid , Michael J. Black , Dimitrios Tzionas

Video-to-audio (V2A) generation shows great potential in fields such as film production. Despite significant advances, current V2A methods relying on global video information struggle with complex scenes and generating audio tailored to…

计算机视觉与模式识别 · 计算机科学 2025-07-15 Yingshan Liang , Keyu Fan , Zhicheng Du , Yiran Wang , Qingyang Shi , Xinyu Zhang , Jiasheng Lu , Peiwu Qin

Physical touch, a fundamental aspect of human social interaction, remains largely absent in real-time virtual communication. We present a haptic-enabled multi-user Virtual Reality (VR) system that facilitates real-time, bi-directional…

人机交互 · 计算机科学 2025-10-23 Premankur Banerjee , Jiaxuan Wang , Lauren Tomita , Mia P Montiel , Heather Culbertson

A Human Computer Interface (HCI) System for playing games is designed here for more natural communication with the machines. The system presented here is a vision-based system for detection of long voluntary eye blinks and interpretation of…

人机交互 · 计算机科学 2010-02-11 S. Sumathi , S. K. Srivatsa , M. Uma Maheswari

Large Language Models (LLMs) are transforming personalized search, recommendations, and customer interaction in e-commerce. Customers increasingly shop across multiple devices, from voice-only assistants to multimodal displays, each…

信息检索 · 计算机科学 2025-11-20 Mariya Hendriksen , Svitlana Vakulenko , Jordan Massiah , Gabriella Kazai , Emine Yilmaz

Understanding objects through multiple sensory modalities is fundamental to human perception, enabling cross-sensory integration and richer comprehension. For AI and robotic systems to replicate this ability, access to diverse, high-quality…

计算机视觉与模式识别 · 计算机科学 2025-04-04 Samuel Clarke , Suzannah Wistreich , Yanjie Ze , Jiajun Wu

Composed Image Retrieval (CIR) has demonstrated significant potential by enabling flexible multimodal queries that combine a reference image and modification text. However, CIR inherently prioritizes semantic matching, struggling to…

计算机视觉与模式识别 · 计算机科学 2026-04-08 Yuxin Yang , Yinan Zhou , Yuxin Chen , Ziqi Zhang , Zongyang Ma , Chunfeng Yuan , Bing Li , Jun Gao , Weiming Hu

We propose a novel visual SLAM method that integrates text objects tightly by treating them as semantic features via fully exploring their geometric and semantic prior. The text object is modeled as a texture-rich planar patch whose…

计算机视觉与模式识别 · 计算机科学 2023-07-04 Boying Li , Danping Zou , Yuan Huang , Xinghan Niu , Ling Pei , Wenxian Yu

MolecularWebXR is our new website for education, science communication and scientific peer discussion in chemistry and biology built on WebXR. It democratizes multi-user, inclusive virtual reality (VR) experiences that are deeply immersive…

Preoperative medical imaging is an essential part of surgical planning. The data from medical imaging devices, such as CT and MRI scanners, consist of stacks of 2D images in DICOM format. Conversely, advances in 3D data visualization…

多媒体 · 计算机科学 2022-10-27 Menghe Zhang , Weichen Liu , Nadir Weibel , Jurgen Schulze

Patch-based models, e.g., Vision Transformers (ViTs) and Mixers, have shown impressive results on various visual recognition tasks, alternating classic convolutional networks. While the initial patch-based models (ViTs) treated all patches…

计算机视觉与模式识别 · 计算机科学 2022-12-14 Hyunwoo Kang , Sangwoo Mo , Jinwoo Shin

Accurately reconstructing human behavior in close-interaction scenarios is crucial for enabling realistic virtual interactions in augmented reality, precise motion analysis in sports, and natural collaborative behavior in human-robot tasks.…

计算机视觉与模式识别 · 计算机科学 2026-04-16 Qi Xia , Peishan Cong , Ziyi Wang , Yujing Sun , Qin Sun , Xinge Zhu , Mao Ye , Ruigang Yang , Yuexin Ma

Large Vision Language Models (LVLMs) have demonstrated impressive zero-shot capabilities in various vision-language dialogue scenarios. However, the absence of fine-grained visual object detection hinders the model from understanding the…

计算与语言 · 计算机科学 2024-04-15 Junyu Lu , Dixiang Zhang , Songxin Zhang , Zejian Xie , Zhuoyang Song , Cong Lin , Jiaxing Zhang , Bingyi Jing , Pingjian Zhang

Despite the growing accessibility of augmented reality (AR) for visualization, existing computer-aided design (CAD) systems remain confined to traditional screens or require complex setups or predefined parameters, limiting immersion and…

人机交互 · 计算机科学 2025-09-22 Oğuz Arslan , Artun Akdoğan , Mustafa Doga Dogan

Existing object removal tools often rely on manual masks or text prompts, making precise removal difficult for non-expert users in complex scenes and often leading to incomplete removal or unnatural background completion. To address this…

计算机视觉与模式识别 · 计算机科学 2026-05-15 Ledun Zhang , Yatu Ji , Xufei Zhuang , Xinying Yao

Video Large Language Models (Video LLMs) have shown remarkable progress in understanding and reasoning about visual content, particularly in tasks involving text recognition and text-based visual question answering (Text VQA). However,…

LiDAR-camera fusion can enhance the performance of 3D object detection by utilizing complementary information between depth-aware LiDAR points and semantically rich images. Existing voxel-based methods face significant challenges when…

计算机视觉与模式识别 · 计算机科学 2025-03-05 Ziying Song , Guoxin Zhang , Jun Xie , Lin Liu , Caiyan Jia , Shaoqing Xu , Zhepeng Wang

The development of large language models (LLMs) has greatly advanced the field of multimodal understanding, leading to the emergence of large multimodal models (LMMs). In order to enhance the level of visual comprehension, recent studies…

计算机视觉与模式识别 · 计算机科学 2023-12-19 Ao Zhang , Yuan Yao , Wei Ji , Zhiyuan Liu , Tat-Seng Chua