中文
相关论文

相关论文: ClickAIXR: On-Device Multimodal Vision-Language In…

200 篇论文

Interactive segmentation enables users to extract binary masks of target objects through simple interactions such as clicks, scribbles, and boxes. However, existing methods often support only limited interaction forms and struggle to…

计算机视觉与模式识别 · 计算机科学 2025-06-18 Xi Chen , Hengshuang Zhao

Research on 3D Vision-Language Models (3D-VLMs) is gaining increasing attention, which is crucial for developing embodied AI within 3D scenes, such as visual navigation and embodied question answering. Due to the high density of visual…

计算机视觉与模式识别 · 计算机科学 2025-02-04 Hongyan Zhi , Peihao Chen , Junyan Li , Shuailei Ma , Xinyu Sun , Tianhang Xiang , Yinjie Lei , Mingkui Tan , Chuang Gan

Real-world machine learning models require rigorous evaluation before deployment, especially in safety-critical domains like autonomous driving and surveillance. The evaluation of machine learning models often focuses on data slices, which…

计算机视觉与模式识别 · 计算机科学 2025-05-07 Xinyuan Yan , Xiwei Xuan , Jorge Piazentin Ono , Jiajing Guo , Vikram Mohanty , Shekar Arvind Kumar , Liang Gou , Bei Wang , Liu Ren

The advent of immersive Virtual Reality applications has transformed various domains, yet their integration with advanced artificial intelligence technologies like Visual Language Models remains underexplored. This study introduces a…

机器人学 · 计算机科学 2024-08-06 Mikhail Konenkov , Artem Lykov , Daria Trinitatova , Dzmitry Tsetserukou

Large Multimodal Models (LMMs) have shown strong potential for assisting users in tasks, such as programming, content creation, and information access, yet their interaction remains largely limited to traditional interfaces such as desktops…

人机交互 · 计算机科学 2026-02-12 Liuchuan Yu , Yongqi Zhang , Lap-Fai Yu

There is a growing interest in applying large language models (LLMs) in robotic tasks, due to their remarkable reasoning ability and extensive knowledge learned from vast training corpora. Grounding LLMs in the physical world remains an…

机器人学 · 计算机科学 2024-04-11 Wenqiang Lai , Yuan Gao , Tin Lun Lam

Recent advancements in dialogue systems have highlighted the significance of integrating multimodal responses, which enable conveying ideas through diverse modalities rather than solely relying on text-based interactions. This enrichment…

计算与语言 · 计算机科学 2024-07-08 Chang-Sheng Kao , Yun-Nung Chen

Large Language Model (LLM)-driven digital humans have sparked a series of recent studies on co-speech gesture generation systems. However, existing approaches struggle with real-time synthesis and long-text comprehension. This paper…

图形学 · 计算机科学 2025-06-03 Yueqian Guo , Tianzhao Li , Xin Lyu , Jiehaolin Chen , Zhaohan Wang , Sirui Xiao , Yurun Chen , Yezi He , Helin Li , Fan Zhang

Existing image-to-video generation methods often produce physically implausible motions and lack precise control over object dynamics. While prior approaches have incorporated physics simulators, they remain confined to 2D planar motions…

计算机视觉与模式识别 · 计算机科学 2026-04-28 Tianyidan Xie , Zhentao Huang , Mingjie Wang , Xin Huang , Jun Zhou , Minglun Gong , Zili Yi

The recent emergence and popularity of consumer-grade augmented reality (AR) glasses from major technology companies highlight their potential to become the next daily computing platform. A dominant design trend in this context is the…

人机交互 · 计算机科学 2026-03-30 Ziming Li , Hongji Li , Jialin Wang , Pan Hui , Hai-Ning Liang

Individuals with fine motor impairments, such as those caused by conditions like Parkinson's disease, cerebral palsy, or dyspraxia, face significant challenges in interacting with traditional computer interfaces. Historically, scripted…

人机交互 · 计算机科学 2024-11-20 Shridhar Mehendale , Ankit Walishetti

The integration of Large Language Models (LLMs) like GPT-4 with Extended Reality (XR) technologies offers the potential to build truly immersive XR environments that interact with human users through natural language, e.g., generating and…

多媒体 · 计算机科学 2025-03-21 Jiangong Chen , Xiaoyi Wu , Tian Lan , Bin Li

Volumetric medical imaging technologies produce detailed 3D representations of anatomical structures. However, effective medical data visualization and exploration pose significant challenges, especially for individuals with limited medical…

人机交互 · 计算机科学 2025-07-01 Qixuan Liu , Shi Qiu , Yinqiao Wang , Xiwen Wu , Kenneth Siu Ho Chok , Chi-Wing Fu , Pheng-Ann Heng

Site-specific outdoor AR experiences are typically authored using static 3D models, but are deployed in physical environments that change over time. As a result, virtual content may become misaligned with its intended real-world referents,…

人机交互 · 计算机科学 2025-08-12 Nels Numan , Jessica Van Brummelen , Ziwen Lu , Anthony Steed

The fusion of language and vision in large vision-language models (LVLMs) has revolutionized deep learning-based object detection by enhancing adaptability, contextual reasoning, and generalization beyond traditional architectures. This…

计算机视觉与模式识别 · 计算机科学 2025-10-01 Ranjan Sapkota , Manoj Karkee

As XR devices become widespread, 3D interaction has become commonplace, and UI developers are increasingly required to consider usability to deliver better user experiences. The HCI community has long studied target-pointing performance,…

人机交互 · 计算机科学 2026-01-08 Tatsuya Okuno , Haruto Shimizu , Nobuhito Kasahara , Taiyu Honma , Shota Yamanaka , Homei Miyashita

Extended Reality (XR) is a rapidly growing field offering unique immersive experiences, social networking, learning, and collaboration opportunities. The continuous advancements in XR technology and industry efforts are gradually moving…

人机交互 · 计算机科学 2023-08-23 Pascal Knierim , Thomas Kosch

We introduce V-Agent, a novel multi-agent platform designed for advanced video search and interactive user-system conversations. By fine-tuning a vision-language model (VLM) with a small video preference dataset and enhancing it with a…

计算机视觉与模式识别 · 计算机科学 2026-01-08 SunYoung Park , Jong-Hyeon Lee , Youngjune Kim , Daegyu Sung , Younghyun Yu , Young-rok Cha , Jeongho Ju

Advances in artificial intelligence and human-computer interaction will likely lead to extended reality (XR) becoming pervasive. While XR can provide users with interactive, engaging, and immersive experiences, non-player characters are…

人机交互 · 计算机科学 2024-06-21 Efe Bozkir , Süleyman Özdel , Ka Hei Carrie Lau , Mengdi Wang , Hong Gao , Enkelejda Kasneci

Search engines enable the retrieval of unknown information with texts. However, traditional methods fall short when it comes to understanding unfamiliar visual content, such as identifying an object that the model has never seen before.…

计算机视觉与模式识别 · 计算机科学 2024-10-29 Zhixin Zhang , Yiyuan Zhang , Xiaohan Ding , Xiangyu Yue