中文
相关论文

相关论文: Voice-Interactive Surgical Agent for Multimodal Pa…

200 篇论文

Deriving personalized insights from popular wearable trackers requires complex numerical reasoning that challenges standard LLMs, necessitating tool-based approaches like code generation. Large language model (LLM) agents present a…

Wearable Augmented Reality (AR) technologies are gaining recognition for their potential to transform surgical navigation systems. As these technologies evolve, selecting the right interaction method to control the system becomes crucial.…

人机交互 · 计算机科学 2024-12-25 Hamraz Javaheri , Omid Ghamarnejad , Paul Lukowicz , Gregor Alexander Stavrou , Jakob Karolus

Vision-language-action models (VLAs) have garnered significant attention for their potential in advancing robotic manipulation. However, previous approaches predominantly rely on the general comprehension capabilities of vision-language…

计算机视觉与模式识别 · 计算机科学 2025-06-25 Yuqi Wang , Xinghang Li , Wenxuan Wang , Junbo Zhang , Yingyan Li , Yuntao Chen , Xinlong Wang , Zhaoxiang Zhang

Remote sighted assistance (RSA) has emerged as a conversational assistive technology, where remote sighted workers, i.e., agents, provide real-time assistance to users with vision impairments via video-chat-like communication. Researchers…

人机交互 · 计算机科学 2022-02-04 Jingyi Xie , Rui Yu , Sooyeon Lee , Yao Lyu , Syed Masum Billah , John M. Carroll

Training mental health clinicians to conduct standardized clinical assessments is challenging due to a lack of scalable, realistic practice opportunities, which can impact data quality in clinical trials. To address this gap, we introduce a…

人机交互 · 计算机科学 2025-12-30 Veronica Bossio Botero , Vijay Yadav , Jacob Ouyang , Anzar Abbas , Michelle Worthington

In recent years, Visual Question Localized-Answering in robotic surgery (Surgical-VQLA) has gained significant attention for its potential to assist medical students and junior doctors in understanding surgical scenes. Recently, the rapid…

计算机视觉与模式识别 · 计算机科学 2025-09-23 Pengfei Hao , Hongqiu Wang , Shuaibo Li , Zhaohu Xing , Guang Yang , Kaishun Wu , Lei Zhu

Large vision language models (VLMs) have demonstrated significant potential for integration into daily life, making it crucial for them to incorporate human values when making decisions in real-world situations. This paper introduces VIVA,…

计算与语言 · 计算机科学 2024-10-11 Zhe Hu , Yixiao Ren , Jing Li , Yu Yin

Large language models (LLMs) and vision-language models (VLMs) have the potential to transform biological research by enabling autonomous experimentation. Yet, their application remains constrained by rigid protocol design, limited…

机器人学 · 计算机科学 2025-07-03 Yibo Qiu , Zan Huang , Zhiyu Wang , Handi Liu , Yiling Qiao , Yifeng Hu , Shu'ang Sun , Hangke Peng , Ronald X Xu , Mingzhai Sun

Agentic visual analytics (VA) represents an emerging class of systems in which large language model (LLM)-driven agents autonomously plan, execute, evaluate, and iterate across the full visual analytics pipeline. By shifting users from…

数据库 · 计算机科学 2026-04-20 Tianqi Luo , Leixian Shen , Yuyu Luo

We introduce iFlyBot-VLA, a large-scale Vision-Language-Action (VLA) model trained under a novel framework. The main contributions are listed as follows: (1) a latent action model thoroughly trained on large-scale human and robotic…

计算机视觉与模式识别 · 计算机科学 2025-11-05 Yuan Zhang , Chenyu Xue , Wenjie Xu , Chao Ji , Jiajia wu , Jia Pan

Industrial workflows demand adaptive and trustworthy assistance that can operate under limited computing, connectivity, and strict privacy constraints. In this work, we present MICA (Multi-Agent Industrial Coordination Assistant), a…

人工智能 · 计算机科学 2026-03-10 Di Wen , Kunyu Peng , Junwei Zheng , Yufan Chen , Yitian Shi , Jiale Wei , Ruiping Liu , Kailun Yang , Rainer Stiefelhagen

Procedural activity assistants potentially support humans in a variety of settings, from our daily lives, e.g., cooking or assembling flat-pack furniture, to professional situations, e.g., manufacturing or biological experiments. Despite…

计算与语言 · 计算机科学 2025-10-02 Kimihiro Hasegawa , Wiradee Imrattanatrai , Masaki Asada , Ken Fukuda , Teruko Mitamura

Radiology visual question answering (RVQA) provides precise answers to questions about chest X-ray images, alleviating radiologists' workload. While recent methods based on multimodal large language models (MLLMs) and retrieval-augmented…

人工智能 · 计算机科学 2025-08-06 Ziruo Yi , Jinyu Liu , Ting Xiao , Mark V. Albert

Data science and engineering workflows often span multiple stages, from warehousing to orchestration, using tools like BigQuery, dbt, and Airbyte. As vision language models (VLMs) advance in multimodal understanding and code generation,…

Data scarcity remains a fundamental barrier to achieving fully autonomous surgical robots. While large scale vision language action (VLA) models have shown impressive generalization in household and industrial manipulation by leveraging…

Multimodal artificial intelligence (AI) systems have the potential to enhance clinical decision-making by interpreting various types of medical data. However, the effectiveness of these models across all medical fields is uncertain. Each…

Recent advances in AI combine large language models (LLMs) with vision encoders that bring forward unprecedented technical capabilities to leverage for a wide range of healthcare applications. Focusing on the domain of radiology,…

Recent advances in Multimodal Large Language Models (MLLMs) have driven rapid progress in Vision-Language-Action (VLA) models for robotic manipulation. Although effective in many scenarios, current approaches largely rely on explicit…

In dynamic environments such as warehouses, hospitals, and homes, robots must seamlessly transition between gross motion and precise manipulations to complete complex tasks. However, current Vision-Language-Action (VLA) frameworks, largely…

机器人学 · 计算机科学 2026-03-03 Xiongfeng Peng , Jiaqian Yu , Dingzhe Li , Yixiang Jin , Lu Xu , Yamin Mao , Chao Zhang , Weiming Li , Sujin Jang , Dongwook Lee , Daehyun Ji

Robotic manipulation in open-world environments requires reasoning across semantics, geometry, and long-horizon action dynamics. Existing hierarchical Vision-Language-Action (VLA) frameworks typically use 2D representations to connect…

机器人学 · 计算机科学 2026-03-17 You Wu , Zixuan Chen , Cunxu Ou , Wenxuan Wang , Wenbo Huang , Lin Cao , Yangtao Chen , Weichao Qiu , Xingyue Quan , Jieqi Shi , Jing Huo , Yang Gao