中文
相关论文

相关论文: SARAH: Spatially Aware Real-time Agentic Humans

200 篇论文

We introduce a speech-guided embodied agent framework for video-guided skull base surgery that dynamically executes perception and image-guidance tasks in response to surgeon queries. The proposed system integrates natural language…

计算机视觉与模式识别 · 计算机科学 2026-04-17 Jecia Z. Y. Mao , Francis X. Creighton , Russell H. Taylor , Manish Sahu

Progress in Embodied AI has made it possible for end-to-end-trained agents to navigate in photo-realistic environments with high-level reasoning and zero-shot or language-conditioned behavior, but benchmarks are still dominated by…

We present SPEAR, a continuous receiver-to-receiver acoustic neural warping field for spatial acoustic effects prediction in an acoustic 3D space with a single stationary audio source. Unlike traditional source-to-receiver modelling methods…

声音 · 计算机科学 2024-06-18 Yuhang He , Shitong Xu , Jia-Xing Zhong , Sangyun Shin , Niki Trigoni , Andrew Markham

Despite significant advances in talking avatar generation, existing methods face critical challenges: insufficient text-following capability for diverse actions, lack of temporal alignment between actions and audio content, and dependency…

计算机视觉与模式识别 · 计算机科学 2026-01-21 Ziqiao Peng , Yi Chen , Yifeng Ma , Guozhen Zhang , Zhiyao Sun , Zixiang Zhou , Youliang Zhang , Zhengguang Zhou , Zhaoxin Fan , Hongyan Liu , Yuan Zhou , Qinglin Lu , Jun He

High-fidelity physics simulation is essential for scalable robotic learning, but the sim-to-real gap persists, especially for tasks involving complex, dynamic, and discontinuous interactions like physical contacts. Explicit system…

机器人学 · 计算机科学 2026-01-21 Changwei Jing , Jai Krishna Bandi , Jianglong Ye , Yan Duan , Pieter Abbeel , Xiaolong Wang , Sha Yi

In this paper, we present the design of a multimodal interaction framework for intelligent virtual agents in wearable mixed reality environments, especially for interactive applications at museums, botanical gardens, and similar places.…

人机交互 · 计算机科学 2025-03-26 Ghazanfar Ali , Hong-Quan Le , Junho Kim , Seoung-won Hwang , Jae-In Hwang

We present Roomify, a spatially-grounded transformation system that generates themed virtual environments anchored to users' physical rooms while maintaining spatial structure and functional semantics. Current VR approaches face a…

人机交互 · 计算机科学 2026-03-06 Xueyang Wang , Qinxuan Cen , Weitao Bi , Yunxiang Ma , Xin Yi , Robert Xiao , Xinyi Fu , Hewu Li

Audio-driven video generation aims to synthesize realistic videos that align with input audio recordings, akin to the human ability to visualize scenes from auditory input. However, existing approaches predominantly focus on exploring…

图形学 · 计算机科学 2026-03-17 Kien T. Pham , Yingqing He , Yazhou Xing , Qifeng Chen , Long Chen

Joint Attention (JA), a crucial social skill for developing shared focus, is often impaired in children with Autism Spectrum Disorder (ASD), affecting social communication and highlighting the need for early intervention. Addressing gaps in…

Augmented Reality (AR) systems, while enhancing task performance through real-time guidance, pose risks of inducing cognitive tunneling-a hyperfocus on virtual content that compromises situational awareness (SA) in safety-critical…

机器学习 · 计算机科学 2025-09-03 Zhehan Qu , Tianyi Hu , Christian Fronk , Maria Gorlatova

Human language learners are exposed to a trickle of informative, context-sensitive language, but a flood of raw sensory data. Through both social language use and internal processes of rehearsal and practice, language learners are able to…

In Extended Reality (XR), rendering sound that accurately simulates real-world acoustics is pivotal in creating lifelike and believable virtual experiences. However, existing XR spatial audio rendering methods often struggle with real-time…

Embodied Conversational Agents that make use of co-speech gestures can enhance human-machine interactions in many ways. In recent years, data-driven gesture generation approaches for ECAs have attracted considerable research attention, and…

人机交互 · 计算机科学 2022-10-14 Yuan He , André Pereira , Taras Kucherenko

Locating a target based on auditory and visual cues$\unicode{x2013}$such as finding a car in a crowded parking lot or identifying a speaker in a virtual meeting$\unicode{x2013}$requires balancing effort, time, and accuracy under…

人机交互 · 计算机科学 2026-02-04 Hyunsung Cho , Xuejing Luo , Byungjoo Lee , David Lindlbauer , Antti Oulasvirta

Humans can robustly recognize and localize objects by integrating visual and auditory cues. While machines are able to do the same now with images, less work has been done with sounds. This work develops an approach for dense semantic…

计算机视觉与模式识别 · 计算机科学 2020-03-10 Arun Balajee Vasudevan , Dengxin Dai , Luc Van Gool

Audio-visual embodied navigation, as a hot research topic, aims training a robot to reach an audio target using egocentric visual (from the sensors mounted on the robot) and audio (emitted from the target) input. The audio-visual…

声音 · 计算机科学 2022-10-06 Yinfeng Yu , Lele Cao , Fuchun Sun , Xiaohong Liu , Liejun Wang

We present EmbodiedHead, a speech-driven talking-head framework that equips LLMs with real-time visual avatars for conversation. A practical embodied avatar must achieve real-time generation, unified listening-speaking behavior, and high…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Yu Zhang , Kaiyuan Shen , Yang Li

Skeletal Action recognition from an egocentric view is important for applications such as interfaces in AR/VR glasses and human-robot interaction, where the device has limited resources. Most of the existing skeletal action recognition…

计算机视觉与模式识别 · 计算机科学 2023-09-20 Junan Lin , Zhichao Sun , Enjie Cao , Taein Kwon , Mahdi Rad , Marc Pollefeys

Modeling spatial-temporal interactions among neighboring agents is at the heart of multi-agent problems such as motion forecasting and crowd navigation. Despite notable progress, it remains unclear to which extent modern representations can…

机器学习 · 计算机科学 2025-06-12 Ahmad Rahimi , Po-Chien Luan , Yuejiang Liu , Frano Rajič , Alexandre Alahi

Interaction between humanoids involves bidirectional and nonverbal reactivity, coordination and synchrony. Toward socially aware robots and interactive virtual agents, we present SalsaAgent, a language model that generates expressive,…

计算机视觉与模式识别 · 计算机科学 2026-05-29 Payam Jome Yazdian , Zoe Stanley , Angelica Lim