English
Related papers

Related papers: From Instruction to Event: Sound-Triggered Mobile …

200 papers

We model acoustic dynamics in space and time from synthetic sensor data. The tasks are (i) to predict and extrapolate the spatiotemporal dynamics, and (ii) reconstruct the acoustic state from partial observations. To achieve this, we…

Fluid Dynamics · Physics 2024-11-12 Defne Ege Ozan , Luca Magri

One of the long-term goals of artificial intelligence is to build an agent that can communicate intelligently with human in natural language. Most existing work on natural language learning relies heavily on training over a pre-collected…

Computation and Language · Computer Science 2017-05-30 Haichao Zhang , Haonan Yu , Wei Xu

In an era of human-computer interaction with increasingly agentic AI systems capable of connecting with users conversationally, speech is an important modality for commanding agents. By recognizing and using speech emotions (i.e., how a…

Human-Computer Interaction · Computer Science 2025-04-14 Ilhan Aslan , Timothy Merritt , Stine S. Johansen , Niels van Berkel

Obtaining reliable feedback from the environment is a fundamental capability for intelligent agents to evaluate the correctness of their actions and to accumulate reusable knowledge. However, most existing approaches rely on predefined…

Artificial Intelligence · Computer Science 2026-01-09 Hong Su

Continuous advancements in robotics and AI are driving the integration of robots from industry into everyday environments. However, dynamic and unpredictable human activities in daily lives would directly or indirectly conflict with robot…

Robotics · Computer Science 2025-09-08 Dongping Li , Shaoting Peng , John Pohovey , Katherine Rose Driggs-Campbell

Multi-agent models often describe populations segregated either in the physical space, i.e. subdivided in metapopulations, or in the ecology of opinions, i.e. partitioned in echo chambers. Here we show how the interplay between homophily…

Physics and Society · Physics 2017-02-23 Michele Starnini , Mattia Frasca , Andrea Baronchelli

Conversational agents have traditionally been developed for either task-oriented dialogue (TOD) or open-ended chitchat, with limited progress in unifying the two. Yet, real-world conversations naturally involve fluid transitions between…

Computation and Language · Computer Science 2025-11-13 Yejin Yoon , Yuri Son , Namyoung So , Minseo Kim , Minsoo Cho , Chanhee Park , Seungshin Lee , Taeuk Kim

Mobile agents that can leverage help from humans can potentially accomplish more complex tasks than they could entirely on their own. We develop "Help, Anna!" (HANNA), an interactive photo-realistic simulator in which an agent fulfills…

Human-Computer Interaction · Computer Science 2019-11-25 Khanh Nguyen , Hal Daumé

Agent-based IoT applications have recently been proposed in several domains, such as health care, smart cities and agriculture. Deploying these applications in specific settings has been very challenging for many reasons including the…

Multiagent Systems · Computer Science 2018-02-13 Nathalia Nascimento , Paulo Alencar , Carlos Lucena , Donald Cowan

Building a persona-based conversation agent is challenging owing to the lack of large amounts of speaker-specific conversation data for model training. This paper addresses the problem by proposing a multi-task learning approach to training…

Computation and Language · Computer Science 2017-10-23 Yi Luan , Chris Brockett , Bill Dolan , Jianfeng Gao , Michel Galley

The rapid advances in audio analysis underscore its vast potential for humancomputer interaction, environmental monitoring, and public safety; yet, existing audioonly datasets often lack spatial context. To address this gap, we present two…

Sound · Computer Science 2025-12-10 Shuaihang Yuan , Congcong Wen , Muhammad Shafique , Anthony Tzes , Yi Fang

Audio-visual navigation enables embodied agents to navigate toward sound-emitting targets by leveraging both auditory and visual cues. However, most existing approaches rely on precomputed room impulse responses (RIRs) for binaural audio…

Computer Vision and Pattern Recognition · Computer Science 2026-04-02 Yichen Zeng , Hebaixu Wang , Meng Liu , Yu Zhou , Chen Gao , Kehan Chen , Gongping Huang

Generating long-form audio-visual stories from a short user prompt remains challenging due to an intent-execution gap, where high-level narrative intent must be preserved across coherent, shot-level multimodal generation over long horizons.…

Computer Vision and Pattern Recognition · Computer Science 2026-02-04 Wenzhang Sun , Zhenyu Wang , Zhangchi Hu , Chunfeng Wang , Hao Li , Wei Chen

Haptic interfaces have untapped the sense of touch to assist multimodal music learning. We have recently seen various improvements of interface design on tactile feedback and force guidance aiming to make instrument learning more effective.…

Human-Computer Interaction · Computer Science 2019-06-05 Yian Zhang , Yinmiao Li , Daniel Chin , Gus Xia

Acoustic events are sounds with well-defined spectro-temporal characteristics which can be associated with the physical objects generating them. Acoustic scenes are collections of such acoustic events in no specific temporal order. Given…

Sound · Computer Science 2022-06-28 Rahil Parikh , Harshavardhan Sundar , Ming Sun , Chao Wang , Spyros Matsoukas

Recent advances in duplex speech models have enabled natural, low-latency speech-to-speech interactions. However, existing models are restricted to a fixed role and voice, limiting their ability to support structured, role-driven real-world…

Computation and Language · Computer Science 2026-02-09 Rajarshi Roy , Jonathan Raiman , Sang-gil Lee , Teodor-Dumitru Ene , Robert Kirby , Sungwon Kim , Jaehyeon Kim , Bryan Catanzaro

Humans can robustly recognize and localize objects by using visual and/or auditory cues. While machines are able to do the same with visual data already, less work has been done with sounds. This work develops an approach for scene…

Sound · Computer Science 2022-03-01 Dengxin Dai , Arun Balajee Vasudevan , Jiri Matas , Luc Van Gool

Automatic audio event recognition plays a pivotal role in making human robot interaction more closer and has a wide applicability in industrial automation, control and surveillance systems. Audio event is composed of intricate phonic…

Audio and Speech Processing · Electrical Eng. & Systems 2023-04-12 Tushar Sandhan , Sukanya Sonowal , Jin Young Choi

Understanding how humans control unstable systems is central to many research problems, with applications ranging from quiet standing to aircraft landing. Increasingly much evidence appears in favor of event-driven control hypothesis: human…

Biological Physics · Physics 2014-06-17 Arkady Zgonnikov , Ihor Lubashevsky , Shigeru Kanemoto , Toru Miyazawa , Takashi Suzuki

Risk management resulting from the actions and states of the different elements making up a operating room is a major concern during a surgical procedure. Agent-based simulation shows an interest through its interaction concepts,…

Artificial Intelligence · Computer Science 2020-07-23 Bruno Perez , Julien Henriet , Christophe Lang , Laurent Philippe