English
Related papers

Related papers: A Multi-Agent AI Framework for Immersive Audiobook…

200 papers

Large Audio Language Models (LALMs) excel at perception but struggle with complex reasoning requiring precise acoustic measurements. While external tools can extract fine-grained features like exact tempo or pitch, effective integration…

Sound · Computer Science 2026-02-17 Siqian Tong , Xuan Li , Yiwei Wang , Baolong Bi , Yujun Cai , Shenghua Liu , Yuchen He , Chengpeng Hao

Human auditory perception is shaped by moving sound sources in 3D space, yet prior work in generative sound modelling has largely been restricted to mono signals or static spatial audio. In this work, we introduce a framework for generating…

Sound · Computer Science 2025-09-29 Yunyi Liu , Shaofan Yang , Kai Li , Xu Li

The rise in capability and ubiquity of generative artificial intelligence (AI) technologies has enabled its application to the field of Socially Interactive Agents (SIAs). Despite rising interest in modern AI-powered components used for…

Human-Computer Interaction · Computer Science 2024-10-29 Spencer Lin , Basem Rizk , Miru Jun , Andy Artze , Caitlin Sullivan , Sharon Mozgai , Scott Fisher

Language is a cornerstone of cultural identity, yet globalization and the dominance of major languages have placed nearly 3,000 languages at risk of extinction. Existing AI-driven translation models prioritize efficiency but often fail to…

Computation and Language · Computer Science 2025-06-10 Mahfuz Ahmed Anik , Abdur Rahman , Azmine Toushik Wasi , Md Manjurul Ahsan

With the advancement of generative models, the synthesis of different sensory elements such as music, visuals, and speech has achieved significant realism. However, the approach to generate multi-sensory outputs has not been fully explored,…

Computer Vision and Pattern Recognition · Computer Science 2024-08-22 Minheng Ni , Chenfei Wu , Huaying Yuan , Zhengyuan Yang , Ming Gong , Lijuan Wang , Zicheng Liu , Wangmeng Zuo , Nan Duan

Employing voice-based emotion recognition function in artificial intelligence (AI) product will improve the user experience. Most of researches that have been done only focus on the speech collected under controlled conditions. The…

Audio and Speech Processing · Electrical Eng. & Systems 2018-03-06 Fei Tao , Gang Liu , Qingen Zhao

In architectural interior design, miscommunication frequently arises as clients lack design knowledge, while designers struggle to explain complex spatial relationships, leading to delayed timelines and financial losses. Recent advancements…

Artificial Intelligence · Computer Science 2026-03-17 Ren Jian Lim , Rushi Dai

Diffusion models have emerged as powerful deep generative techniques, producing high-quality and diverse samples in applications in various domains including audio. While existing reviews provide overviews, there remains limited in-depth…

Sound · Computer Science 2026-01-16 Ge Zhu , Yutong Wen , Zhiyao Duan

Every individual carries a unique and personal life story shaped by their memories and experiences. However, these memories are often scattered and difficult to organize into a coherent narrative, a challenge that defines the task of…

Human-Computer Interaction · Computer Science 2025-09-30 Shayan Talaei , Meijin Li , Kanu Grover , James Kent Hippler , Diyi Yang , Amin Saberi

We tackle a task where an agent learns to navigate in a 2D maze-like environment called XWORLD. In each session, the agent perceives a sequence of raw-pixel frames, a natural language command issued by a teacher, and a set of rewards. The…

Computation and Language · Computer Science 2017-05-23 Haonan Yu , Haichao Zhang , Wei Xu

Deep Audio Analyzer is an open source speech framework that aims to simplify the research and the development process of neural speech processing pipelines, allowing users to conceive, compare and share results in a fast and reproducible…

Sound · Computer Science 2023-10-31 Valerio Francesco Puglisi , Oliver Giudice , Sebastiano Battiato

Developing algorithms for sound classification, detection, and localization requires large amounts of flexible and realistic audio data, especially when leveraging modern machine learning and beamforming techniques. However, most existing…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-23 Luca Barbisan , Marco Levorato , Fabrizio Riente

Sound event localization frameworks based on deep neural networks have shown increased robustness with respect to reverberation and noise in comparison to classical parametric approaches. In particular, recurrent architectures that…

This paper investigates the quality of multi-agent dialogues in simulations powered by Large Language Models (LLMs). Analyzing dialogues and memory over multiple sessions revealed significant issues such as repetition, inconsistency, and…

Computation and Language · Computer Science 2024-08-13 KuanChao Chu , Yi-Pei Chen , Hideki Nakayama

This paper presents an audio visual automatic speech recognition (AV-ASR) system using a Transformer-based architecture. We particularly focus on the scene context provided by the visual information, to ground the ASR. We extract…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-01 Georgios Paraskevopoulos , Srinivas Parthasarathy , Aparna Khare , Shiva Sundaram

Dialogue models falter in noisy, multi-speaker environments, often producing irrelevant responses and awkward turn-taking. We present AV-Dialog, the first multimodal dialog framework that uses both audio and visual cues to track the target…

Computation and Language · Computer Science 2025-11-17 Tuochao Chen , Bandhav Veluri , Hongyu Gong , Shyamnath Gollakota

This study addresses the challenge that generative models struggle to balance flexibility, stability, and controllability in complex interactive scenarios. It proposes a controllable generation framework for dynamic interactive content…

Human-Computer Interaction · Computer Science 2026-02-27 Rui Liu

Proactive AR agents promise context-aware assistance, but their interactions often rely on explicit voice prompts or responses, which can be disruptive or socially awkward. We introduce Sensible Agent, a framework designed for unobtrusive…

Human-Computer Interaction · Computer Science 2025-11-03 Geonsun Lee , Min Xia , Nels Numan , Xun Qian , David Li , Yanhe Chen , Achin Kulshrestha , Ishan Chatterjee , Yinda Zhang , Dinesh Manocha , David Kim , Ruofei Du

Audio-visual navigation of an agent towards locating an audio goal is a challenging task especially when the audio is sporadic or the environment is noisy. In this paper, we present CAVEN, a Conversation-based Audio-Visual Embodied…

Computer Vision and Pattern Recognition · Computer Science 2023-12-29 Xiulong Liu , Sudipta Paul , Moitreya Chatterjee , Anoop Cherian

Agentic AI systems use specialized agents to handle tasks within complex workflows, enabling automation and efficiency. However, optimizing these systems often requires labor-intensive, manual adjustments to refine roles, tasks, and…

Computation and Language · Computer Science 2024-12-24 Kamer Ali Yuksel , Hassan Sawaf