English
Related papers

Related papers: Input-Envelope-Output: Auditable Generative Music …

200 papers

Recently, diffusion models have achieved great success in mono-channel audio generation. However, when it comes to stereo audio generation, the soundscapes often have a complex scene of multiple objects and directions. Controlling stereo…

Sound · Computer Science 2025-02-26 Peiwen Sun , Sitong Cheng , Xiangtai Li , Zhen Ye , Huadai Liu , Honggang Zhang , Wei Xue , Yike Guo

Diffusion models have emerged as powerful deep generative techniques, producing high-quality and diverse samples in applications in various domains including audio. While existing reviews provide overviews, there remains limited in-depth…

Sound · Computer Science 2026-01-16 Ge Zhu , Yutong Wen , Zhiyao Duan

Recent progress in auditory intelligence has yielded high-performing systems for sound event detection (SED), acoustic scene classification (ASC), automated audio captioning (AAC), and audio question answering (AQA). Yet these tasks remain…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-12 Hyeonuk Nam

Since 2023, generative AI has rapidly advanced in the music domain. Despite significant technological advancements, music-generative models raise critical ethical challenges, including a lack of transparency and accountability, along with…

AI-enabled systems are increasingly introduced into educational contexts, yet their effectiveness depends less on technological sophistication than on the quality of pedagogical mediation, ethical constraints, and context-sensitive design.…

Human-Computer Interaction · Computer Science 2026-05-21 Pier Paolo Benedetti

This paper addresses the problem of safe autonomous navigation in unknown obstacle-filled environments using only local sensory information. We propose a smooth feedback controller derived from an unconstrained penalty-based formulation…

Systems and Control · Electrical Eng. & Systems 2025-11-14 Lyes Smaili , Soulaimane Berkane

Audio has rapidly become a primary interface for foundation models, powering real-time voice assistants. Ensuring safety in audio systems is inherently more complex than just "unsafe text spoken aloud": real-world risks can hinge on…

Sound · Computer Science 2026-04-13 Mintong Kang , Chen Fang , Bo Li

Generating long-form audio-visual stories from a short user prompt remains challenging due to an intent-execution gap, where high-level narrative intent must be preserved across coherent, shot-level multimodal generation over long horizons.…

Computer Vision and Pattern Recognition · Computer Science 2026-02-04 Wenzhang Sun , Zhenyu Wang , Zhangchi Hu , Chunfeng Wang , Hao Li , Wei Chen

Despite significant advancements in neural text-to-audio generation, challenges persist in controllability and evaluation. This paper addresses these issues through the Sound Scene Synthesis challenge held as part of the Detection and…

In the context of environmental sound classification, the adaptability of systems is key: which sound classes are interesting depends on the context and the user's needs. Recent advances in text-to-audio retrieval allow for zero-shot audio…

Sound · Computer Science 2023-08-21 Saksham Singh Kushwaha , Magdalena Fuentes

When the disturbance input matrix is nonlinear, existing disturbance observer design methods rely on the solvability of a partial differential equation or the existence of an output function with a uniformly well-defined disturbance…

Systems and Control · Electrical Eng. & Systems 2024-06-21 Yujie Wang , Xiangru Xu

Moving around in the world is naturally a multisensory experience, but today's embodied agents are deaf---restricted to solely their visual perception of the environment. We introduce audio-visual navigation for complex, acoustically and…

Computer Vision and Pattern Recognition · Computer Science 2020-08-25 Changan Chen , Unnat Jain , Carl Schissler , Sebastia Vicenc Amengual Gari , Ziad Al-Halah , Vamsi Krishna Ithapu , Philip Robinson , Kristen Grauman

Recent advances in the Active Speaker Detection (ASD) problem build upon a two-stage process: feature extraction and spatio-temporal context aggregation. In this paper, we propose an end-to-end ASD workflow where feature learning and…

Computer Vision and Pattern Recognition · Computer Science 2022-07-26 Juan Leon Alcazar , Moritz Cordes , Chen Zhao , Bernard Ghanem

Autonomous soundscape augmentation systems typically use trained models to pick optimal maskers to effect a desired perceptual change. While acoustic information is paramount to such systems, contextual information, including participant…

Sound · Computer Science 2024-07-03 Kenneth Ooi , Karn N. Watcharasupat , Bhan Lam , Zhen-Ting Ong , Woon-Seng Gan

The generation of sounding videos has seen significant advancements with the advent of diffusion models. However, existing methods often lack the fine-grained control needed to generate viewpoint-specific content from larger, immersive…

Current mobile manipulation research predominantly follows an instruction-driven paradigm, where agents rely on predefined textual commands to execute tasks. However, this setting confines agents to a passive role, limiting their autonomy…

Robotics · Computer Science 2026-04-16 Hao Ju , Shaofei Huang , Hongyu Li , Zihan Ding , Si Liu , Meng Wang , Zhedong Zheng

Visually impaired people face significant challenges when attempting to interact with and understand complex environments, and traditional assistive technologies often struggle to quickly provide necessary contextual understanding and…

Human-Computer Interaction · Computer Science 2025-05-02 Bhanuja Ainary

In recent years the automotive industry has been strongly promoting the development of smart cars, equipped with multi-modal sensors to gather information about the surroundings, in order to aid human drivers or make autonomous decisions.…

Audio and Speech Processing · Electrical Eng. & Systems 2023-01-31 Jun Yin , Stefano Damiano , Marian Verhelst , Toon van Waterschoot , Andre Guntoro

Despite advances in text and visual generation, creating coherent long-form audio narratives remains challenging. Existing frameworks often exhibit limitations such as mismatched character settings with voice performance, insufficient…

Sound · Computer Science 2026-05-21 Yiming Ren , Xuenan Xu , Ziyang Zhang , Wen Wu , Baoxiang Li , Chao Zhang

As artificial agents display increasingly sophisticated emotion-like behaviors, frameworks for assessing whether such systems risk instantiating consciousness remain limited. This contribution asks whether synthetic emotion-like control can…

Artificial Intelligence · Computer Science 2026-03-05 Hermann Borotschnig