English
Related papers

Related papers: Input-Envelope-Output: Auditable Generative Music …

200 papers

We introduce ImmerseDiffusion, an end-to-end generative audio model that produces 3D immersive soundscapes conditioned on the spatial, temporal, and environmental conditions of sound objects. ImmerseDiffusion is trained to generate…

Sound · Computer Science 2025-02-11 Mojtaba Heydari , Mehrez Souden , Bruno Conejo , Joshua Atkins

With the development of AI-Generated Content (AIGC), text-to-audio models are gaining widespread attention. However, it is challenging for these models to generate audio aligned with human preference due to the inherent information density…

Sound · Computer Science 2024-02-02 Huan Liao , Haonan Han , Kai Yang , Tianjiao Du , Rui Yang , Zunnan Xu , Qinmei Xu , Jingquan Liu , Jiasheng Lu , Xiu Li

Zero-shot learning enables models to generalise to unseen classes by leveraging semantic information, bridging the gap between training and testing sets with non-overlapping classes. While much research has focused on zero-shot learning in…

Sound · Computer Science 2025-07-03 Ysobel Sims , Alexandre Mendes , Stephan Chalup

Text-to-3D generative AI systems create navigable environments from natural language prompts, but unlike text-to-image generation, evaluation requires embodied exploration of spatial coherence, scale, and navigability. We present the first…

Human-Computer Interaction · Computer Science 2026-03-17 Aung Pyae

AI agents that build user interfaces on the fly assembling buttons, forms, and data displays from structured protocol payloads are becoming common in production systems. The trouble is that a payload can pass every schema check and still…

Artificial Intelligence · Computer Science 2026-03-06 Mohd Safwan Uddin , Saba Hajira

End-to-end (E2E) spoken dialogue systems are increasingly replacing cascaded pipelines for voice-based human-AI interaction, processing raw audio directly without intermediate transcription. Existing benchmarks primarily evaluate these…

The diagnosis of Autism Spectrum Disorder (ASD) in children is commonly accompanied by a diagnosis of sensory processing disorders as well. Abnormalities are usually reported in multiple sensory processing domains, showing a higher…

Robotics · Computer Science 2019-01-07 Hifza Javed , Rachael Burns , Myounghoon Jeon , Ayanna M. Howard , Chung Hyuk Park

We address the problem of detecting initial system--environment correlations when the environment is not directly accessible. Most existing approaches rely on full state tomography or multiple system preparations, which can be…

Quantum Physics · Physics 2026-02-24 Ali Abu-Nada , Russell Ceballos , Lian-Ao Wu

Generative AI, specifically text-to-image models, have revolutionized interior architectural design by enabling the rapid translation of conceptual ideas into visual representations from simple text prompts. While generative AI can produce…

Human-Computer Interaction · Computer Science 2025-06-19 Richa Gupta , Alexander Htet Kyaw

Audio is indispensable for real-world video, yet generation models have largely overlooked audio components. Current approaches to producing audio-visual content often rely on cascaded pipelines, which increase cost, accumulate errors, and…

Imagine interconnected objects with embedded artificial intelligence (AI), empowered to sense the environment, see it, hear it, touch it, interact with it, and move. As future networks of intelligent objects come to life, tremendous new…

The extended state observer (ESO) plays an important role in the design of feedback control for nonlinear systems. However, its high-gain nature creates a challenge in engineering practice in cases where the output measurement is corrupted…

Systems and Control · Electrical Eng. & Systems 2020-09-30 Krzysztof Łakomy , Rafal Madonski

A major focus of recent research in spoken language understanding (SLU) has been on the end-to-end approach where a single model can predict intents directly from speech inputs without intermediate transcripts. However, this approach…

Computation and Language · Computer Science 2021-06-15 Sujeong Cha , Wangrui Hou , Hyun Jung , My Phung , Michael Picheny , Hong-Kwang Kuo , Samuel Thomas , Edmilson Morais

Recent advances in text-to-audio generation enable models to translate natural-language descriptions into diverse musical output. However, the robustness of these systems under semantically equivalent prompt variations remains largely…

Sound · Computer Science 2026-05-06 Jiahui Wu

AI systems for music generation are increasingly common and easy to use, granting people without any musical background the ability to create music. Because of this, generative-AI has been marketed and celebrated as a means of democratizing…

Sound · Computer Science 2025-08-13 Liam Pram , Fabio Morreale

This study addresses the challenge that generative models struggle to balance flexibility, stability, and controllability in complex interactive scenarios. It proposes a controllable generation framework for dynamic interactive content…

Human-Computer Interaction · Computer Science 2026-02-27 Rui Liu

Generative design, an AI-assisted technology for optimizing design through algorithmic processes, is propelling advancements across numerous fields. As the use of immersive environments such as Augmented Reality (AR) continues to rise,…

Human-Computer Interaction · Computer Science 2025-03-28 Sora Kang , Kaiwen Yu , Xinyi Zhou , Joonhwan Lee

Earable acoustic sensing offers a powerful and non-invasive modality for capturing fine-grained auditory and physiological signals directly from the ear canal, enabling continuous and context-aware monitoring of cognitive states. As earable…

Human-Computer Interaction · Computer Science 2025-12-23 Xijia Wei , Ting Dang , Khaldoon Al-Naimi , Yang Liu , Fahim Kawsar , Alessandro Montanari

This research introduces an innovative AI-driven multi-agent framework specifically designed for creating immersive audiobooks. Leveraging neural text-to-speech synthesis with FastSpeech 2 and VALL-E for expressive narration and…

Sound · Computer Science 2025-05-09 Shaja Arul Selvamani , Nia D'Souza Ganapathy

Conventional scalp-based EEG systems are cumbersome to use, requiring extensive setup, restrictive wiring, and conductive gels that can dry out and limit long-term monitoring, while also carrying social stigma. As a result, there is…

Neurons and Cognition · Quantitative Biology 2026-04-27 Min Suk Lee , Abhinav Uppal , Ananya Thota , Chetan Pathrabe , Rommani Mondal , Akshay Paul , Yuchen Xu , Gert Cauwenberghs