English
Related papers

Related papers: Input-Envelope-Output: Auditable Generative Music …

200 papers

Encoder-decoder based neural architectures serve as the basis of state-of-the-art approaches in end-to-end open domain dialog systems. Since most of such systems are trained with a maximum likelihood~(MLE) objective they suffer from issues…

Advanced diffusion models like RPG, Stable Diffusion 3 and FLUX have made notable strides in compositional text-to-image generation. However, these methods typically exhibit distinct strengths for compositional generation, with some…

Computer Vision and Pattern Recognition · Computer Science 2025-02-06 Xinchen Zhang , Ling Yang , Guohao Li , Yaqi Cai , Jiake Xie , Yong Tang , Yujiu Yang , Mengdi Wang , Bin Cui

Sensor-based interactive systems -- e.g., "smart" speakers, webcams, and RFID tags -- allow us to embed computational functionality into physical environments. They also expose users to real and perceived privacy risks: users know that…

Human-Computer Interaction · Computer Science 2026-04-02 Youngwook Do , Yuxi Wu , Gregory D. Abowd , Sauvik Das

This paper presents a context-aware framework for feature selection and classification procedures to realize a fast and accurate audio event annotation and classification. The context-aware design starts with exploring feature extraction…

Sound · Computer Science 2023-03-08 M. Mehrdad Morsali , Hoda Mohammadzade , Saeed Bagheri Shouraki

Generative audio models are rapidly advancing in both capabilities and public utilization -- several powerful generative audio models have readily available open weights, and some tech companies have released high quality generative audio…

In large-scale networks of uncertain dynamical systems, where communication is limited and there is a strong interaction among subsystems, learning local models and control policies offers great potential for designing high-performance…

Systems and Control · Electrical Eng. & Systems 2021-11-08 Andrea Carron , Jerome Sieber , Melanie N. Zeilinger

Many promising applications of multimodal wearables require continuous sensing and heavy computation, yet users reject such devices due to privacy concerns. This paper shares our experiences building an ear-mounted voice-and-vision wearable…

Human-Computer Interaction · Computer Science 2025-11-26 Yonatan Tussa , Andy Heredia , Nirupam Roy

Humans possess a remarkable ability to integrate auditory and visual information, enabling a deeper understanding of the surrounding environment. This early fusion of audio and visual cues, demonstrated through cognitive psychology and…

Computer Vision and Pattern Recognition · Computer Science 2023-12-05 Shentong Mo , Pedro Morgado

AudioSet is a widely used benchmark in the audio research community and has significantly advanced various audio-related tasks. However, persistent issues with label accuracy and completeness remain critical bottlenecks that limit…

Sound · Computer Science 2025-08-25 Yulin Sun , Qisheng Xu , Yi Su , Qian Zhu , Yong Dou , Xinwang Liu , Kele Xu

Recent advances in Internet-of-Things (IoT) technologies have sparked significant interest towards developing learning-based sensing applications on embedded edge devices. These efforts, however, are being challenged by the complexities of…

Systems and Control · Electrical Eng. & Systems 2024-02-23 Abdulrahman Bukhari , Seyedmehdi Hosseinimotlagh , Hyoseung Kim

Open generative models are vitally important for the community, allowing for fine-tunes and serving as baselines when presenting new models. However, most current text-to-audio models are private and not accessible for artists and…

Sound · Computer Science 2024-08-01 Zach Evans , Julian D. Parker , CJ Carr , Zack Zukowski , Josiah Taylor , Jordi Pons

Music genre classification shapes how listeners discover music, how platforms design recommendations, and how sociologists study cultural taste. Yet existing genre labels are inconsistent in granularity: they exaggerate boundaries between…

Physics and Society · Physics 2026-04-30 Makoto Takeuchi

Multimodal Large Language Models (MLLMs) excel in Open-Vocabulary (OV) emotion recognition but often neglect fine-grained acoustic modeling. Existing methods typically use global audio encoders, failing to capture subtle, local temporal…

Multimedia · Computer Science 2026-03-24 Liyun Zhang , Xuanmeng Sha , Shuqiong Wu , Fengkai Liu

Immersive spatial audio has become increasingly critical for applications ranging from AR/VR to home entertainment and automotive sound systems. However, existing generative methods remain constrained to low-dimensional formats such as…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-21 Zining Liang , Runbang Wang , Xuzhou Ye , Qiuqiang Kong

This paper addresses the design of robust dynamic output feedback control for highly uncertain systems in which the unknown disturbance might be excited by the derivative of the control input. This context appears in many industrial…

Systems and Control · Computer Science 2016-10-20 Mazen Alamir , Jean Dobrowolski , Amgad tarek Mohammed

Autism spectrum disorder (ASD) is a neurodevelopmental condition characterized by challenges in social communication, repetitive behavior, and sensory processing. One important research area in ASD is evaluating children's behavioral…

Sound · Computer Science 2025-06-03 Tiantian Feng , Anfeng Xu , Xuan Shi , Somer Bishop , Shrikanth Narayanan

Instruction-guided text-to-speech (ITTS) enables users to control speech generation through natural language prompts, offering a more intuitive interface than traditional TTS. However, the alignment between user style instructions and…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-28 Yi-Cheng Lin , Huang-Cheng Chou , Tzu-Chieh Wei , Kuan-Yu Chen , Hung-yi Lee

This paper investigates the distributed event-triggered control problem for a class of uncertain pure-feedback nonlinear multi-agent systems (MASs) with polluted feedback. Under the setting of event-triggered control, substantial challenges…

Systems and Control · Electrical Eng. & Systems 2023-02-28 Libei Sun , Zhirong Zhang , Xinjian Huang , Xiucai Huang

Virtual acoustic environments enable the creation and simulation of realistic and eco-logically valid daily-life situations vital for hearing research and audiology. Reverberant indoor environments are particularly important. For real-time…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-03 Stefan Fichna , Steven van de Par , Bernhard U. Seeber , Stephan D. Ewert

Robust design of autonomous systems under uncertainty is an important yet challenging problem. This work proposes a robust controller that consists of a state estimator and a tube based predictive control law. The class of linear systems…

Systems and Control · Electrical Eng. & Systems 2022-10-11 Tianchen Ji , Junyi Geng , Katherine Driggs-Campbell