English
Related papers

Related papers: Input-Envelope-Output: Auditable Generative Music …

200 papers

This paper presents InfiniteAudio, a simple yet effective strategy for generating infinite-length audio using diffusion-based text-to-audio methods. Current approaches face memory constraints because the output size increases with input…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-04 Chaeyoung Jung , Hojoon Ki , Ji-Hoon Kim , Junmo Kim , Joon Son Chung

As large language models (LLMs) become increasingly integrated into daily life, audio has emerged as a key interface for human-AI interaction. However, this convenience also introduces new vulnerabilities, making audio a potential attack…

Sound · Computer Science 2026-02-05 Hiskias Dingeto , Taeyoun Kwon , Dasol Choi , Bodam Kim , DongGeon Lee , Haon Park , JaeHoon Lee , Jongho Shin

We address the problem of dynamic output feedback stabilization at an unobservable target point. The challenge lies in according the antagonistic nature of the objective and the properties of the system: the system tends to be less…

Optimization and Control · Mathematics 2022-02-02 Lucas Brivadis , Jean-Paul Gauthier , Ludovic Sacchelli , Ulysse Serres

The design of embedded systems, that are ubiquitously used in mobile devices and cars, is becoming continuously more complex such that efficient system-level design methods are becoming crucial. My research aims at developing systems that…

Artificial Intelligence · Computer Science 2019-05-15 Philipp Wanko

AI systems are becoming active participants in organizational and knowledge work. They increasingly interact with humans, coordinate workflows, and operate in multi-agent arrangements. Understanding their effects therefore requires more…

Artificial Intelligence · Computer Science 2026-05-19 Yingjie Zhang , Chun Feng , Weizhang Zhu , Tianshu Sun

Soundscape augmentation or "masking" introduces wanted sounds into the acoustic environment to improve acoustic comfort. Usually, the masker selection and playback strategies are either arbitrary or based on simple rules (e.g. -3 dBA),…

Audio and Speech Processing · Electrical Eng. & Systems 2023-08-16 Bhan Lam , Zhen-Ting Ong , Kenneth Ooi , Wen-Hui Ong , Trevor Wong , Karn N. Watcharasupat , Woon-Seng Gan

In this work, we propose a compositional scheme based on small-gain reasoning to synthesize safety controllers for interconnected stochastic hybrid systems. In our proposed setting, we first offer an augmented scheme that characterizes each…

Systems and Control · Electrical Eng. & Systems 2026-04-14 Mahdieh Zaker , Omid Akbarzadeh , Behrad Samari , Abolfazl Lavaei

We propose a benchmark for evaluating compositionality in audio representations. Audio compositionality refers to representing sound scenes in terms of constituent sources and attributes, and combining them systematically. While central to…

Sound · Computer Science 2026-03-17 Chuyang Chen , Bea Steers , Brian McFee , Juan Bello

With the increasing popularity of Internet of Things (IoT) devices, securing sensitive user data has emerged as a major challenge. These devices often collect confidential information, such as audio and visual data, through peripheral…

Cryptography and Security · Computer Science 2023-12-21 Peterson Yuhala , Jämes Ménétrey , Pascal Felber , Marcelo Pasin , Valerio Schiavoni

This paper introduces a novel approach to speech restoration by integrating a context-related conditioning strategy. Specifically, we employ the diffusion-based generative restoration model, UNIVERSE++, as a backbone to evaluate the…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-13 Soo-Whan Chung , Min-Seok Choi

Can large language model agents discover hidden safety objectives through experience alone? We introduce EPO-Safe (Experiential Prompt Optimization for Safe Agents), a framework where an LLM iteratively generates action plans, receives…

Artificial Intelligence · Computer Science 2026-04-28 Víctor Gallego

During social interactions, understanding the intricacies of the context can be vital, particularly for socially anxious individuals. While previous research has found that the presence of a social interaction can be detected from ambient…

Human-Computer Interaction · Computer Science 2024-07-22 Varun Reddy , Zhiyuan Wang , Emma Toner , Max Larrazabal , Mehdi Boukhechba , Bethany A. Teachman , Laura E. Barnes

Embodied agents require robust navigation systems to operate in unstructured environments, making the robustness of Simultaneous Localization and Mapping (SLAM) models critical to embodied agent autonomy. While real-world datasets are…

Computer Vision and Pattern Recognition · Computer Science 2024-06-25 Xiaohao Xu , Tianyi Zhang , Sibo Wang , Xiang Li , Yongqi Chen , Ye Li , Bhiksha Raj , Matthew Johnson-Roberson , Xiaonan Huang

Expressive music performance rendering involves interpreting symbolic scores with variations in timing, dynamics, articulation, and instrument-specific techniques, resulting in performances that capture musical can emotional intent. We…

Audio and Speech Processing · Electrical Eng. & Systems 2025-02-12 Huan Zhang , Akira Maezawa , Simon Dixon

This paper proposes a multi-agent artificial intelligence system that generates response-oriented media content in real time based on audio-derived emotional signals. Unlike conventional speech emotion recognition studies that focus…

Artificial Intelligence · Computer Science 2026-01-21 HyeYoung Lee

The safety of learning-enabled cyber-physical systems is compromised by the well-known vulnerabilities of deep neural networks to out-of-distribution (OOD) inputs. Existing literature has sought to monitor the safety of such systems by…

Machine Learning · Computer Science 2025-04-21 Vivian Lin , Ramneet Kaur , Yahan Yang , Souradeep Dutta , Yiannis Kantaros , Anirban Roy , Susmit Jha , Oleg Sokolsky , Insup Lee

Large-scale multimodal generative modeling has created milestones in text-to-image and text-to-video generation. Its application to audio still lags behind for two main reasons: the lack of large-scale datasets with high-quality text-audio…

A site-specific Type-II codebook design is proposed for downlink massive multiple-input multiple-output (MIMO) limited-feedback beamforming. The key idea is to embed a learned site-specific propagation prior into the Type-II channel state…

Signal Processing · Electrical Eng. & Systems 2026-04-24 Cheng-Jie Zhao , Zhaolin Wang , Zongyao Zhao , Yuanwei Liu

We introduce a new system for data-driven audio sound model design built around two different neural network architectures, a Generative Adversarial Network(GAN) and a Recurrent Neural Network (RNN), that takes advantage of the unique…

Sound · Computer Science 2022-06-28 Lonce Wyse , Purnima Kamath , Chitralekha Gupta

We investigate a coherent feedback squeezer that uses quantum coherent feedback (measurement-free) control. Our squeezer is simple, easy to implement, robust to the gain fluctuation, and broadband compared to the existing squeezers because…

‹ Prev 1 3 4 5 6 7 10 Next ›