中文
相关论文

相关论文: BREATH: A Bio-Radar Embodied Agent for Tonal and H…

200 篇论文

Binaural stereo audio is recorded by imitating the way the human ear receives sound, which provides people with an immersive listening experience. Existing approaches leverage autoencoders and directly exploit visual spatial information to…

声音 · 计算机科学 2023-11-15 Zhaojian Li , Bin Zhao , Yuan Yuan

Modeling late reverberation in real-time interactive applications is a challenging task when multiple sound sources and listeners are present in the same environment. This is especially problematic when the environment is geometrically…

声音 · 计算机科学 2025-10-14 Matteo Scerbo , Sebastian J. Schlecht , Randall Ali , Lauri Savioja , Enzo De Sena

We introduce a new system for data-driven audio sound model design built around two different neural network architectures, a Generative Adversarial Network(GAN) and a Recurrent Neural Network (RNN), that takes advantage of the unique…

声音 · 计算机科学 2022-06-28 Lonce Wyse , Purnima Kamath , Chitralekha Gupta

World models have demonstrated impressive performance on robotic learning tasks. Many such tasks inherently demand multimodal reasoning; for example, filling a bottle with water will lead to visual information alone being ambiguous or…

机器人学 · 计算机科学 2025-12-10 Fan Zhang , Michael Gienger

Music is both an auditory and an embodied phenomenon, closely linked to human motion and naturally expressed through dance. However, most existing audio representations neglect this embodied dimension, limiting their ability to capture…

声音 · 计算机科学 2026-01-30 Xuanchen Wang , Heng Wang , Weidong Cai

Identifying novel hypotheses is essential to scientific research, yet this process risks being overwhelmed by the sheer volume and complexity of available information. Existing automated methods often struggle to generate novel and…

This study presents the progress of our recent research regarding wireless measurements of the human body. First, we explain radar imaging algorithms for screening passengers at security checkpoints that can process data faster than…

信号处理 · 电气工程与系统科学 2020-12-04 Takuya Sakamoto

We propose Beat Transformer, a novel Transformer encoder architecture for joint beat and downbeat tracking. Different from previous models that track beats solely based on the spectrogram of an audio mixture, our model deals with demixed…

声音 · 计算机科学 2022-09-16 Jingwei Zhao , Gus Xia , Ye Wang

The present paper develops recursive algorithms to track shifts in the resonance frequency of linear systems in real time. To date, automatic resonance tracking has been limited to non-model-based approaches, which rely solely on the phase…

最优化与控制 · 数学 2020-12-22 Thomas Vasileiou

This paper introduces a generative model designed for multimodal control over text-to-image foundation generative AI models such as Stable Diffusion, specifically tailored for engineering design synthesis. Our model proposes parametric,…

人工智能 · 计算机科学 2024-12-09 Rui Zhou , Yanxia Zhang , Chenyang Yuan , Frank Permenter , Nikos Arechiga , Matt Klenk , Faez Ahmed

Recent advancements in music generation have garnered significant attention, yet existing approaches face critical limitations. Some current generative models can only synthesize either the vocal track or the accompaniment track. While some…

音频与语音处理 · 电气工程与系统科学 2025-03-04 Ziqian Ning , Huakang Chen , Yuepeng Jiang , Chunbo Hao , Guobin Ma , Shuai Wang , Jixun Yao , Lei Xie

The recent success of the generative model shows that leveraging the multi-modal embedding space can manipulate an image using text information. However, manipulating an image with other sources rather than text, such as sound, is not easy…

图形学 · 计算机科学 2021-12-02 Seung Hyun Lee , Wonseok Roh , Wonmin Byeon , Sang Ho Yoon , Chan Young Kim , Jinkyu Kim , Sangpil Kim

Electrical waves in the heart form rotating spiral or scroll waves during life-threatening arrhythmias such as atrial or ventricular fibrillation. The wave dynamics are typically modeled using coupled partial differential equations, which…

医学物理 · 物理学 2024-10-07 Tanish Baranwal , Jan Lebert , Jan Christoph

Clinical image interpretation is inherently multi-step and tool-centric: clinicians iteratively combine visual evidence with patient context, quantify findings, and refine their decisions through a sequence of specialized procedures. While…

人工智能 · 计算机科学 2026-03-09 Lin Fan , Pengyu Dai , Zhipeng Deng , Haolin Wang , Xun Gong , Yefeng Zheng , Yafei Ou

Lyric interpretations can help people understand songs and their lyrics quickly, and can also make it easier to manage, retrieve and discover songs efficiently from the growing mass of music archives. In this paper we propose BART-fusion, a…

声音 · 计算机科学 2022-08-25 Yixiao Zhang , Junyan Jiang , Gus Xia , Simon Dixon

Drawing real world social inferences usually requires taking into account information from multiple modalities. Language is a particularly powerful source of information in social settings, especially in novel situations where language can…

Recent approaches in music generation rely on disentangled representations, often labeled as structure and timbre or local and global, to enable controllable synthesis. Yet the underlying properties of these embeddings remain underexplored.…

The complex nature of musical emotion introduces inherent bias in both recognition and generation, particularly when relying on a single audio encoder, emotion classifier, or evaluation metric. In this work, we conduct a study on Music…

音频与语音处理 · 电气工程与系统科学 2025-05-01 Yuanchao Li , Azalea Gui , Dimitra Emmanouilidou , Hannes Gamper

The human somatosensory system integrates multimodal sensory feedback, including tactile, proprioceptive, and thermal signals, to enable comprehensive perception and effective interaction with the environment. Inspired by the biological…

机器人学 · 计算机科学 2025-09-03 Fengyi Wang , Xiangyu Fu , Nitish Thakor , Gordon Cheng

Music Emotion Recognition involves the automatic identification of emotional elements within music tracks, and it has garnered significant attention due to its broad applicability in the field of Music Information Retrieval. It can also be…

声音 · 计算机科学 2023-08-29 Kexin Zhu , Xulong Zhang , Jianzong Wang , Ning Cheng , Jing Xiao