中文
相关论文

相关论文: AIRCADE: an Anechoic and IR Convolution-based Aura…

200 篇论文

Recent progress in audio-language modeling, such as automated audio captioning, has benefited from training on synthetic data generated with the aid of large-language models. However, such approaches for environmental sound captioning have…

声音 · 计算机科学 2024-10-17 Mithun Manivannan , Vignesh Nethrapalli , Mark Cartwright

The integration of multimodal Electronic Health Records (EHR) data has significantly advanced clinical predictive capabilities. Existing models, which utilize clinical notes and multivariate time-series EHR data, often fall short of…

计算与语言 · 计算机科学 2025-02-27 Yinghao Zhu , Changyu Ren , Zixiang Wang , Xiaochen Zheng , Shiyun Xie , Junlan Feng , Xi Zhu , Zhoujun Li , Liantao Ma , Chengwei Pan

This paper introduces a new training strategy to improve speech dereverberation systems using minimal acoustic information and reverberant (wet) speech. Most existing algorithms rely on paired dry/wet data, which is difficult to obtain, or…

音频与语音处理 · 电气工程与系统科学 2025-02-12 Louis Bahrman , Mathieu Fontaine , Gael Richard

Medical Ultrasound (US), despite its wide use, is characterized by artifacts and operator dependency. Those attributes hinder the gathering and utilization of US datasets for the training of Deep Neural Networks used for Computer-Assisted…

图像与视频处理 · 电气工程与系统科学 2021-05-06 Maria Tirindelli , Christine Eilers , Walter Simson , Magdalini Paschali , Mohammad Farid Azampour , Nassir Navab

This paper proposes a multimodal emotion recognition system based on hybrid fusion that classifies the emotions depicted by speech utterances and corresponding images into discrete classes. A new interpretability technique has been…

计算机视觉与模式识别 · 计算机科学 2023-01-10 Puneet Kumar , Sarthak Malik , Balasubramanian Raman

We present a novel approach to improve the performance of learning-based speech dereverberation using accurate synthetic datasets. Our approach is designed to recover the reverb-free signal from a reverberant speech signal. We show that…

音频与语音处理 · 电气工程与系统科学 2022-12-13 Rohith Aralikatti , Zhenyu Tang , Dinesh Manocha

Compound Expression Recognition (CER) is vital for effective interpersonal interactions. Human emotional expressions are inherently complex due to the presence of compound expressions, requiring the consideration of both local and global…

计算机视觉与模式识别 · 计算机科学 2024-07-29 Xuxiong Liu , Kang Shen , Jun Yao , Boyan Wang , Minrui Liu , Liuwei An , Zishun Cui , Weijie Feng , Xiao Sun

Text data is commonly utilized as a primary input to enhance Speech Emotion Recognition (SER) performance and reliability. However, the reliance on human-transcribed text in most studies impedes the development of practical SER systems,…

音频与语音处理 · 电气工程与系统科学 2025-03-25 Yuanchao Li , Peter Bell , Catherine Lai

This paper presents a joint source separation algorithm that simultaneously reduces acoustic echo, reverberation and interfering sources. Target speeches are separated from the mixture by maximizing independence with respect to the other…

声音 · 计算机科学 2021-04-12 Yueyue Na , Ziteng Wang , Zhang Liu , Biao Tian , Qiang Fu

Typically, neural network-based speech dereverberation models are trained on paired data, composed of a dry utterance and its corresponding reverberant utterance. The main limitation of this approach is that such models can only be trained…

音频与语音处理 · 电气工程与系统科学 2022-03-30 Hannah Muckenhirn , Aleksandr Safin , Hakan Erdogan , Felix de Chaumont Quitry , Marco Tagliasacchi , Scott Wisdom , John R. Hershey

This paper introduces a generalization of the empirical interpolation method (EIM) and the reduced basis method (RBM) in order to allow their combination with data mining and data assimilation. The purpose is to be able to derive sound…

数值分析 · 数学 2017-05-09 Y. Maday , O. Mula

Compound Expression Recognition (CER) plays a crucial role in interpersonal interactions. Due to the existence of Compound Expressions , human emotional expressions are complex, requiring consideration of both local and global facial…

计算机视觉与模式识别 · 计算机科学 2024-03-20 Jun Yu , Jichao Zhu , Wangyuan Zhu

Automatic lyric transcription (ALT) is a nascent field of study attracting increasing interest from both the speech and music information retrieval communities, given its significant application potential. However, ALT with audio data alone…

音频与语音处理 · 电气工程与系统科学 2023-02-20 Xiangming Gu , Longshen Ou , Danielle Ong , Ye Wang

In this study, we aim to determine if generalized sounds and music can share a common emotional space, improving predictions of emotion in terms of arousal and valence. We propose the use of multiple datasets as a multi-domain learning…

声音 · 计算机科学 2024-08-15 Federico Simonetta , Francesca Certo , Stavros Ntalampiras

In this paper, we study different approaches for classifying emotions from speech using acoustic and text-based features. We propose to obtain contextualized word embeddings with BERT to represent the information contained in speech…

机器学习 · 计算机科学 2024-03-28 Leonardo Pepino , Pablo Riera , Luciana Ferrer , Agustin Gravano

Multilevel compositional data are data that are repeatedly measured or clustered within groups and are non-negative and sum to a constant value. These data arise in various settings, such as intensive, longitudinal studies using ecological…

统计方法学 · 统计学 2025-02-21 Flora Le , Tyman E. Stanford , Dorothea Dumuid , Joshua F. Wiley

Temporal causal representation learning is a powerful tool for uncovering complex patterns in observational studies, which are often represented as low-dimensional time series. However, in many real-world applications, data are…

机器学习 · 计算机科学 2025-07-21 Jianhong Chen , Meng Zhao , Mostafa Reisi Gahrooei , Xubo Yue

Most recent advances in audio dereverberation focus almost exclusively on speech, leaving percussive and drum signals largely unexplored despite their importance in music production. Percussive dereverberation poses distinct challenges due…

声音 · 计算机科学 2026-05-12 Dimos Makris , András Barják , Maximos Kaliakatsos-Papakostas

Lyric interpretations can help people understand songs and their lyrics quickly, and can also make it easier to manage, retrieve and discover songs efficiently from the growing mass of music archives. In this paper we propose BART-fusion, a…

声音 · 计算机科学 2022-08-25 Yixiao Zhang , Junyan Jiang , Gus Xia , Simon Dixon

The advance of technology for transmitting Data-over-Sound in various IoT and telecommunication applications has led to the concept of machine-to-machine over-the-air acoustic signalling. Reverberation can have a detrimental effect on such…

音频与语音处理 · 电气工程与系统科学 2019-08-14 Amogh Matt , Dan Stowell