English
Related papers

Related papers: NeuralSound: Learning-based Modal Sound Synthesis …

200 papers

This paper elucidates a model for acoustic single and multi-tone classification in resource constrained edge devices. The proposed model is of State-of-the-art Fast Accurate Stable Tiny Gated Recurrent Neural Network. This model has…

Machine Learning · Computer Science 2021-12-07 Raghav Rawat , Shreyash Gupta , Shreyas Mohapatra , Sujata Priyambada Mishra , Sreesankar Rajagopal

Accurately modeling sound propagation with complex real-world environments is essential for Novel View Acoustic Synthesis (NVAS). While previous studies have leveraged visual perception to estimate spatial acoustics, the combined use of…

Multimedia · Computer Science 2025-03-18 Hadam Baek , Hannie Shin , Jiyoung Seo , Chanwoo Kim , Saerom Kim , Hyeongbok Kim , Sangpil Kim

Despite the recent developments in the field of cross-modal retrieval, there has been less research focusing on low-resource languages due to the lack of manually annotated datasets. In this paper, we propose a noise-robust cross-lingual…

Computer Vision and Pattern Recognition · Computer Science 2022-08-29 Yabing Wang , Jianfeng Dong , Tianxiang Liang , Minsong Zhang , Rui Cai , Xun Wang

Multimodal transfer learning aims to transform pretrained representations of diverse modalities into a common domain space for effective multimodal fusion. However, conventional systems are typically built on the assumption that all…

Computer Vision and Pattern Recognition · Computer Science 2023-09-28 Yanan Wang , Donghuo Zeng , Shinya Wada , Satoshi Kurihara

This paper introduces a new multi-modal model based on the Transformer architecture and tensor product fusion strategy, combining BERT's text vectors and ViT's image vectors to classify students' psychological conditions, with an accuracy…

Computer Vision and Pattern Recognition · Computer Science 2024-11-19 Ao Xiang , Zongqing Qi , Han Wang , Qin Yang , Danqing Ma

We propose a learnable content adaptive front end for audio signal processing. Before the modern advent of deep learning, we used fixed representation non-learnable front-ends like spectrogram or mel-spectrogram with/without neural…

Sound · Computer Science 2024-12-24 Prateek Verma , Chris Chafe

Data-driven modeling of complex physical systems is receiving a growing amount of attention in the simulation and machine learning communities. Since most physical simulations are based on compute-intensive, iterative implementations of…

Sound · Computer Science 2024-03-20 Martin Spitznagel , Janis Keuper

Reconstructing the room transfer functions needed to calculate the complex sound field in a room has several important real-world applications. However, an unpractical number of microphones is often required. Recently, in addition to…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-12 Francesca Ronchini , Luca Comanducci , Mirco Pezzoli , Fabio Antonacci , Augusto Sarti

Sequential models achieve state-of-the-art results in audio, visual and textual domains with respect to both estimating the data distribution and generating high-quality samples. Efficient sampling for this class of models has however…

Tone Transfer is a novel deep-learning technique for interfacing a sound source with a synthesizer, transforming the timbre of audio excerpts while keeping their musical form content. Due to its good audio quality results and continuous…

Sound · Computer Science 2023-10-10 Franco Caspe , Andrew McPherson , Mark Sandler

Computed ultrasound tomography in echo mode generates maps of tissue speed of sound (SoS) from the shift of echoes when detected under varying steering angles. It solves a linearized inverse problem that requires regularization to…

Medical Physics · Physics 2026-04-23 Parisa Salemi Yolgunlu , Jules Blom , Naiara Korta Martiartu , Michael Jaeger

Previous works (Donahue et al., 2018a; Engel et al., 2019a) have found that generating coherent raw audio waveforms with GANs is challenging. In this paper, we show that it is possible to train GANs reliably to generate high quality…

Audio and Speech Processing · Electrical Eng. & Systems 2019-12-10 Kundan Kumar , Rithesh Kumar , Thibault de Boissiere , Lucas Gestin , Wei Zhen Teoh , Jose Sotelo , Alexandre de Brebisson , Yoshua Bengio , Aaron Courville

The goal of multi-modal learning is to use complimentary information on the relevant task provided by the multiple modalities to achieve reliable and robust performance. Recently, deep learning has led significant improvement in multi-modal…

Computer Vision and Pattern Recognition · Computer Science 2018-11-05 Jaekyum Kim , Junho Koh , Yecheol Kim , Jaehyung Choi , Youngbae Hwang , Jun Won Choi

Radiative transfer calculations are essential for modeling planetary atmospheres. However, standard methods are computationally demanding and impose accuracy-speed trade-offs. High computational costs force numerical simplifications in…

Earth and Planetary Astrophysics · Physics 2025-11-03 Isaac Malsky , Tiffany Kataria , Natasha E. Batalha , Matthew Graham

The introduction of the neural implicit representation has notably propelled the advancement of online dense reconstruction techniques. Compared to traditional explicit representations, such as TSDF, it improves the mapping completeness and…

Computer Vision and Pattern Recognition · Computer Science 2025-09-16 Yuqing Lan , Chenyang Zhu , Shuaifeng Zhi , Jiazhao Zhang , Zhoufeng Wang , Renjiao Yi , Yijie Wang , Kai Xu

We present a physics-informed voiced backend renderer for singing-voice synthesis. Given synthetic single-channel audio and a fund-amental--frequency trajectory, we train a time-domain Webster model as a physics-informed neural network to…

Sound · Computer Science 2026-03-03 Minhui Lu , Joshua D. Reiss

In this paper, we present a deep learning based wireless transceiver. We describe in detail the corresponding artificial neural network architecture, the training process, and report on excessive over-the-air measurement results. We employ…

Signal Processing · Electrical Eng. & Systems 2019-05-28 Johannes Schmitz , Caspar von Lengerke , Nikita Airee , Arash Behboodi , Rudolf Mathar

In recent years, Deep Learning has been successfully applied to multimodal learning problems, with the aim of learning useful joint representations in data fusion applications. When the available modalities consist of time series data such…

Computer Vision and Pattern Recognition · Computer Science 2017-04-12 Xitong Yang , Palghat Ramesh , Radha Chitta , Sriganesh Madhvanath , Edgar A. Bernal , Jiebo Luo

This paper presents a Multi-modal Emotion Recognition (MER) system designed to enhance emotion recognition accuracy in challenging acoustic conditions. Our approach combines a modified and extended Hierarchical Token-semantic Audio…

Sound · Computer Science 2025-07-30 Ohad Cohen , Gershon Hazan , Sharon Gannot

While both the data volume and heterogeneity of the digital music content is huge, it has become increasingly important and convenient to build a recommendation or search system to facilitate surfacing these content to the user or consumer…