English
Related papers

Related papers: Multimodal Deep Learning Method for Real-Time Spat…

200 papers

Reconstructing high-fidelity magnetic resonance (MR) images from under-sampled k-space is a commonly used strategy to reduce scan time. The posterior sampling of diffusion models based on the real measurement data holds significant promise…

Image and Video Processing · Electrical Eng. & Systems 2024-07-04 Jiayue Chu , Chenhe Du , Xiyue Lin , Yuyao Zhang , Hongjiang Wei

Measuring the acoustic characteristics of a space is often done by capturing its impulse response (IR), a representation of how a full-range stimulus sound excites it. This work generates an IR from a single image, which can then be applied…

Sound · Computer Science 2021-08-17 Nikhil Singh , Jeff Mentch , Jerry Ng , Matthew Beveridge , Iddo Drori

This paper presents a Multi-modal Emotion Recognition (MER) system designed to enhance emotion recognition accuracy in challenging acoustic conditions. Our approach combines a modified and extended Hierarchical Token-semantic Audio…

Sound · Computer Science 2025-07-30 Ohad Cohen , Gershon Hazan , Sharon Gannot

In multimedia applications such as films and video games, spatial audio techniques are widely employed to enhance user experiences by simulating 3D sound: transforming mono audio into binaural formats. However, this process is often complex…

Multimedia · Computer Science 2025-02-14 Xiaojing Liu , Ogulcan Gurelli , Yan Wang , Joshua Reiss

Low Dose Computed Tomography suffers from a high amount of noise and/or undersampling artefacts in the reconstructed image. In the current article, a Deep Learning technique is exploited as a regularization term for the iterative…

Image and Video Processing · Electrical Eng. & Systems 2019-06-04 Shabab Bazrafkan , Vincent Van Nieuwenhove , Joris Soons , Jan De Beenhouwer , Jan Sijbers

In this work, we propose a training algorithm for an audio-visual automatic speech recognition (AV-ASR) system using deep recurrent neural network (RNN).First, we train a deep RNN acoustic model with a Connectionist Temporal Classification…

Computer Vision and Pattern Recognition · Computer Science 2016-11-10 Abhinav Thanda , Shankar M Venkatesan

Rotating-view thick-slice acquisition is highly SNR-efficient for mesoscale diffusion MRI (dMRI) but requires numerous rotating views to satisfy Nyquist sampling, resulting in long scan time. We propose a self-supervised Spatial-Angular…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Yinzhe Wu , Hongyu Rui , Fanwen Wang , Jiahao Huang , Zi Wang , Guang Yang

Multimodal image super-resolution (SR) is the reconstruction of a high resolution image given a low-resolution observation with the aid of another image modality. While existing deep multimodal models do not incorporate domain knowledge…

Computer Vision and Pattern Recognition · Computer Science 2020-09-08 Iman Marivani , Evaggelia Tsiligianni , Bruno Cornelis , Nikos Deligiannis

Interactions with virtual assistants typically start with a predefined trigger phrase followed by the user command. To make interactions with the assistant more intuitive, we explore whether it is feasible to drop the requirement that users…

Computation and Language · Computer Science 2024-03-27 Dominik Wagner , Alexander Churchill , Siddharth Sigtia , Panayiotis Georgiou , Matt Mirsamadi , Aarshee Mishra , Erik Marchi

Audio-visual speech recognition (AVSR) has gained remarkable success for ameliorating the noise-robustness of speech recognition. Mainstream methods focus on fusing audio and visual inputs to obtain modality-invariant representations.…

Sound · Computer Science 2023-02-03 Chen Chen , Yuchen Hu , Qiang Zhang , Heqing Zou , Beier Zhu , Eng Siong Chng

A real-world application or setting involves interaction between different modalities (e.g., video, speech, text). In order to process the multimodal information automatically and use it for an end application, Multimodal Representation…

Computer Vision and Pattern Recognition · Computer Science 2022-11-08 Abhinav Joshi , Naman Gupta , Jinang Shah , Binod Bhattarai , Ashutosh Modi , Danail Stoyanov

We propose a physically-motivated deep learning framework to solve a general version of the challenging indoor lighting estimation problem. Given a single LDR image with a depth map, our method predicts spatially consistent lighting at any…

Computer Vision and Pattern Recognition · Computer Science 2023-05-09 Zhengqin Li , Li Yu , Mikhail Okunev , Manmohan Chandraker , Zhao Dong

In this paper, we present methods in deep multimodal learning for fusing speech and visual modalities for Audio-Visual Automatic Speech Recognition (AV-ASR). First, we study an approach where uni-modal deep networks are trained separately…

Computation and Language · Computer Science 2015-01-23 Youssef Mroueh , Etienne Marcheret , Vaibhava Goel

Vision-language models (VLM) excel at general understanding yet remain weak at dynamic spatial reasoning (DSR), i.e., reasoning about the evolvement of object geometry and relationship in 3D space over time, largely due to the scarcity of…

Computer Vision and Pattern Recognition · Computer Science 2025-12-24 Shengchao Zhou , Yuxin Chen , Yuying Ge , Wei Huang , Jiehong Lin , Ying Shan , Xiaojuan Qi

Multi-source localization is an important and challenging technique for multi-talker conversation analysis. This paper proposes a novel supervised learning method using deep neural networks to estimate the direction of arrival (DOA) of all…

Audio and Speech Processing · Electrical Eng. & Systems 2021-11-30 Aswin Shanmugam Subramanian , Chao Weng , Shinji Watanabe , Meng Yu , Dong Yu

In the domain of RIS-based indoor localization, our work introduces two distinct approaches to address real-world challenges. The first method is based on deep learning, employing a Long Short-Term Memory (LSTM) network. The second, a novel…

Signal Processing · Electrical Eng. & Systems 2024-05-06 Rafael A. Aguiar , Nuno Paulino , Luís M. Pessoa

Super-resolution (SR) for image enhancement has great importance in medical image applications. Broadly speaking, there are two types of SR, one requires multiple low resolution (LR) images from different views of the same object to be…

Image and Video Processing · Electrical Eng. & Systems 2018-10-17 Jin Zhu , Guang Yang , Pietro Lio

This paper proposes SOLVR, a unified pipeline for learning based LiDAR-Visual re-localisation which performs place recognition and 6-DoF registration across sensor modalities. We propose a strategy to align the input sensor modalities by…

Computer Vision and Pattern Recognition · Computer Science 2025-02-07 Joshua Knights , Sebastián Barbas Laina , Peyman Moghadam , Stefan Leutenegger

We present a method for improving the quality of synthetic room impulse responses for far-field speech recognition. We bridge the gap between the fidelity of synthetic room impulse responses (RIRs) and the real room impulse responses using…

Sound · Computer Science 2021-11-15 Anton Ratnarajah , Zhenyu Tang , Dinesh Manocha

Changes in room acoustics, such as modifications to surface absorption or the insertion of a scattering object, significantly impact measured room impulse responses (RIRs). These changes can affect the performance of systems used in echo…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-03 Karolina Prawda