English
Related papers

Related papers: Image2Reverb: Cross-Modal Reverb Impulse Response …

200 papers

We propose a multimodal deep learning model for VR auralization that generates spatial room impulse responses (SRIRs) in real time to reconstruct scene-specific auditory perception. Employing SRIRs as the output reduces computational…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-08 Zhiyu Li , Xinwen Yue , Shenghui Zhao , Jing Wang

A technique is presented for producing synthetic images from numerical simulations whereby the image resolution is adapted around prominent features. In so doing, adaptive image ray-tracing (AIR) improves the efficiency of a calculation by…

Instrumentation and Methods for Astrophysics · Physics 2015-05-20 E. R. Parkin

State-of-the-art deep-learning-based voice activity detectors (VADs) are often trained with anechoic data. However, real acoustic environments are generally reverberant, which causes the performance to significantly deteriorate. To mitigate…

Sound · Computer Science 2021-06-28 Amir Ivry , Israel Cohen , Baruch Berdugo

We introduce an inversion based method, denoted as IMAge-Guided model INvErsion (IMAGINE), to generate high-quality and diverse images from only a single training sample. We leverage the knowledge of image semantics from a pre-trained…

Computer Vision and Pattern Recognition · Computer Science 2021-04-14 Pei Wang , Yijun Li , Krishna Kumar Singh , Jingwan Lu , Nuno Vasconcelos

We introduce a new audio processing technique that increases the sampling rate of signals such as speech or music using deep convolutional neural networks. Our model is trained on pairs of low and high-quality audio examples; at test-time,…

Sound · Computer Science 2017-08-03 Volodymyr Kuleshov , S. Zayd Enam , Stefano Ermon

We propose Im2Wav, an image guided open-domain audio generation system. Given an input image or a sequence of images, Im2Wav generates a semantically relevant sound. Im2Wav is based on two Transformer language models, that operate over a…

Sound · Computer Science 2023-02-28 Roy Sheffer , Yossi Adi

Controlling the style and characteristics of speech synthesis is crucial for adapting the output to specific contexts and user requirements. Previous Text-to-speech (TTS) works have focused primarily on the technical aspects of producing…

Sound · Computer Science 2025-09-04 Jiawei Zhang , Tian-Hao Zhang , Jun Wang , Jiaran Gao , Xinyuan Qian , Xu-Cheng Yin

Optoacoustic image formation is conventionally based upon ultrasound time-of-flight readings from multiple detection positions. Herein, we exploit acoustic scattering to physically encode the position of optical absorbers in the acquired…

Biological Physics · Physics 2019-10-30 Xose Luis Dean-Ben , Ali Ozbek , Hernan Lopez-Schier , Daniel Razansky

Text-guided image generation has witnessed unprecedented progress due to the development of diffusion models. Beyond text and image, sound is a vital element within the sphere of human perception, offering vivid representations and…

Graphics · Computer Science 2023-06-21 Yue Yang , Kaipeng Zhang , Yuying Ge , Wenqi Shao , Zeyue Xue , Yu Qiao , Ping Luo

We present a novel approach that improves the performance of reverberant speech separation. Our approach is based on an accurate geometric acoustic simulator (GAS) which generates realistic room impulse responses (RIRs) by modeling both…

Audio and Speech Processing · Electrical Eng. & Systems 2021-07-21 Rohith Aralikatti , Anton Ratnarajah , Zhenyu Tang , Dinesh Manocha

Knowing the geometrical and acoustical parameters of a room may benefit applications such as audio augmented reality, speech dereverberation or audio forensics. In this paper, we study the problem of jointly estimating the total surface…

Sound · Computer Science 2021-07-30 Prerak Srivastava , Antoine Deleforge , Emmanuel Vincent

The characteristics of a sound field are intrinsically linked to the geometric and spatial properties of the environment surrounding a sound source and a listener. The physics of sound propagation is captured in a time-domain signal known…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-21 Christopher Ick , Gordon Wichern , Yoshiki Masuyama , François Germain , Jonathan Le Roux

Implicit neural representation (INR) has proven to be accurate and efficient in various domains. In this work, we explore how different neural networks can be designed as a new texture INR, which operates in a continuous manner rather than…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Albert Kwok , Zheyuan Hu , Dounia Hammou

Remote sensing provides valuable information about objects or areas from a distance in either active (e.g., RADAR and LiDAR) or passive (e.g., multispectral and hyperspectral) modes. The quality of data acquired by remotely sensed imaging…

Image and Video Processing · Electrical Eng. & Systems 2022-11-22 Benhood Rasti , Yi Chang , Emanuele Dalsasso , Loïc Denis , Pedram Ghamisi

Inverse rendering seeks to recover 3D geometry, surface material, and lighting from captured images, enabling advanced applications such as novel-view synthesis, relighting, and virtual object insertion. However, most existing techniques…

Computer Vision and Pattern Recognition · Computer Science 2025-03-19 Chih-Hao Lin , Jia-Bin Huang , Zhengqin Li , Zhao Dong , Christian Richardt , Tuotuo Li , Michael Zollhöfer , Johannes Kopf , Shenlong Wang , Changil Kim

In this paper, we propose a way of synthesizing realistic images directly with natural language description, which has many useful applications, e.g. intelligent image manipulation. We attempt to accomplish such synthesis: given a source…

Computer Vision and Pattern Recognition · Computer Science 2017-07-24 Hao Dong , Simiao Yu , Chao Wu , Yike Guo

This paper presents a robust regression approach for image binarization under significant background variations and observation noises. The work is motivated by the need of identifying foreground regions in noisy microscopic image or…

Computer Vision and Pattern Recognition · Computer Science 2018-07-18 Garret Vo , Chiwoo Park

We present a novel algorithm for high resolution coherent imaging of sound sources in random scattering media using time resolved measurements of the acoustic pressure at an array of receivers. The sound waves travel a long distance between…

Numerical Analysis · Mathematics 2017-12-15 Liliana Borcea , Ilker Kocyigit

We consider imaging the reflectivity of scatterers from intensity-only data recorded by a single moving transducer that both emits and receives signals, forming a synthetic aperture. By exploiting frequency illumination diversity, we obtain…

Computational Physics · Physics 2019-05-15 Miguel Moscoso , Alexei Novikov , George Papanicolaou , Chrysoula Tsogka

In today's tech-driven world, significant advancements in artificial intelligence and virtual reality have emerged. These developments drive research into exploring their intersection in the realm of soundscape. Not only do these…

Sound · Computer Science 2025-04-11 Rima Ayoubi , Laurent Lescop , Sang Bum Park
‹ Prev 1 8 9 10 Next ›