中文
相关论文

相关论文: Room Impulse Response Synthesis via Differentiable…

200 篇论文

The recent advances in deep learning indicate significant progress in the field of single image super-resolution. With the advent of these techniques, high-resolution image with high peak signal to noise ratio (PSNR) and excellent…

图像与视频处理 · 电气工程与系统科学 2020-04-09 Meenu Ajith , Aswathy Rajendra Kurup , Manel Martínez-Ramón

In time-varying fading channels, channel coefficients are estimated using pilot symbols that are transmitted every coherence interval. For channels with high Doppler spread, the rapid channel variations over time will require considerable…

信息论 · 计算机科学 2022-03-24 Sandesh Rao Mattu , Lakshmi Narasimhan T , A. Chockalingam

We present a novel approach for super-resolution that utilizes implicit neural representation (INR) to effectively reconstruct and enhance low-resolution videos and images. By leveraging the capacity of neural networks to implicitly encode…

计算机视觉与模式识别 · 计算机科学 2025-03-07 Mary Aiyetigbo , Wanqi Yuan , Feng Luo , Nianyi Li

Deep-learning-based approaches to depth estimation are rapidly advancing, offering superior performance over existing methods. To estimate the depth in real-world scenarios, depth estimation models require the robustness of various noise…

计算机视觉与模式识别 · 计算机科学 2022-04-06 Zhengyang Lu , Ying Chen

Real-time Deep Neural Network (DNN) inference with low-latency requirement has become increasingly important for numerous applications in both cloud computing (e.g., Apple's Siri) and edge computing (e.g., Google/Waymo's driverless car).…

分布式、并行与集群计算 · 计算机科学 2020-02-11 Weiwen Jiang , Edwin H. -M. Sha , Xinyi Zhang , Lei Yang , Qingfeng Zhuge , Yiyu Shi , Jingtong Hu

Recent research on image denoising has progressed with the development of deep learning architectures, especially convolutional neural networks. However, real-world image denoising is still very challenging because it is not possible to…

图像与视频处理 · 电气工程与系统科学 2019-05-28 Dong-Wook Kim , Jae Ryun Chung , Seung-Won Jung

This report details MERL's system for room impulse response (RIR) estimation submitted to the Generative Data Augmentation Workshop at ICASSP 2025 for Augmenting RIR Data (Task 1) and Improving Speaker Distance Estimation (Task 2). We first…

音频与语音处理 · 电气工程与系统科学 2025-04-22 Christopher Ick , Gordon Wichern , Yoshiki Masuyama , François G. Germain , Jonathan Le Roux

Implicit neural representations (INRs) have recently emerged as a powerful tool that provides an accurate and resolution-independent encoding of data. Their robustness as general approximators has been shown in a wide variety of data…

Discrete Diffusion Language Models (DLMs) offer a promising non-autoregressive alternative for text generation, yet effective mechanisms for inference-time control remain relatively underexplored. Existing approaches include sampling-level…

计算与语言 · 计算机科学 2026-01-30 Eden Avrahami , Eliya Nachmani

This paper presents a novel approach for the joint design of a reconfigurable intelligent surface (RIS) and a transmitter-receiver pair that are trained together as a set of deep neural networks (DNNs) to optimize the end-to-end…

网络与互联网体系结构 · 计算机科学 2021-12-09 Tugba Erpek , Yalin E. Sagduyu , Ahmed Alkhateeb , Aylin Yener

Present-day Deep Reinforcement Learning (RL) systems show great promise towards building intelligent agents surpassing human-level performance. However, the computational complexity associated with the underlying deep neural networks (DNNs)…

机器学习 · 计算机科学 2021-09-20 Adarsh Kumar Kosta , Malik Aqeel Anwar , Priyadarshini Panda , Arijit Raychowdhury , Kaushik Roy

In this paper we introduce StoRIR - a stochastic room impulse response generation method dedicated to audio data augmentation in machine learning applications. This technique, in contrary to geometrical methods like image-source or ray…

音频与语音处理 · 电气工程与系统科学 2020-08-18 Piotr Masztalski , Mateusz Matuszewski , Karol Piaskowski , Michał Romaniuk

In this paper, we describe how to efficiently implement an acoustic room simulator to generate large-scale simulated data for training deep neural networks. Even though Google Room Simulator in [1] was shown to be quite effective in…

声音 · 计算机科学 2019-01-03 Chanwoo Kim , Ehsan Variani , Arun Narayanan , Michiel Bacchiani

Convolutional recurrent networks (CRN) integrating a convolutional encoder-decoder (CED) structure and a recurrent structure have achieved promising performance for monaural speech enhancement. However, feature representation across…

声音 · 计算机科学 2024-12-02 Shengkui Zhao , Bin Ma , Karn N. Watcharasupat , Woon-Seng Gan

Implicit Neural Representation (INR) is an innovative approach for representing complex shapes or objects without explicitly defining their geometry or surface structure. Instead, INR represents objects as continuous functions. Previous…

计算机视觉与模式识别 · 计算机科学 2024-04-25 Hanqiu Chen , Hang Yang , Stephen Fitzmeyer , Cong Hao

The characteristics of a sound field are intrinsically linked to the geometric and spatial properties of the environment surrounding a sound source and a listener. The physics of sound propagation is captured in a time-domain signal known…

音频与语音处理 · 电气工程与系统科学 2025-05-21 Christopher Ick , Gordon Wichern , Yoshiki Masuyama , François Germain , Jonathan Le Roux

Recurrent neural networks (RNNs) such as Long Short Term Memory (LSTM) networks have become popular in a variety of applications such as image processing, data classification, speech recognition, and as controllers in autonomous systems. In…

机器学习 · 计算机科学 2020-07-21 Sara Mohammadinejad , Brandon Paulsen , Chao Wang , Jyotirmoy V. Deshmukh

Deep neural networks are among the most widely applied machine learning tools showing outstanding performance in a broad range of tasks. We present a method for folding a deep neural network of arbitrary size into a single neuron with…

机器学习 · 计算机科学 2021-09-15 Florian Stelzer , André Röhm , Raul Vicente , Ingo Fischer , Serhiy Yanchuk

Fourier-encoded implicit neural representations (INRs) have shown strong capability in modeling continuous signals from discrete samples. However, conventional Fourier feature mappings use a fixed set of frequencies over the entire spatial…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Ligen Shi , Jun Qiu , Yuhang Zheng , Zengyu Pang , Chang Liu

In this paper, we present a generic and robust multimodal synthesis system that produces highly natural speech and facial expression simultaneously. The key component of this system is the Duration Informed Attention Network (DurIAN), an…

计算与语言 · 计算机科学 2019-09-09 Chengzhu Yu , Heng Lu , Na Hu , Meng Yu , Chao Weng , Kun Xu , Peng Liu , Deyi Tuo , Shiyin Kang , Guangzhi Lei , Dan Su , Dong Yu