English
Related papers

Related papers: Audio Super-Resolution with Latent Bridge Models

200 papers

Snapshot compressive spectral imaging reconstruction aims to reconstruct three-dimensional spatial-spectral images from a single-shot two-dimensional compressed measurement. Existing state-of-the-art methods are mostly based on deep…

Image and Video Processing · Electrical Eng. & Systems 2024-08-27 Zongliang Wu , Ruiying Lu , Ying Fu , Xin Yuan

Deep learning-based dMRI super-resolution methods can effectively enhance image resolution by leveraging the learning capabilities of neural networks on large datasets. However, these methods tend to learn a fixed scale mapping between…

Image and Video Processing · Electrical Eng. & Systems 2024-08-15 Ruoyou Wu , Jian Cheng , Cheng Li , Juan Zou , Jing Yang , Wenxin Fan , Yong Liang , Shanshan Wang

Audio-Visual Speech Recognition (AVSR) integrates acoustic and visual information to enhance robustness in adverse acoustic conditions. Recent advances in Large Language Models (LLMs) have yielded competitive automatic speech recognition…

Sound · Computer Science 2026-03-05 Fei Su , Cancan Li , Juan Liu , Wei Ju , Hongbin Suo , Ming Li

In this paper, we consider the problem of multi-resolution compressed sensing (MR-CS) reconstruction, which has received little attention in the literature. Instead of always reconstructing the signal at the original high resolution (HR),…

Information Theory · Computer Science 2016-01-21 Xing Wang , Jie Liang

Consistency Models (CMs) have significantly accelerated the sampling process in diffusion models, yielding impressive results in synthesizing high-resolution images. To explore and extend these advancements to point-cloud-based 3D shape…

Computer Vision and Pattern Recognition · Computer Science 2025-12-30 Bi'an Du , Wei Hu , Renjie Liao

Speech Emotion Recognition (SER) is becoming a key role in global business today to improve service efficiency, like call center services. Recent SERs were based on a deep learning approach. However, the efficiency of deep learning depends…

We develop a large language model (LLM) based automatic speech recognition (ASR) system that can be contextualized by providing keywords as prior information in text prompts. We adopt decoder-only architecture and use our in-house LLM,…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-14 Kento Nozawa , Takashi Masuko , Toru Taniguchi

Video super-resolution (SR) aims at generating a sequence of high-resolution (HR) frames with plausible and temporally consistent details from their low-resolution (LR) counterparts. The key challenge for video SR lies in the effective…

Computer Vision and Pattern Recognition · Computer Science 2020-01-08 Longguang Wang , Yulan Guo , Li Liu , Zaiping Lin , Xinpu Deng , Wei An

Deep learning (DL) architectures for superresolution (SR) normally contain tremendous parameters, which has been regarded as the crucial advantage for obtaining satisfying performance. However, with the widespread use of mobile phones for…

Image and Video Processing · Electrical Eng. & Systems 2023-07-19 Biao Li , Jiabin Liu , Bo Wang , Zhiquan Qi , Yong Shi

The workload of real-time rendering is steeply increasing as the demand for high resolution, high refresh rates, and high realism rises, overwhelming most graphics cards. To mitigate this problem, one of the most popular solutions is to…

Graphics · Computer Science 2023-10-17 Zhihua Zhong , Jingsen Zhu , Yuxin Dai , Chuankun Zheng , Yuchi Huo , Guanlin Chen , Hujun Bao , Rui Wang

Self-supervised learning (SSL) models have become crucial in speech processing, with recent advancements concentrating on developing architectures that capture representations across multiple timescales. The primary goal of these…

Audio and Speech Processing · Electrical Eng. & Systems 2024-11-01 Theo Clark , Benedetta Cevoli , Eloy de Jong , Timofey Abramski , Jamie Dougherty

A low-resolution digital surface model (DSM) features distinctive attributes impacted by noise, sensor limitations and data acquisition conditions, which failed to be replicated using simple interpolation methods like bicubic. This causes…

Image and Video Processing · Electrical Eng. & Systems 2024-04-08 Daniel Panangian , Ksenia Bittner

Existing reference (RF)-based super-resolution (SR) models try to improve perceptual quality in SR under the assumption of the availability of high-resolution RF images paired with low-resolution (LR) inputs at testing. As the RF images…

Computer Vision and Pattern Recognition · Computer Science 2021-04-07 Mohammad Saeed Rad , Thomas Yu , Behzad Bozorgtabar , Jean-Philippe Thiran

Prior material creation methods had limitations in producing diverse results mainly because reconstruction-based methods relied on real-world measurements and generation-based methods were trained on relatively small material datasets. To…

Computer Vision and Pattern Recognition · Computer Science 2024-07-02 Linxuan Xin , Zheng Zhang , Jinfu Wei , Wei Gao , Duan Gao

Existing hologram super-resolution (HSR) methods primarily focus on angle-of-view expansion. Adapting them for volumetric spatial up-sampling introduces severe quadratic depth distortion, degrading 3D focal accuracy. We propose CV-HoloSR, a…

Graphics · Computer Science 2026-04-14 Youchan No , Jaehong Lee , Daejun Choi , Dae Youl Park , Duksu Kim

This article describes a density ratio approach to integrating external Language Models (LMs) into end-to-end models for Automatic Speech Recognition (ASR). Applied to a Recurrent Neural Network Transducer (RNN-T) ASR model trained on a…

Audio and Speech Processing · Electrical Eng. & Systems 2020-03-02 Erik McDermott , Hasim Sak , Ehsan Variani

Voice communication using the air conduction microphone in noisy environments suffers from the degradation of speech audibility. Bone conduction microphones (BCM) are robust against ambient noises but suffer from limited effective bandwidth…

Sound · Computer Science 2021-12-28 Yuang Li , Yuntao Wang , Xin Liu , Yuanchun Shi , Shao-fu Shih

While burst Low-Resolution (LR) images are useful for improving their Super Resolution (SR) image compared to a single LR image, prior burst SR methods are trained in a deterministic manner, which produces a blurry SR image. Since such…

Computer Vision and Pattern Recognition · Computer Science 2025-07-21 Kento Kawai , Takeru Oba , Kyotaro Tokoro , Kazutoshi Akita , Norimichi Ukita

Reference-based Image Super-Resolution (RefSR) aims to restore a low-resolution (LR) image by utilizing the semantic and texture information from an additional reference high-resolution (reference HR) image. Existing diffusion-based RefSR…

Computer Vision and Pattern Recognition · Computer Science 2025-08-15 Zhenning Shi , Zizheng Yan , Yuhang Yu , Clara Xue , Jingyu Zhuang , Qi Zhang , Jinwei Chen , Tao Li , Qingnan Fan

Recently, Mamba-based super-resolution (SR) methods have demonstrated the ability to capture global receptive fields with linear complexity, addressing the quadratic computational cost of Transformer-based SR approaches. However, existing…

Computer Vision and Pattern Recognition · Computer Science 2026-04-24 Sichen Guo , Wenjie Li , Yuanyang Liu , Guangwei Gao , Jian Yang , Chia-Wen Lin