English
Related papers

Related papers: Self-Supervised Compression and Artifact Correctio…

200 papers

There is a growing interest in cost-effective self-supervised fine-tuning (SSFT) of self-supervised learning (SSL)-based speech models to obtain task-specific representations. These task-specific representations are used for robust…

Computation and Language · Computer Science 2024-03-12 Amit Meghanani , Thomas Hain

The rapid advancement of generative models has enabled highly realistic audio deepfakes, yet current detectors suffer from a critical bias problem, leading to poor generalization across unseen datasets. This paper proposes Artifact-Focused…

Media compression standards have reached a plateau in terms of the rate-distortion-complexity trade-off, limiting the ability to offload expensive AI perception to the cloud in applications like robotics, wearables, and remote sensing.…

Image and Video Processing · Electrical Eng. & Systems 2026-05-29 Dan Jacobellis , Neeraja J. Yadwadkar

With the continuous development of underwater vision technology, more and more remote sensing images could be obtained. In the underwater scene, sonar sensors are currently the most effective remote perception devices, and the sonar images…

Computer Vision and Pattern Recognition · Computer Science 2021-12-16 Xiaoteng Zhou , Changli Yu , Xin Yuan , Yi Wu , Haijun Feng , Citong Luo

Image restoration algorithms such as super resolution (SR) are indispensable pre-processing modules for object detection in low quality images. Most of these algorithms assume the degradation is fixed and known a priori. However, in…

Computer Vision and Pattern Recognition · Computer Science 2022-08-08 Ziteng Cui , Yingying Zhu , Lin Gu , Guo-Jun Qi , Xiaoxiao Li , Renrui Zhang , Zenghui Zhang , Tatsuya Harada

Acoustic scene classification (ASC) predominantly relies on supervised approaches. However, acquiring labeled data for training ASC models is often costly and time-consuming. Recently, self-supervised learning (SSL) has emerged as a…

Sound · Computer Science 2024-08-28 Yiqiang Cai , Shengchen Li , Xi Shao

Recent progress in synthetic aperture sonar (SAS) technology and processing has led to significant advances in underwater imaging, outperforming previously common approaches in both accuracy and efficiency. There are, however, inherent…

Numerical Analysis · Mathematics 2017-06-28 John McKay , Anne Gelb , Vishal Monga , Raghu Raj

Spatial frequency analysis and transforms serve a central role in most engineered image and video lossy codecs, but are rarely employed in neural network (NN)-based approaches. We propose a novel NN-based image coding framework that…

Image and Video Processing · Electrical Eng. & Systems 2023-01-04 Hyomin Choi , Fabien Racape , Shahab Hamidi-Rad , Mateen Ulhaq , Simon Feltman

Oceanic processes at fine scales are crucial yet difficult to observe accurately due to limitations in satellite and in-situ measurements. The Surface Water and Ocean Topography (SWOT) mission provides high-resolution Sea Surface Height…

Atmospheric and Oceanic Physics · Physics 2025-03-28 Eugenio Cutolo , Carlos Granero-Belinchon , Ptashanna Thiraux , Jinbo Wang , Ronan Fablet

Current camera image and signal processing pipelines (ISPs), including deep-trained versions, tend to apply a single filter that is uniformly applied to the entire image. This is despite the fact that most acquired camera images have…

Computer Vision and Pattern Recognition · Computer Science 2023-06-28 Chandrajit Bajaj , Yi Wang , Yunhao Yang

Side-scan sonar (SSS) is a lightweight acoustic sensor that is frequently deployed on autonomous underwater vehicles (AUVs) to provide high-resolution seafloor images. However, using side-scan images to perform simultaneous localization and…

Robotics · Computer Science 2023-12-22 Jun Zhang , Yiping Xie , Li Ling , John Folkesson

Automatic speech recognition (ASR) in the cloud allows the use of larger models and more powerful multi-channel signal processing front-ends compared to on-device processing. However, it also adds an inherent latency due to the transmission…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-16 Lukas Drude , Jahn Heymann , Andreas Schwarz , Jean-Marc Valin

Deep learning-based joint source-channel coding (DeepJSCC) has emerged as a promising technique in 6G for enhancing the efficiency and reliability of data transmission across diverse modalities, particularly in low signal-to-noise ratio…

Signal Processing · Electrical Eng. & Systems 2025-09-09 Kaiyi Chi , Yinghui He , Qianqian Yang , Zhiping Jiang , Yuanchao Shu , Zhiqin Wang , Jun Luo , Jiming Chen

Self-supervised learning has proved to be a powerful approach to learn image representations without the need of large labeled datasets. For underwater robotics, it is of great interest to design computer vision algorithms to improve…

Computer Vision and Pattern Recognition · Computer Science 2022-04-21 Alan Preciado-Grijalva , Bilal Wehbe , Miguel Bande Firvida , Matias Valdenegro-Toro

In real-world applications, such as sharing photos on social media platforms, images are always not only sub-sampled but also heavily compressed thus often containing various artefacts. Simple methods for enhancing the resolution of such…

Image and Video Processing · Electrical Eng. & Systems 2022-11-23 Hongming Luo , Fei Zhou , Guangsen Liao , Guoping Qiu

Recent progress in learning-based image compression has demonstrated that end-to-end optimization can substantially outperform traditional codecs by jointly learning compact latent representations and probabilistic entropy models. However,…

Image and Video Processing · Electrical Eng. & Systems 2026-03-12 Sofia Iliopoulou , Dimitris Ampeliotis , Athanassios Skodras

MR data are acquired in the frequency domain, known as k-space. Acquiring high-quality and high-resolution MR images can be time-consuming, posing a significant challenge when multiple sequences providing complementary contrast information…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Georgia Kanli , Daniele Perlo , Selma Boudissa , Radovan Jirik , Olivier Keunen

Scene Parsing is a crucial step to enable autonomous systems to understand and interact with their surroundings. Supervised deep learning methods have made great progress in solving scene parsing problems, however, come at the cost of…

Computer Vision and Pattern Recognition · Computer Science 2019-03-26 Keng-Chi Liu , Yi-Ting Shen , Jan P. Klopp , Liang-Gee Chen

Autoregressive (AR) video diffusion models enable long-form video generation but remain expensive due to repeated multi-step denoising. Existing training-free acceleration methods rely on binary cache-or-recompute decisions, overlooking…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Hanshuai Cui , Zhiqing Tang , Zhi Yao , Fanshuai Meng , Weijia Jia , Wei Zhao

Accelerated MRI involves collecting partial $k$-space measurements to reduce acquisition time, patient discomfort, and motion artifacts, and typically uses regular undersampling patterns or human-designed schemes. Recent works have studied…

Image and Video Processing · Electrical Eng. & Systems 2026-05-20 Siddhant Gautam , Angqi Li , Nicole Seiberlich , Jeffrey A. Fessler , Saiprasad Ravishankar