English
Related papers

Related papers: EigeNet: Geometry-Informed Multi-Modal Learning fo…

200 papers

Load forecasting plays a pivotal role in the safe and stable operation of power systems. Conventional deep learning methods often struggle to adapt to few-shot scenarios frequently encountered in industrial applications. Existing…

Signal Processing · Electrical Eng. & Systems 2026-05-12 Yuxuan Chen , Shuo Dai , Ruoyi Xu , Haipeng Xie

Multi-frame infrared small target detection (IRSTD) plays a crucial role in low-altitude and maritime surveillance. The hybrid architecture combining CNNs and Transformers shows great promise for enhancing multi-frame IRSTD performance. In…

Computer Vision and Pattern Recognition · Computer Science 2025-03-05 Zhihua Shen , Siyang Chen , Han Wang , Tongsu Zhang , Xiaohu Zhang , Xiangpeng Xu , Xia Yang

The generation of room impulse responses (RIRs) using deep neural networks has attracted growing research interest due to its applications in virtual and augmented reality, audio postproduction, and related fields. Most existing approaches…

Sound · Computer Science 2025-07-17 Silvia Arellano , Chunghsin Yeh , Gautam Bhattacharya , Daniel Arteaga

This paper describes a novel Deep Learning method for the design of IIR parametric filters for automatic audio equalization. A simple and effective neural architecture, named BiasNet, is proposed to determine the IIR equalizer parameters.…

Audio and Speech Processing · Electrical Eng. & Systems 2021-10-06 Giovanni Pepe , Leonardo Gabrielli , Stefano Squartini , Carlo Tripodi , Nicolò Strozzi

In this paper, we explore a novel framework, EGIInet (Explicitly Guided Information Interaction Network), a model for View-guided Point cloud Completion (ViPC) task, which aims to restore a complete point cloud from a partial one with a…

Computer Vision and Pattern Recognition · Computer Science 2024-07-24 Hang Xu , Chen Long , Wenxiao Zhang , Yuan Liu , Zhen Cao , Zhen Dong , Bisheng Yang

Few-shot segmentation focuses on the generalization of models to segment unseen object instances with limited training samples. Although tremendous improvements have been achieved, existing methods are still constrained by two factors. (1)…

Computer Vision and Pattern Recognition · Computer Science 2021-10-26 Xianghui Yang , Bairun Wang , Kaige Chen , Xinchi Zhou , Shuai Yi , Wanli Ouyang , Luping Zhou

Target speech extraction remains difficult for compact devices because monaural neural models lack spatial evidence and classical beamformers lose resolving power when the microphone aperture is only a few centimetres. We present IsoNet, a…

Sound · Computer Science 2026-05-18 Dinanath Padhya , Sajen Maharjan , Binita Adhikari , Ishwor Raj Pokharel

Unsupervised learning for geometric perception (depth, optical flow, etc.) is of great interest to autonomous systems. Recent works on unsupervised learning have made considerable progress on perceiving geometry; however, they usually…

Computer Vision and Pattern Recognition · Computer Science 2019-04-08 Yue Meng , Yongxi Lu , Aman Raj , Samuel Sunarjo , Rui Guo , Tara Javidi , Gaurav Bansal , Dinesh Bharadia

Existing learning-based methods effectively reconstruct HDR images from multi-exposure LDR inputs with extended dynamic range and improved detail, but they rely more on empirical design rather than theoretical foundation, which can impact…

Image and Video Processing · Electrical Eng. & Systems 2025-07-08 Xinyue Li , Zhangkai Ni , Wenhan Yang

Electroencephalography (EEG) is a widely used non-invasive technique for monitoring brain activity, but low signal-to-noise ratios (SNR) due to various artifacts often compromise its utility. Conventional artifact removal methods require…

Signal Processing · Electrical Eng. & Systems 2025-11-06 Shantanu Sarkar , Piotr Nabrzyski , Saurabh Prasad , Jose Luis Contreras-Vidal

Prediction of room impulse responses (RIRs) is essential for room acoustics, spatial audio, and immersive applications, yet conventional simulations and measurements remain computationally expensive and time-consuming. This work proposes a…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-30 Imran Muhammad , Gerald Schuller

Deep image registration has demonstrated exceptional accuracy and fast inference. Recent advances have adopted either multiple cascades or pyramid architectures to estimate dense deformation fields in a coarse-to-fine manner. However, due…

Computer Vision and Pattern Recognition · Computer Science 2024-07-19 Xinxing Cheng , Xi Jia , Wenqi Lu , Qiufu Li , Linlin Shen , Alexander Krull , Jinming Duan

Surgical instrument segmentation is extremely important for computer-assisted surgery. Different from common object segmentation, it is more challenging due to the large illumination and scale variation caused by the special surgical…

Computer Vision and Pattern Recognition · Computer Science 2020-05-25 Zhen-Liang Ni , Gui-Bin Bian , Guan-An Wang , Xiao-Hu Zhou , Zeng-Guang Hou , Xiao-Liang Xie , Zhen Li , Yu-Han Wang

Scene observation from multiple perspectives would bring a more comprehensive visual experience. However, in the context of acquiring multiple views in the dark, the highly correlated views are seriously alienated, making it challenging to…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Hao Luo , Baoliang Chen , Lingyu Zhu , Peilin Chen , Shiqi Wang

Neural Radiance Field (NeRF) technology has made significant strides in creating novel viewpoints. However, its effectiveness is hampered when working with sparsely available views, often leading to performance dips due to overfitting.…

Computer Vision and Pattern Recognition · Computer Science 2025-09-10 Yuru Xiao , Xianming Liu , Deming Zhai , Kui Jiang , Junjun Jiang , Xiangyang Ji

For few-shot semantic segmentation, the primary task is to extract class-specific intrinsic information from limited labeled data. However, the semantic ambiguity and inter-class similarity of previous methods limit the accuracy of…

Computer Vision and Pattern Recognition · Computer Science 2023-12-12 Xiaoyi Bao , Jie Qin , Siyang Sun , Yun Zheng , Xingang Wang

Cardiac Cine Magnetic Resonance Imaging (MRI) provides an accurate assessment of heart morphology and function in clinical practice. However, MRI requires long acquisition times, with recent deep learning-based methods showing great promise…

Image and Video Processing · Electrical Eng. & Systems 2024-07-04 Siying Xu , Kerstin Hammernik , Andreas Lingg , Jens Kuebler , Patrick Krumm , Daniel Rueckert , Sergios Gatidis , Thomas Kuestner

Incremental Learning (IL) aims to learn new tasks while preserving previously acquired knowledge. Integrating the zero-shot learning capabilities of pre-trained vision-language models into IL methods has marked a significant advancement.…

Computer Vision and Pattern Recognition · Computer Science 2026-03-20 Haihua Luo , Xuming Ran , Jiangrong Shen , Timo Hämäläinen , Zhonghua Chen , Qi Xu , Fengyu Cong

Birds-eye-view (BEV) semantic segmentation is critical for autonomous driving for its powerful spatial representation ability. It is challenging to estimate the BEV semantic maps from monocular images due to the spatial gap, since it is…

Computer Vision and Pattern Recognition · Computer Science 2022-07-22 Shi Gong , Xiaoqing Ye , Xiao Tan , Jingdong Wang , Errui Ding , Yu Zhou , Xiang Bai

Majority models of remote sensing image changing detection can only get great effect in a specific resolution data set. With the purpose of improving change detection effectiveness of the model in the multi-resolution data set, a weighted…

Computer Vision and Pattern Recognition · Computer Science 2022-05-04 Yu Jiang , Lei Hu , Yongmei Zhang , Xin Yang