中文
相关论文

相关论文: EigeNet: Geometry-Informed Multi-Modal Learning fo…

200 篇论文

Deep neural networks for semantic segmentation rely on large-scale annotated datasets, leading to an annotation bottleneck that motivates few shot semantic segmentation (FSS) which aims to generalize to novel classes with minimal labeled…

计算机视觉与模式识别 · 计算机科学 2025-12-17 Ourui Fu , Hangzhou He , Kaiwen Li , Xinliang Zhang , Lei Zhu , Shuang Zeng , Zhaoheng Xie , Yanye Lu

In the field of human-computer interaction and psychological assessment, speech emotion recognition (SER) plays an important role in deciphering emotional states from speech signals. Despite advancements, challenges persist due to system…

声音 · 计算机科学 2025-02-04 Alaa Nfissi , Wassim Bouachir , Nizar Bouguila , Brian Mishara

Referring Image Segmentation (RIS) aims to segment an object described in natural language from an image, with the main challenge being a text-to-pixel correlation. Previous methods typically rely on single-modality features, such as vision…

计算机视觉与模式识别 · 计算机科学 2024-05-21 Yichen Yan , Xingjian He , Sihan Chen , Shichen Lu , Jing Liu

Multi-view stereo omnidirectional distance estimation usually needs to build a cost volume with many hypothetical distance candidates. The cost volume building process is often computationally heavy considering the limited resources a…

计算机视觉与模式识别 · 计算机科学 2024-05-10 Conner Pulling , Je Hon Tan , Yaoyu Hu , Sebastian Scherer

Accurate segmentation is crucial for clinical applications, but existing models often assume fixed, high-resolution inputs and degrade significantly when faced with lower-resolution data in real-world scenarios. To address this limitation,…

图像与视频处理 · 电气工程与系统科学 2025-07-22 Simon Winther Albertsen , Hjalte Svaneborg Bjørnstrup , Mostafa Mehdipour Ghazi

Advances in deep learning techniques have allowed recent work to reconstruct the shape of a single object given only one RBG image as input. Building on common encoder-decoder architectures for this task, we propose three extensions: (1)…

计算机视觉与模式识别 · 计算机科学 2020-08-06 Stefan Popov , Pablo Bauszat , Vittorio Ferrari

In the fifth-generation communication system (5G), multipath-assisted positioning (MAP) has emerged as a promising approach. With the enhancement of signal resolution, multipath component (MPC) are no longer regarded as noise but rather as…

信号处理 · 电气工程与系统科学 2025-06-05 Ye Tian , Xueting Xu , Ao Peng

Visible-Infrared person re-identification (VI-ReID) is a challenging matching problem due to large modality varitions between visible and infrared images. Existing approaches usually bridge the modality gap with only feature-level…

计算机视觉与模式识别 · 计算机科学 2021-02-25 Haojie Liu , Shun Ma , Daoxun Xia , Shaozi Li

Recent progress in semantic segmentation is driven by deep Convolutional Neural Networks and large-scale labeled image datasets. However, data labeling for pixel-wise segmentation is tedious and costly. Moreover, a trained model can only…

计算机视觉与模式识别 · 计算机科学 2019-03-07 Chi Zhang , Guosheng Lin , Fayao Liu , Rui Yao , Chunhua Shen

Infrared small target detection (IRSTD) remains challenging due to the scarcity of useful target cues and the presence of severe background clutter. Most current methods rely on conventional feature learning and local interaction modeling,…

计算机视觉与模式识别 · 计算机科学 2026-05-11 Qianwen Ma , Yang Xu , Shangwei Deng , Xiaobo Li , Haofeng Hu

A good motion retargeting cannot be reached without reasonable consideration of source-target differences on both the skeleton and shape geometry levels. In this work, we propose a novel Residual RETargeting network (R2ET) structure, which…

计算机视觉与模式识别 · 计算机科学 2023-03-16 Jiaxu Zhang , Junwu Weng , Di Kang , Fang Zhao , Shaoli Huang , Xuefei Zhe , Linchao Bao , Ying Shan , Jue Wang , Zhigang Tu

Rings like gold, thuds like wood! The sound we hear in a scene is shaped not only by the spatial layout of the environment but also by the materials of the objects and surfaces within it. For instance, a room with wooden walls will produce…

计算机视觉与模式识别 · 计算机科学 2026-04-24 Mahnoor Fatima Saad , Sagnik Majumder , Kristen Grauman , Ziad Al-Halah

Radiology report generation (RRG) has gained increasing research attention because of its huge potential to mitigate medical resource shortages and aid the process of disease decision making by radiologists. Recent advancements in RRG are…

计算机视觉与模式识别 · 计算机科学 2024-03-05 Jun Wang , Abhir Bhalerao , Terry Yin , Simon See , Yulan He

We propose im2nerf, a learning framework that predicts a continuous neural object representation given a single input image in the wild, supervised by only segmentation output from off-the-shelf recognition methods. The standard approach to…

计算机视觉与模式识别 · 计算机科学 2022-09-12 Lu Mi , Abhijit Kundu , David Ross , Frank Dellaert , Noah Snavely , Alireza Fathi

High-resolution medical images can provide more detailed information for better diagnosis. Conventional medical image super-resolution relies on a single task which first performs the extraction of the features and then upscaling based on…

图像与视频处理 · 电气工程与系统科学 2025-04-25 Xiaoyan Kui , Zexin Ji , Beiji Zou , Yang Li , Yulan Dai , Liming Chen , Pierre Vera , Su Ruan

This work presents SGCDet, a novel multi-view indoor 3D object detection framework based on adaptive 3D volume construction. Unlike previous approaches that restrict the receptive field of voxels to fixed locations on images, we introduce a…

计算机视觉与模式识别 · 计算机科学 2025-07-25 Runmin Zhang , Zhu Yu , Si-Yuan Cao , Lingyu Zhu , Guangyi Zhang , Xiaokai Bai , Hui-Liang Shen

One of the main challenges since the advancement of convolutional neural networks is how to connect the extracted feature map to the final classification layer. VGG models used two sets of fully connected layers for the classification part…

计算机视觉与模式识别 · 计算机科学 2024-03-12 Mohammad Rahimzadeh , AmirAli Askari , Soroush Parvin , Elnaz Safi , Mohammad Reza Mohammadi

In this paper, we address the challenge of generating novel views of real-world objects with limited multi-view images through our proposed approach, FewShotNeRF. Our method utilizes meta-learning to acquire optimal initialization,…

计算机视觉与模式识别 · 计算机科学 2024-08-12 Piraveen Sivakumar , Paul Janson , Jathushan Rajasegaran , Thanuja Ambegoda

Room impulse responses are a core resource for dereverberation, robust speech recognition, source localization, and room acoustics estimation. We present RIR-Mega, a large collection of simulated RIRs described by a compact, machine…

音频与语音处理 · 电气工程与系统科学 2025-10-29 Mandip Goswami

The phase shift information (PSI) overhead poses a critical challenge to enabling real-time intelligent reflecting surface (IRS)-assisted wireless systems, particularly under dynamic and resource-constrained conditions. In this paper, we…

信号处理 · 电气工程与系统科学 2025-05-08 Xianhua Yu , Dong Li , Bowen Gu , Xiaoye Jing , Wen Wu , Tuo Wu , Kan Yu