English
Related papers

Related papers: SITS-DECO: A Generative Decoder Is All You Need Fo…

200 papers

Vision Transformers (ViTs) have shown remarkable performance and scalability across various computer vision tasks. To apply single-scale ViTs to image segmentation, existing methods adopt a convolutional adapter to generate multi-scale…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Tommie Kerssies , Niccolò Cavagnero , Alexander Hermans , Narges Norouzi , Giuseppe Averta , Bastian Leibe , Gijs Dubbelman , Daan de Geus

Due to strict rate and reliability demands, wireless image transmission remains difficult for both classical layered designs and joint source-channel coding (JSCC), especially under low latency. Diffusion-based generative decoders can…

Machine Learning · Computer Science 2026-01-13 Jingwen Fu , Ming Xiao , Mikael Skoglund , Dong In Kim

Recently diffusion models have shown improvement in synthetic image quality as well as better control in generation. We motivate and present Gen2Det, a simple modular pipeline to create synthetic training data for object detection for free…

Computer Vision and Pattern Recognition · Computer Science 2023-12-08 Saksham Suri , Fanyi Xiao , Animesh Sinha , Sean Chang Culatana , Raghuraman Krishnamoorthi , Chenchen Zhu , Abhinav Shrivastava

Massive MIMO systems can enhance spectral and energy efficiency, but they require accurate channel state information (CSI), which becomes costly as the number of antennas increases. While machine learning (ML) autoencoders show promise for…

Signal Processing · Electrical Eng. & Systems 2025-11-12 Hao Luo , Saeed R. Khosravirad , Ahmed Alkhateeb

Generative modelling paradigms based on denoising diffusion processes have emerged as a leading candidate for conditional sampling in inverse problems. In many real-world applications, we often have access to large, expensively trained…

Open-access multispectral imagery from missions like Landsat 8-9 and Sentinel-2 has fueled the development of geospatial foundation models (GFMs) for humanitarian and environmental applications. Yet, their deployment remains limited by (i)…

Computer Vision and Pattern Recognition · Computer Science 2025-10-08 Ibrahim Salihu Yusuf , Iffanice Houndayi , Rym Oualha , Mohamed Aziz Cherif , Kobby Panford-Quainoo , Arnu Pretorius

Residential floor plan generation requires not only geometric fidelity but also spatial configurational logic: shared living spaces should be integrative, while private spaces should remain segregated. Existing generators increasingly use…

Machine Learning · Computer Science 2026-05-13 Zhuoyang Jiang , Dongqing Zhang

The introduction of DETR represents a new paradigm for object detection. However, its decoder conducts classification and box localization using shared queries and cross-attention layers, leading to suboptimal results. We observe that…

Computer Vision and Pattern Recognition · Computer Science 2023-10-25 Manyuan Zhang , Guanglu Song , Yu Liu , Hongsheng Li

Semi-supervised Camouflaged Object Detection (SSCOD) aims to reduce reliance on costly pixel-level annotations by leveraging limited annotated data and abundant unlabeled data. However, existing SSCOD methods based on Teacher-Student…

Computer Vision and Pattern Recognition · Computer Science 2025-08-01 Xihang Hu , Fuming Sun , Jiazhe Liu , Feilong Xu , Xiaoli Zhang

Foundation models have the potential to transform the landscape of remote sensing (RS) data analysis by enabling large computer vision models to be pre-trained on vast amounts of remote sensing data. These models can then be fine-tuned with…

Computer Vision and Pattern Recognition · Computer Science 2024-11-27 Caleb S. Spradlin , Jordan A. Caraballo-Vega , Jian Li , Mark L. Carroll , Jie Gong , Paul M. Montesano

Earth observation (EO) data features diverse sensing platforms with varying spectral bands, spatial resolutions, and sensing modalities. While most prior work has constrained inputs to fixed sensors, a new class of any-sensor foundation…

Machine Learning · Computer Science 2025-08-04 Leonard Waldmann , Ando Shah , Yi Wang , Nils Lehmann , Adam J. Stewart , Zhitong Xiong , Xiao Xiang Zhu , Stefan Bauer , John Chuang

Foundation models (FoMos), referring to large-scale AI models, possess human-like capabilities and are able to perform competitively in the domain of human intelligence. The breakthrough in FoMos has inspired researchers to deploy such…

Networking and Internet Architecture · Computer Science 2023-10-31 Hai Wu , Xu Chen , Kaibin Huang

The ubiquity of time series data creates a strong demand for general-purpose foundation models, yet developing them for classification remains a significant challenge, largely due to the high cost of labeled data. Foundation models capable…

Machine Learning · Computer Science 2025-11-27 Chin-Chia Michael Yeh , Uday Singh Saini , Junpeng Wang , Xin Dai , Xiran Fan , Jiarui Sun , Yujie Fan , Yan Zheng

Low-shot object counters estimate the number of objects in an image using few or no annotated exemplars. Objects are localized by matching them to prototypes, which are constructed by unsupervised image-wide object appearance aggregation.…

Computer Vision and Pattern Recognition · Computer Science 2024-12-02 Jer Pelhan , Alan Lukežič , Vitjan Zavrtanik , Matej Kristan

Generative models have recently gained increasing attention in image generation and editing tasks. However, they often lack a direct connection to object geometry, which is crucial in sensitive domains such as computational anatomy,…

Graphics · Computer Science 2025-04-14 Nian Wu , Nivetha Jayakumar , Jiarui Xing , Miaomiao Zhang

Cross-view object Geo-localization aims to precisely pinpoint the same object across large-scale satellite imagery based on drone images. Due to significant differences in viewpoint and scale, coupled with complex background interference,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-19 Fan Zhang , Haoyuan Ren , Fei Ma , Qiang Yin , Yongsheng Zhou

In this paper, we propose YOSO, a real-time panoptic segmentation framework. YOSO predicts masks via dynamic convolutions between panoptic kernels and image feature maps, in which you only need to segment once for both instance and semantic…

Computer Vision and Pattern Recognition · Computer Science 2023-03-28 Jie Hu , Linyan Huang , Tianhe Ren , Shengchuan Zhang , Rongrong Ji , Liujuan Cao

Understanding dynamic 3D environments is essential for safe autonomous driving, particularly when reasoning about human-centric, nonrigid agents. However, existing weakly supervised occupancy prediction frameworks predominantly assume…

Computer Vision and Pattern Recognition · Computer Science 2026-05-28 Yang Gao , Wuyang Li , Po-Chien Luan , Alexandre Alahi

Recently, some works have tried to combine diffusion and Generative Adversarial Networks (GANs) to alleviate the computational cost of the iterative denoising inference in Diffusion Models (DMs). However, existing works in this line suffer…

Computer Vision and Pattern Recognition · Computer Science 2025-02-26 Yihong Luo , Xiaolong Chen , Xinghua Qu , Tianyang Hu , Jing Tang

The development of analytical software for big Earth observation data faces several challenges. Designers need to balance between conflicting factors. Solutions that are efficient for specific hardware architectures can not be used in other…

‹ Prev 1 4 5 6 7 8 10 Next ›