English
Related papers

Related papers: C-RADIOv4 (Tech Report)

200 papers

Convolutional Neural Networks (ConvNets) at present achieve remarkable performance in image classification tasks. However, current ConvNets cannot guarantee the capabilities of the mammalian visual systems such as invariance to contrast and…

Computer Vision and Pattern Recognition · Computer Science 2021-09-16 E. Ulises Moya-Sánchez , Sebastiá Xambo-Descamps , Abraham Sánchez , Sebastián Salazar-Colores , Ulises Cortés

Channel state information (CSI) reporting is important for multiple-input multiple-output (MIMO) transmitters to achieve high capacity and energy efficiency in frequency division duplex (FDD) mode. CSI reporting for massive MIMO systems…

Information Theory · Computer Science 2019-12-24 Zhenyu Liu , Lin Zhang , Zhi Ding

Real-time object detection has achieved substantial progress through meticulously designed architectures and optimization strategies. However, the pursuit of high-speed inference via lightweight network designs often leads to degraded…

Computer Vision and Pattern Recognition · Computer Science 2025-10-30 Zijun Liao , Yian Zhao , Xin Shan , Yu Yan , Chang Liu , Lei Lu , Xiangyang Ji , Jie Chen

This paper introduces ConvShareViT, a novel deep learning architecture that adapts Vision Transformers (ViTs) to the 4f free-space optical system. ConvShareViT replaces linear layers in multi-head self-attention (MHSA) and Multilayer…

Computer Vision and Pattern Recognition · Computer Science 2025-04-17 Riad Ibadulla , Thomas M. Chen , Constantino Carlos Reyes-Aldasoro

Most current medical vision language models struggle to jointly generate diagnostic text and pixel-level segmentation masks in response to complex visual questions. This represents a major limitation towards clinical application, as…

Computer Vision and Pattern Recognition · Computer Science 2025-10-22 Chengrun Li , Corentin Royer , Haozhe Luo , Bastian Wittmann , Xia Li , Ibrahim Hamamci , Sezgin Er , Anjany Sekuboyina , Bjoern Menze

The RD50-MPW4 is the latest HV-CMOS pixel sensor from the CERN-RD50-CMOS working group, designed to evaluate the HV-CMOS technology in terms of spatial resolution, radiation hardness and timing performance. Fabricated by LFoundry using a…

Albeit recent progress in speaker verification generates powerful models, malicious attacks in the form of spoofed speech, are generally not coped with. Recent results in ASVSpoof2015 and BTAS2016 challenges indicate that spoof-aware…

Audio and Speech Processing · Electrical Eng. & Systems 2020-07-28 Heinrich Dinkel , Nanxin Chen , Yanmin Qian , Kai Yu

In recent years, there have been numerous developments towards solving multimodal tasks, aiming to learn a stronger representation than through a single modality. Certain aspects of the data can be particularly useful in this case - for…

Machine Learning · Statistics 2023-09-06 Cătălina Cangea , Petar Veličković , Pietro Liò

In 6G mobile communications, acquiring accurate and timely channel state information (CSI) becomes increasingly challenging due to the growing antenna array size and bandwidth. To alleviate the CSI feedback burden, the channel knowledge map…

Signal Processing · Electrical Eng. & Systems 2026-03-11 Kequan Zhou , Guangyi Zhang , Hanlei Li , Yunlong Cai , Shengli Liu , Guanding Yu

Representations in the auditory cortex might be based on mechanisms similar to the visual ventral stream; modules for building invariance to transformations and multiple layers for compositionality and selectivity. In this paper we propose…

The development of multi-label deep learning models for retinal disease classification is often hindered by the scarcity of large, expertly annotated clinical datasets due to patient privacy concerns and high costs. The recent release of…

Computer Vision and Pattern Recognition · Computer Science 2025-08-25 Jerry Cao-Xue , Tien Comlekoglu , Keyi Xue , Guanliang Wang , Jiang Li , Gordon Laurie

Self-supervision has shown great potential for audio-visual speech recognition by vastly reducing the amount of labeled data required to build good systems. However, existing methods are either not entirely end-to-end or do not train joint…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-23 Jiachen Lian , Alexei Baevski , Wei-Ning Hsu , Michael Auli

We propose a new technique for multiple-input multiple-output (MIMO) radar with colocated antennas which we call phased-MIMO radar. The new technique enjoys the advantages of MIMO radar without sacrificing the main advantage of phased-array…

Information Theory · Computer Science 2015-05-13 Aboulnasr Hassanien , Sergiy A. Vorobyov

We open-source MiMo-VL-7B-SFT and MiMo-VL-7B-RL, two powerful vision-language models delivering state-of-the-art performance in both general visual understanding and multimodal reasoning. MiMo-VL-7B-RL outperforms Qwen2.5-VL-7B on 35 out of…

We introduce a new convolutional AutoEncoder architecture for user modelling and recommendation tasks with several improvements over the state of the art. Firstly, our model has the flexibility to learn a set of associations and…

Machine Learning · Computer Science 2025-09-10 Antoine Ledent , Petr Kasalický , Rodrigo Alves , Hady W. Lauw

Building upon Diff-A-Riff, a latent diffusion model for musical instrument accompaniment generation, we present a series of improvements targeting quality, diversity, inference speed, and text-driven control. First, we upgrade the…

Sound · Computer Science 2024-10-31 Javier Nistal , Marco Pasini , Stefan Lattner

Recent advances in pre-trained vision transformers have shown promise in parameter-efficient audio-visual learning without audio pre-training. However, few studies have investigated effective methods for aligning multimodal features in…

Computer Vision and Pattern Recognition · Computer Science 2024-06-10 Tanvir Mahmud , Shentong Mo , Yapeng Tian , Diana Marculescu

High resolution automotive radar sensors are required in order to meet the high bar of autonomous vehicles needs and regulations. However, current radar systems are limited in their angular resolution causing a technological gap. An…

Signal Processing · Electrical Eng. & Systems 2022-06-07 Itai Orr , Moshik Cohen , Harel Damari , Meir Halachmi , Zeev Zalevsky

This paper presents RadioGUNet, a UNet-based deep learning framework for pathloss estimation in wireless communication. Unlike other frameworks, it leverages group equivariant convolutional networks, which are known to increase the…

Networking and Internet Architecture · Computer Science 2025-11-25 Ziyue Yang , Feng Liu , Yifei Jin , Konstantinos Vandikas

Recent progress in Multimodal Large Language Models (MLLMs) has highlighted the critical roles of both the visual backbone and the underlying language model. While prior work has primarily focused on scaling these components to billions of…

Computer Vision and Pattern Recognition · Computer Science 2025-08-01 Federico Cocchi , Nicholas Moratelli , Davide Caffagni , Sara Sarto , Lorenzo Baraldi , Marcella Cornia , Rita Cucchiara