English
Related papers

Related papers: Multi-Scale Implicit Transformer with Re-parameter…

200 papers

In multimodal unsupervised image-to-image translation tasks, the goal is to translate an image from the source domain to many images in the target domain. We present a simple method that produces higher quality images than current…

Computer Vision and Pattern Recognition · Computer Science 2019-04-09 Yazeed Alharbi , Neil Smith , Peter Wonka

Rate-Splitting (RS) has recently been shown to provide significant performance benefits in various multi-user transmission scenarios. In parallel, the huge degrees-of-freedom provided by the appealing massive Multiple-Input Multiple-Output…

Information Theory · Computer Science 2017-04-24 Anastasios Papazafeiropoulos , Bruno Clerckx , Tharmalingam Ratnarajah

Implicit neural representation (INR) models signals as continuous functions using neural networks, offering efficient and differentiable optimization for inverse problems across diverse disciplines. However, the representational capacity of…

Computer Vision and Pattern Recognition · Computer Science 2025-11-14 Zhicheng Cai , Hao Zhu , Linsen Chen , Qiu Shen , Xun Cao

Multi-contrast magnetic resonance imaging (MRI) is widely used in clinical practice as each contrast provides complementary information. However, the availability of each imaging contrast may vary amongst patients, which poses challenges to…

Image and Video Processing · Electrical Eng. & Systems 2023-03-31 Jiang Liu , Srivathsa Pasumarthi , Ben Duffy , Enhao Gong , Keshav Datta , Greg Zaharchuk

We introduce dual-decoder Transformer, a new model architecture that jointly performs automatic speech recognition (ASR) and multilingual speech translation (ST). Our models are based on the original Transformer architecture (Vaswani et…

Computation and Language · Computer Science 2020-11-21 Hang Le , Juan Pino , Changhan Wang , Jiatao Gu , Didier Schwab , Laurent Besacier

This paper presents a new Vision Transformer (ViT) architecture Multi-Scale Vision Longformer, which significantly enhances the ViT of \cite{dosovitskiy2020image} for encoding high-resolution images using two techniques. The first is the…

Computer Vision and Pattern Recognition · Computer Science 2021-05-28 Pengchuan Zhang , Xiyang Dai , Jianwei Yang , Bin Xiao , Lu Yuan , Lei Zhang , Jianfeng Gao

Symbol detection for Massive Multiple-Input Multiple-Output (MIMO) is a challenging problem for which traditional algorithms are either impractical or suffer from performance limitations. Several recently proposed learning-based approaches…

Signal Processing · Electrical Eng. & Systems 2019-06-12 Mehrdad Khani , Mohammad Alizadeh , Jakob Hoydis , Phil Fleming

Structural magnetic resonance imaging (sMRI) provides accurate estimates of the brain's structural organization and learning invariant brain representations from sMRI is an enduring issue in neuroscience. Previous deep representation…

Computer Vision and Pattern Recognition · Computer Science 2023-06-21 Ning Jiang , Gongshu Wang , Tianyi Yan

This paper investigates the inverse capabilities and broader utility of multimodal latent spaces within task-specific AI (Artificial Intelligence) models. While these models excel at their designed forward tasks (e.g., text-to-image…

Machine Learning · Computer Science 2025-08-01 Siwoo Park

This paper proposes a modification to RNN-Transducer (RNN-T) models for automatic speech recognition (ASR). In standard RNN-T, the emission of a blank symbol consumes exactly one input frame; in our proposed method, we introduce additional…

Audio and Speech Processing · Electrical Eng. & Systems 2024-04-15 Hainan Xu , Fei Jia , Somshubra Majumdar , Shinji Watanabe , Boris Ginsburg

Diffusion-based image super-resolution (SR) methods are mainly limited by the low inference speed due to the requirements of hundreds or even thousands of sampling steps. Existing acceleration sampling techniques inevitably sacrifice…

Computer Vision and Pattern Recognition · Computer Science 2023-10-19 Zongsheng Yue , Jianyi Wang , Chen Change Loy

Implicit representations are widely used for object reconstruction due to their efficiency and flexibility. In 2021, a novel structure named neural implicit map has been invented for incremental reconstruction. A neural implicit map…

Computer Vision and Pattern Recognition · Computer Science 2022-06-20 Yijun Yuan , Andreas Nuechter

Existing frameworks for image stitching often provide visually reasonable stitchings. However, they suffer from blurry artifacts and disparities in illumination, depth level, etc. Although the recent learning-based stitchings relax such…

Computer Vision and Pattern Recognition · Computer Science 2024-01-23 Minsu Kim , Jaewon Lee , Byeonghun Lee , Sunghoon Im , Kyong Hwan Jin

Purpose: Earth system models (ESMs) integrate the interactions of the atmosphere, ocean, land, ice, and biosphere to estimate the state of regional and global climate under a wide variety of conditions. The ESMs are highly complex; thus,…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Ehsan Zeraatkar , Salah Faroughi , Jelena Tešić

Implicit neural representations are a promising new avenue of representing general signals by learning a continuous function that, parameterized as a neural network, maps the domain of a signal to its codomain; the mapping from spatial…

Machine Learning · Computer Science 2021-11-09 Jaeho Lee , Jihoon Tack , Namhoon Lee , Jinwoo Shin

This paper studies the fair transmission design for an intelligent reflecting surface (IRS) aided rate-splitting multiple access (RSMA). IRS is used to establish a good signal propagation environment and enhance the RSMA transmission…

Information Theory · Computer Science 2024-03-18 Shanshan Zhang , Wen Chen , Qingqing Wu , Ziwei Liu , Shunqing Zhang , Jun Li

Rate-splitting multiple access (RSMA) has been studied for multiuser multiple-input multiple-output (MUMIMO) systems especially in the presence of imperfect channel state information (CSI) at the transmitter. However, its precoding designs…

Information Theory · Computer Science 2025-12-05 Wentao Zhou , Yijie Mao , Di Zhang , Mérouane Debbah , Inkyu Lee

Learning discrete representations of data is a central machine learning task because of the compactness of the representations and ease of interpretation. The task includes clustering and hash learning as special cases. Deep neural networks…

Machine Learning · Statistics 2017-06-15 Weihua Hu , Takeru Miyato , Seiya Tokui , Eiichi Matsumoto , Masashi Sugiyama

Referring Remote Sensing Image Segmentation (RRSIS) is a new challenge that combines computer vision and natural language processing, delineating specific regions in aerial images as described by textual queries. Traditional Referring Image…

Computer Vision and Pattern Recognition · Computer Science 2024-04-03 Sihan Liu , Yiwei Ma , Xiaoqing Zhang , Haowei Wang , Jiayi Ji , Xiaoshuai Sun , Rongrong Ji

For Image Super-Resolution (SR), it is common to train and evaluate scale-specific models composed of an encoder and upsampler for each targeted scale. Consequently, many SR studies encounter substantial training times and complex…

Computer Vision and Pattern Recognition · Computer Science 2024-11-19 Dongheon Lee , Seokju Yun , Youngmin Ro
‹ Prev 1 8 9 10 Next ›