English
Related papers

Related papers: Fusion Segment Transformer: Bi-Directional Attenti…

200 papers

The way features propagate in Fully Convolutional Networks is of momentous importance to capture multi-scale contexts for obtaining precise segmentation masks. This paper proposes a novel series-parallel hybrid paradigm called the Chained…

Computer Vision and Pattern Recognition · Computer Science 2021-09-16 Quan Tang , Fagui Liu , Tong Zhang , Jun Jiang , Yu Zhang

Despite significant progress in 3D object detection, point clouds remain challenging due to sparse data, incomplete structures, and limited semantic information. Capturing contextual relationships between distant objects presents additional…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Md Sohag Mia , Md Nahid Hasan , Muhammad Abdullah Adnan

Medical image segmentation - the prerequisite of numerous clinical needs - has been significantly prospered by recent advances in convolutional neural networks (CNNs). However, it exhibits general limitations on modeling explicit long-range…

Computer Vision and Pattern Recognition · Computer Science 2021-07-13 Yundong Zhang , Huiye Liu , Qiang Hu

Medical image segmentation faces challenges due to variations in anatomical structures. While convolutional neural networks (CNNs) effectively capture local features, they struggle with modeling long-range dependencies. Transformers…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Lalit Maurya , Honghai Liu , Reyer Zwiggelaar

Generative audio technologies now enable highly realistic voice cloning and real-time voice conversion, increasing the risk of impersonation, fraud, and misinformation in communication channels such as phone and video calls. This study…

Sound · Computer Science 2026-01-09 Prajwal Chinchmalatpure , Suyash Chinchmalatpure , Siddharth Chavan

The escalating challenges of managing vast sensor-generated data, particularly in audio applications, necessitate innovative solutions. Current systems face significant computational and storage demands, especially in real-time applications…

Modern semantic segmentation frameworks usually combine low-level and high-level features from pre-trained backbone convolutional models to boost performance. In this paper, we first point out that a simple fusion of low-level and…

Computer Vision and Pattern Recognition · Computer Science 2018-04-12 Zhenli Zhang , Xiangyu Zhang , Chao Peng , Dazhi Cheng , Jian Sun

Recent advances in Text-to-Speech (TTS) and Voice-Conversion (VC) using generative Artificial Intelligence (AI) technology have made it possible to generate high-quality and realistic human-like audio. This poses growing challenges in…

Sound · Computer Science 2025-03-25 Xiang Li , Pin-Yu Chen , Wenqi Wei

We present a lightweight latent diffusion model for vocal-conditioned musical accompaniment generation that addresses critical limitations in existing music AI systems. Our approach introduces a novel soft alignment attention mechanism that…

Sound · Computer Science 2026-01-06 Hei Shing Cheung , Boya Zhang , Jonathan H. Chan

The advent of high-resolution multispectral/hyperspectral sensors, LiDAR DSM (Digital Surface Model) information and many others has provided us with an unprecedented wealth of data for Earth Observation. Multimodal AI seeks to exploit…

Computer Vision and Pattern Recognition · Computer Science 2023-07-10 Nhi Kieu , Kien Nguyen , Sridha Sridharan , Clinton Fookes

With the rapid development of deep generative models (such as Generative Adversarial Networks and Diffusion models), AI-synthesized images are now of such high quality that humans can hardly distinguish them from pristine ones. Although…

Computer Vision and Pattern Recognition · Computer Science 2023-09-06 Yan Ju , Shan Jia , Jialing Cai , Haiying Guan , Siwei Lyu

The proliferation of AI-generated media poses significant challenges to information authenticity and social trust, making reliable detection methods highly demanded. Methods for detecting AI-generated media have evolved rapidly, paralleling…

Computer Vision and Pattern Recognition · Computer Science 2025-02-13 Yueying Zou , Peipei Li , Zekun Li , Huaibo Huang , Xing Cui , Xuannan Liu , Chenghanyu Zhang , Ran He

The malicious misuse and widespread dissemination of AI-generated images pose a significant threat to the authenticity of online information. Current detection methods often struggle to generalize to unseen generative models, and the rapid…

Computer Vision and Pattern Recognition · Computer Science 2026-01-12 Hanyi Wang , Jun Lan , Yaoyu Kang , Huijia Zhu , Weiqiang Wang , Zhuosheng Zhang , Shilin Wang

The rapid advancement of generative AI has enabled the creation of highly realistic forged facial images, posing significant threats to AI security, digital media integrity, and public trust. Face forgery techniques, ranging from face…

Computer Vision and Pattern Recognition · Computer Science 2025-10-29 Xin Zhang , Yuqi Song , Fei Zuo

Deepfake media is becoming widespread nowadays because of the easily available tools and mobile apps which can generate realistic looking deepfake videos/images without requiring any technical knowledge. With further advances in this field…

Computer Vision and Pattern Recognition · Computer Science 2022-08-12 Sohail Ahmed Khan , Duc-Tien Dang-Nguyen

Infrared and visible image fusion, as a hot topic in image processing and image enhancement, aims to produce fused images retaining the detail texture information in visible images and the thermal radiation information in infrared images. A…

Image and Video Processing · Electrical Eng. & Systems 2021-04-15 Zixiang Zhao , Jiangshe Zhang , Shuang Xu , Kai Sun , Chunxia Zhang , Junmin Liu

Infrared and visible image fusion is a powerful technique that combines complementary information from different modalities for downstream semantic perception tasks. Existing learning-based methods show remarkable performance, but are…

Computer Vision and Pattern Recognition · Computer Science 2023-08-09 Zhu Liu , Jinyuan Liu , Benzhuang Zhang , Long Ma , Xin Fan , Risheng Liu

In this article, we present our vision of preamble detection in a physical random access channel for next-generation (Next-G) networks using machine learning techniques. Preamble detection is performed to maintain communication and…

Networking and Internet Architecture · Computer Science 2022-04-25 Sunder Ali Khowaja , Kapal Dev , Parus Khuwaja , Quoc-Viet Pham , Nawab Muhammad Faseeh Qureshi , Paolo Bellavista , Maurizio Magarini

We study the merit of transfer learning for two sound recognition problems, i.e., audio tagging and sound event detection. Employing feature fusion, we adapt a baseline system utilizing only spectral acoustic inputs to also make use of…

Audio and Speech Processing · Electrical Eng. & Systems 2022-09-27 Wim Boes , Hugo Van hamme

Accurate detection of obstacles in 3D is an essential task for autonomous driving and intelligent transportation. In this work, we propose a general multimodal fusion framework FusionPainting to fuse the 2D RGB image and 3D point clouds at…

Computer Vision and Pattern Recognition · Computer Science 2021-08-11 Shaoqing Xu , Dingfu Zhou , Jin Fang , Junbo Yin , Zhou Bin , Liangjun Zhang
‹ Prev 1 8 9 10 Next ›