English
Related papers

Related papers: Transformer Decoders with MultiModal Regularizatio…

200 papers

Human intelligence is multimodal; we integrate visual, linguistic, and acoustic signals to maintain a holistic worldview. Most current pretraining methods, however, are limited to one or two modalities. We present i-Code, a self-supervised…

We introduce V-Trans4Style, an innovative algorithm tailored for dynamic video content editing needs. It is designed to adapt videos to different production styles like documentaries, dramas, feature films, or a specific YouTube channel's…

Computer Vision and Pattern Recognition · Computer Science 2025-01-15 Pooja Guhan , Tsung-Wei Huang , Guan-Ming Su , Subhadra Gopalakrishnan , Dinesh Manocha

In the field of cross-modal retrieval, single encoder models tend to perform better than dual encoder models, but they suffer from high latency and low throughput. In this paper, we present a dual encoder model called BagFormer that…

Information Retrieval · Computer Science 2023-01-02 Haowen Hou , Xiaopeng Yan , Yigeng Zhang , Fengzong Lian , Zhanhui Kang

Moving Object Detection (MOD) is a crucial task for the Autonomous Driving pipeline. MOD is usually handled via 2-stream convolutional architectures that incorporates both appearance and motion cues, without considering the inter-relations…

Computer Vision and Pattern Recognition · Computer Science 2021-06-23 Eslam Mohamed , Ahmad El-Sallab

Recently, deep learning methods have been widely used for tumor segmentation of multimodal medical images with promising results. However, most existing methods are limited by insufficient representational ability, specific modality number…

Image and Video Processing · Electrical Eng. & Systems 2023-07-06 Jun Shi , Hongyu Kan , Shulan Ruan , Ziqi Zhu , Minfan Zhao , Liang Qiao , Zhaohui Wang , Hong An , Xudong Xue

Transformers have recently gained increasing attention in computer vision. However, existing studies mostly use Transformers for feature representation learning, e.g. for image classification and dense predictions, and the generalizability…

Computer Vision and Pattern Recognition · Computer Science 2021-12-08 Shengcai Liao , Ling Shao

Nutrition estimation is crucial for effective dietary management and overall health and well-being. Existing methods often struggle with sub-optimal accuracy and can be time-consuming. In this paper, we propose NuNet, a transformer-based…

Computer Vision and Pattern Recognition · Computer Science 2024-06-05 Zhengyi Kwan , Wei Zhang , Zhengkui Wang , Aik Beng Ng , Simon See

Low-light image enhancement (LLIE) is a fundamental yet challenging task due to the presence of noise, loss of detail, and poor contrast in images captured under insufficient lighting conditions. Recent methods often rely solely on…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Alexandru Brateanu , Raul Balmez , Ciprian Orhei , Codruta Ancuti , Cosmin Ancuti

Detecting fake news has received a lot of attention. Many previous methods concatenate independently encoded unimodal data, ignoring the benefits of integrated multimodal information. Also, the absence of specialized feature extraction for…

Machine Learning · Computer Science 2025-01-24 Eunjee Choi , Jong-Kook Kim

Multi-modal industrial anomaly detection typically relies on separate models for each product category, fundamentally limiting practical scalability. When shifting to a unified paradigm that handles diverse classes simultaneously, detection…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Yangchen Wu , Huiqiang Xie

This paper presents Universal Vision-Language Dense Retrieval (UniVL-DR), which builds a unified model for multi-modal retrieval. UniVL-DR encodes queries and multi-modality resources in an embedding space for searching candidates from…

Information Retrieval · Computer Science 2023-02-07 Zhenghao Liu , Chenyan Xiong , Yuanhuiyi Lv , Zhiyuan Liu , Ge Yu

Cross-modal retrieval of image-text and video-text is a prominent research area in computer vision and natural language processing. However, there has been insufficient attention given to cross-modal retrieval between human motion and text,…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Sheng Yan , Yang Liu , Haoqiang Wang , Xin Du , Mengyuan Liu , Hong Liu

Image clustering is a crucial but challenging task in multimedia machine learning. Recently the combination of clustering with deep learning has achieved promising performance against conventional methods on high-dimensional image data.…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Ruilin Zhang , Haiyang Zheng , Hongpeng Wang

Recognizing food images presents unique challenges due to the variable spatial layout and shape changes of ingredients with different cooking and cutting methods. This study introduces an advanced approach for recognizing ingredients…

Computer Vision and Pattern Recognition · Computer Science 2024-12-06 Kun Fu , Ying Dai

Increasing demand for high field magnetic resonance (MR) scanner indicates the need for high-quality MR images for accurate medical diagnosis. However, cost constraints, instead, motivate a need for algorithms to enhance images from low…

Computer Vision and Pattern Recognition · Computer Science 2018-06-20 Aditya Sharma , Prabhjot Kaur , Aditya Nigam , Arnav Bhavsar

Complicated image registration is a key issue in medical image analysis, and deep learning-based methods have achieved better results than traditional methods. The methods include ConvNet-based and Transformer-based methods. Although…

Computer Vision and Pattern Recognition · Computer Science 2024-10-29 Runshi Zhang , Hao Mo , Junchen Wang , Bimeng Jie , Yang He , Nenghao Jin , Liang Zhu

Anomaly detection in medical images is challenging due to limited annotations and a domain gap compared to natural images. Existing reconstruction methods often rely on frozen pre-trained encoders, which limits adaptation to domain-specific…

Computer Vision and Pattern Recognition · Computer Science 2025-08-28 Luhu Li , Bowen Lin , Mukhtiar Khan , Shujun Fu

Unsupervised learning visible-infrared person re-identification (USL-VI-ReID) aims to learn modality-invariant features from unlabeled cross-modality datasets and reduce the inter-modality gap. However, the existing methods lack…

Computer Vision and Pattern Recognition · Computer Science 2024-12-30 Ruixing Wu , Yiming Yang , Jiakai He , Haifeng Hu

While deep learning methods have shown great success in medical image analysis, they require a number of medical images to train. Due to data privacy concerns and unavailability of medical annotators, it is oftentimes very difficult to…

Image and Video Processing · Electrical Eng. & Systems 2020-10-08 Yue Yang , Pengtao Xie

Detection Transformers (DETR) are renowned object detection pipelines, however computationally efficient multiscale detection using DETR is still challenging. In this paper, we propose a Cross-Resolution Encoding-Decoding (CRED) mechanism…

Computer Vision and Pattern Recognition · Computer Science 2024-10-08 Ashish Kumar , Jaesik Park