中文
相关论文

相关论文: REMM:Rotation-Equivariant Framework for End-to-End…

200 篇论文

Multi-modal Large Language Models (MLLMs) have recently exhibited impressive general-purpose capabilities by leveraging vision foundation models to encode the core concepts of images into representations. These are then combined with…

计算机视觉与模式识别 · 计算机科学 2025-06-12 Sara Ghazanfari , Alexandre Araujo , Prashanth Krishnamurthy , Siddharth Garg , Farshad Khorrami

Renovating the memories in old photos is an intriguing research topic in computer vision fields. These legacy images often suffer from severe and commingled degradations such as cracks, noise, and color-fading, while lack of large-scale…

图像与视频处理 · 电气工程与系统科学 2022-05-12 Runsheng Xu , Zhengzhong Tu , Yuanqi Du , Xiaoyu Dong , Jinlong Li , Zibo Meng , Jiaqi Ma , Hongkai Yu

Referring Expression Comprehension (REC) is a popular multimodal task that aims to accurately detect target objects within a single image based on a given textual expression. However, due to the limitations of earlier models, traditional…

机器学习 · 计算机科学 2025-08-21 Guanghao Jin , Jingpei Wu , Tianpei Guo , Yiyi Niu , Weidong Zhou , Guoyang Liu

The use of local detectors and descriptors in typical computer vision pipelines work well until variations in viewpoint and appearance change become extreme. Past research in this area has typically focused on one of two approaches to this…

计算机视觉与模式识别 · 计算机科学 2022-03-25 Udit Singh Parihar , Aniket Gujarathi , Kinal Mehta , Satyajit Tourani , Sourav Garg , Michael Milford , K. Madhava Krishna

We introduce CommerceMM - a multimodal model capable of providing a diverse and granular understanding of commerce topics associated to the given piece of content (image, text, image+text), and having the capability to generalize to a wide…

计算机视觉与模式识别 · 计算机科学 2022-02-16 Licheng Yu , Jun Chen , Animesh Sinha , Mengjiao MJ Wang , Hugo Chen , Tamara L. Berg , Ning Zhang

One of the challenges of the Optical Music Recognition task is to transcript the symbols of the camera-captured images into digital music notations. Previous end-to-end model which was developed as a Convolutional Recurrent Neural Network…

计算机视觉与模式识别 · 计算机科学 2021-08-05 Aozhi Liu , Lipei Zhang , Yaqi Mei , Baoqiang Han , Zifeng Cai , Zhaohua Zhu , Jing Xiao

Due to the high similarity of disparity between consecutive frames in video sequences, the area where disparity changes is defined as the residual map, which can be calculated. Based on this, we propose RecSM, a network based on residual…

计算机视觉与模式识别 · 计算机科学 2024-06-06 Youchen Zhao , Guorong Luo , Hua Zhong , Haixiong Li

Cross-spectral person re-identification, which aims to associate identities to pedestrians across different spectra, faces a main challenge of the modality discrepancy. In this paper, we address the problem from both image-level and…

计算机视觉与模式识别 · 计算机科学 2023-02-03 Lei Tan , Yukang Zhang , Shengmei Shen , Yan Wang , Pingyang Dai , Xianming Lin , Yongjian Wu , Rongrong Ji

Equivariant and invariant deep learning models have been developed to exploit intrinsic symmetries in data, demonstrating significant effectiveness in certain scenarios. However, these methods often suffer from limited representation…

计算机视觉与模式识别 · 计算机科学 2025-05-27 Yulu Bai , Jiahong Fu , Qi Xie , Deyu Meng

Cross-modal retrieval across image and text modalities is a challenging task due to its inherent ambiguity: An image often exhibits various situations, and a caption can be coupled with diverse images. Set-based embedding has been studied…

计算机视觉与模式识别 · 计算机科学 2023-07-25 Dongwon Kim , Namyup Kim , Suha Kwak

Referring expression grounding is an important and challenging task in computer vision. To avoid the laborious annotation in conventional referring grounding, unpaired referring grounding is introduced, where the training data only contains…

计算机视觉与模式识别 · 计算机科学 2022-06-07 Hengcan Shi , Munawar Hayat , Jianfei Cai

Finding correspondences between images or 3D scans is at the heart of many computer vision and image retrieval applications and is often enabled by matching local keypoint descriptors. Various learning approaches have been applied in the…

计算机视觉与模式识别 · 计算机科学 2018-05-10 Georgios Georgakis , Srikrishna Karanam , Ziyan Wu , Jan Ernst , Jana Kosecka

Feature matching in omnidirectional vision systems is a challenging problem, mainly because complicated optical systems make the theoretical modelling of invariance and construction of invariant feature descriptors hard or even impossible.…

计算机视觉与模式识别 · 计算机科学 2011-12-30 Jonathan Masci , Davide Migliore , Michael M. Bronstein , Jürgen Schmidhuber

Cross-modal retrieval is gaining increasing efficacy and interest from the research community, thanks to large-scale training, novel architectural and learning designs, and its application in LLMs and multimodal LLMs. In this paper, we move…

计算机视觉与模式识别 · 计算机科学 2025-03-05 Davide Caffagni , Sara Sarto , Marcella Cornia , Lorenzo Baraldi , Rita Cucchiara

We have witnessed the discovery of many techniques for network representation learning in recent years, ranging from encoding the context in random walks to embedding the lower order connections, to finding latent space representations with…

机器学习 · 计算机科学 2017-10-10 Yanlei Yu , Zhiwu Lu , Jiajun Liu , Guoping Zhao , Ji-Rong Wen , Kai Zheng

Reranking is a critical component in many information retrieval pipelines. Despite remarkable progress in text-only settings, multimodal reranking remains challenging, particularly when the candidate set contains hybrid text and image…

信息检索 · 计算机科学 2026-05-26 Yupei Yang , Lin Yang , Wanxi Deng , Lin Qu , Shikui Tu , Lei Xu

Incorporating equivariance as an inductive bias into deep learning architectures to take advantage of the data symmetry has been successful in multiple applications, such as chemistry and dynamical systems. In particular, roto-translations…

机器学习 · 计算机科学 2026-01-06 Ahmed A. Elhag , T. Konstantin Rusch , Francesco Di Giovanni , Michael Bronstein

Unified Multimodal Models (UMMs) integrate multimodal understanding and generation, yet they are limited to maintaining visual consistency and disambiguating visual cues when referencing details across multiple input images. In this work,…

计算机视觉与模式识别 · 计算机科学 2026-01-12 Pengcheng Xu , Peng Tang , Donghao Luo , Xiaobin Hu , Weichu Cui , Qingdong He , Zhennan Chen , Jiangning Zhang , Charles Ling , Boyu Wang

The rapid advancement of Multimodal Large Language Models (MLLMs) has extended CLIP-based frameworks to produce powerful, universal embeddings for retrieval tasks. However, existing methods primarily focus on natural images, offering…

计算机视觉与模式识别 · 计算机科学 2025-11-03 Weijian Jian , Yajun Zhang , Dawei Liang , Chunyu Xie , Yixiao He , Dawei Leng , Yuhui Yin

Rotation-invariance is a desired property of machine-learning models for medical image analysis and in particular for computational pathology applications. We propose a framework to encode the geometric structure of the special Euclidean…

计算机视觉与模式识别 · 计算机科学 2020-02-21 Maxime W. Lafarge , Erik J. Bekkers , Josien P. W. Pluim , Remco Duits , Mitko Veta