中文
相关论文

相关论文: Language Guided Local Infiltration for Interactive…

200 篇论文

Despite remarkable advancements in text-to-image person re-identification (TIReID) facilitated by the breakthrough of cross-modal embedding models, existing methods often struggle to distinguish challenging candidate images due to intrinsic…

机器学习 · 计算机科学 2025-06-16 Yang Qin , Chao Chen , Zhihang Fu , Dezhong Peng , Xi Peng , Peng Hu

Many image restoration (IR) tasks require both pixel-level fidelity and high-level semantic understanding to recover realistic photos with fine-grained details. However, previous approaches often struggle to effectively leverage both the…

计算机视觉与模式识别 · 计算机科学 2025-12-22 Cuixin Yang , Rongkang Dong , Kin-Man Lam

Image search stands as a pivotal task in multimedia and computer vision, finding applications across diverse domains, ranging from internet search to medical diagnostics. Conventional image search systems operate by accepting textual or…

多媒体 · 计算机科学 2024-04-30 Hongyi Zhu , Jia-Hong Huang , Stevan Rudinac , Evangelos Kanoulas

We present a novel framework, Localized Image Stylization with Audio (LISA) which performs audio-driven localized image stylization. Sound often provides information about the specific context of the scene and is closely related to a…

计算机视觉与模式识别 · 计算机科学 2022-11-22 Seung Hyun Lee , Chanyoung Kim , Wonmin Byeon , Sang Ho Yoon , Jinkyu Kim , Sangpil Kim

Recently, AutoRegressive (AR) models for the whole image generation empowered by transformers have achieved comparable or even better performance to Generative Adversarial Networks (GANs). Unfortunately, directly applying such AR models to…

计算机视觉与模式识别 · 计算机科学 2021-10-19 Chenjie Cao , Yuxin Hong , Xiang Li , Chengrong Wang , Chengming Xu , XiangYang Xue , Yanwei Fu

In this paper, we primarily address the issue of dialogue-form context query within the interactive text-to-image retrieval task. Our methodology, PlugIR, actively utilizes the general instruction-following capability of LLMs in two ways.…

计算机视觉与模式识别 · 计算机科学 2024-07-26 Saehyung Lee , Sangwon Yu , Junsung Park , Jihun Yi , Sungroh Yoon

Composed Image Retrieval (CIR) aims to retrieve target images from candidate set using a hybrid-modality query consisting of a reference image and a relative caption that describes the user intent. Recent studies attempt to utilize…

信息检索 · 计算机科学 2024-12-17 Zelong Sun , Dong Jing , Guoxing Yang , Nanyi Fei , Zhiwu Lu

This paper addresses the task of interactive, conversational text-to-image retrieval. Our DIR-TIR framework progressively refines the target image search through two specialized modules: the Dialog Refiner Module and the Image Refiner…

计算机视觉与模式识别 · 计算机科学 2025-11-19 Zongwei Zhen , Biqing Zeng

Most existing low-light image enhancement (LLIE) methods rely on pre-trained model priors, low-light inputs, or both, while neglecting the semantic guidance available from normal-light images. This limitation hinders their effectiveness in…

计算机视觉与模式识别 · 计算机科学 2026-05-04 Xiaoran Sun , Liyan Wang , Yeying Jin , Kin-man Lam , Zhixun Su , Yang Yang , Jinshan Pan , Cong Wang

Image-Text Retrieval (ITR) finds broad applications in healthcare, aiding clinicians and radiologists by automatically retrieving relevant patient cases in the database given the query image and/or report, for more efficient clinical…

计算机视觉与模式识别 · 计算机科学 2025-03-11 Meng Zheng , Jiajin Zhang , Benjamin Planche , Zhongpai Gao , Terrence Chen , Ziyan Wu

The goal of Text-to-Image Person Retrieval (TIPR) is to retrieve specific person images according to the given textual descriptions. A primary challenge in this task is bridging the substantial representational gap between visual and…

计算与语言 · 计算机科学 2025-01-20 Delong Liu , Haiwen Li , Zhicheng Zhao , Yuan Dong

Image retrieval with hybrid-modality queries, also known as composing text and image for image retrieval (CTI-IR), is a retrieval task where the search intention is expressed in a more complex query format, involving both vision and text…

计算机视觉与模式识别 · 计算机科学 2022-04-26 Yida Zhao , Yuqing Song , Qin Jin

Current multimodal information retrieval studies mainly focus on single-image inputs, which limits real-world applications involving multiple images and text-image interleaved content. In this work, we introduce the text-image interleaved…

计算与语言 · 计算机科学 2025-02-19 Xin Zhang , Ziqi Dai , Yongqi Li , Yanzhao Zhang , Dingkun Long , Pengjun Xie , Meishan Zhang , Jun Yu , Wenjie Li , Min Zhang

With the rapid rise of Artificial Intelligence Generated Content (AIGC), image manipulation has become increasingly accessible, posing significant challenges for image forgery detection and localization (IFDL). In this paper, we study how…

计算机视觉与模式识别 · 计算机科学 2026-03-16 Shaofeng Guo , Jiequan Cui , Richang Hong

Most existing Low-light Image Enhancement (LLIE) methods either directly map Low-Light (LL) to Normal-Light (NL) images or use semantic or illumination maps as guides. However, the ill-posed nature of LLIE and the difficulty of semantic…

计算机视觉与模式识别 · 计算机科学 2024-07-18 Han Zhou , Wei Dong , Xiaohong Liu , Shuaicheng Liu , Xiongkuo Min , Guangtao Zhai , Jun Chen

Large-scale Text-to-Image (T2I) diffusion models demonstrate significant generation capabilities based on textual prompts. Based on the T2I diffusion models, text-guided image editing research aims to empower users to manipulate generated…

计算机视觉与模式识别 · 计算机科学 2024-05-03 Chuanming Tang , Kai Wang , Fei Yang , Joost van de Weijer

Deceptive images can be shared in seconds with social networking services, posing substantial risks. Tampering traces, such as boundary artifacts and high-frequency information, have been significantly emphasized by massive networks in the…

计算机视觉与模式识别 · 计算机科学 2024-01-02 Xuntao Liu , Yuzhou Yang , Qichao Ying , Zhenxing Qian , Xinpeng Zhang , Sheng Li

Text-to-image person retrieval (TIPR) aims to identify the target person using textual descriptions, facing challenge in modality heterogeneity. Prior works have attempted to address it by developing cross-modal global or local alignment…

计算机视觉与模式识别 · 计算机科学 2025-10-21 Min Cao , Xinyu Zhou , Ding Jiang , Bo Du , Mang Ye , Min Zhang

Composed Image Retrieval (CIR) is a complex task that aims to retrieve images based on a multimodal query. Typical training data consists of triplets containing a reference image, a textual description of desired modifications, and the…

计算机视觉与模式识别 · 计算机科学 2025-03-26 Chuong Huynh , Jinyu Yang , Ashish Tawari , Mubarak Shah , Son Tran , Raffay Hamid , Trishul Chilimbi , Abhinav Shrivastava

Interactive Text-to-image retrieval (I-TIR) is an important enabler for a wide range of state-of-the-art services in domains such as e-commerce and education. However, current methods rely on finetuned Multimodal Large Language Models…

信息检索 · 计算机科学 2025-07-11 Zijun Long , Kangheng Liang , Gerardo Aragon-Camarasa , Richard Mccreadie , Paul Henderson
‹ 上一页 1 2 3 10 下一页 ›