中文
相关论文

相关论文: Self-Supervised Cross-Modal Text-Image Time Series…

200 篇论文

Text-to-image retrieval (T2I retrieval) remains challenging because cross-modal embeddings often behave as bags of concepts, underrepresenting structured visual relationships such as pose and viewpoint. We proposeVisualize-then-Retrieve…

计算机视觉与模式识别 · 计算机科学 2026-04-28 Di Wu , Yixin Wan , Kai-Wei Chang

This paper introduces a novel spatiotemporal feature representation model designed to address the limitations of traditional methods in multidimensional time series (MTS) analysis. The proposed approach converts MTS into one-dimensional…

机器学习 · 计算机科学 2024-10-10 Xu Yan , Yaoting Jiang , Wenyi Liu , Didi Yi , Jianjun Wei

We address the problem of cross-modal information retrieval in the domain of remote sensing. In particular, we are interested in two application scenarios: i) cross-modal retrieval between panchromatic (PAN) and multi-spectral imagery, and…

图像与视频处理 · 电气工程与系统科学 2021-04-22 Ushasi Chaudhuri , Biplab Banerjee , Avik Bhattacharya , Mihai Datcu

The increasing prevalence of video clips has sparked growing interest in text-video retrieval. Recent advances focus on establishing a joint embedding space for text and video, relying on consistent embedding representations to compute…

计算机视觉与模式识别 · 计算机科学 2024-03-28 Jiamian Wang , Guohao Sun , Pichao Wang , Dongfang Liu , Sohail Dianat , Majid Rabbani , Raghuveer Rao , Zhiqiang Tao

Finding target persons in full scene images with a query of text description has important practical applications in intelligent video surveillance.However, different from the real-world scenarios where the bounding boxes are not available,…

计算机视觉与模式识别 · 计算机科学 2024-02-27 Shizhou Zhang , De Cheng , Wenlong Luo , Yinghui Xing , Duo Long , Hao Li , Kai Niu , Guoqiang Liang , Yanning Zhang

Driven by large scale datasets and LLM based architectures, automatic speech recognition (ASR) systems have achieved remarkable improvements in accuracy. However, challenges persist for domain-specific terminology, and short utterances…

音频与语音处理 · 电气工程与系统科学 2025-09-30 Jinming Chen , Lu Wang , Zheshu Song , Wei Deng

Diffusion models have shown great potential in generating realistic image detail. However, adapting these models to video super-resolution (VSR) remains challenging due to their inherent stochasticity and lack of temporal modeling. Previous…

计算机视觉与模式识别 · 计算机科学 2025-08-05 Yong Liu , Jinshan Pan , Yinchuan Li , Qingji Dong , Chao Zhu , Yu Guo , Fei Wang

Recently, the diffusion-based generative paradigm has achieved impressive general image generation capabilities with text prompts due to its accurate distribution modeling and stable training process. However, generating diverse remote…

图像与视频处理 · 电气工程与系统科学 2024-10-31 Jialin Luo , Yuanzhi Wang , Ziqi Gu , Yide Qiu , Shuaizhen Yao , Fuyun Wang , Chunyan Xu , Wenhua Zhang , Dan Wang , Zhen Cui

In this paper, we investigate the cross-media retrieval between images and text, i.e., using image to search text (I2T) and using text to search images (T2I). Existing cross-media retrieval methods usually learn one couple of projections,…

计算机视觉与模式识别 · 计算机科学 2015-06-24 Yunchao Wei , Yao Zhao , Zhenfeng Zhu , Shikui Wei , Yanhui Xiao , Jiashi Feng , Shuicheng Yan

The abundance of multimodal data (e.g. social media posts) has inspired interest in cross-modal retrieval methods. Popular approaches rely on a variety of metric learning losses, which prescribe what the proximity of image and text should…

计算机视觉与模式识别 · 计算机科学 2020-09-24 Christopher Thomas , Adriana Kovashka

Unsupervised human motion segmentation (HMS) can be effectively achieved using subspace clustering techniques. However, traditional methods overlook the role of temporal semantic exploration in HMS. This paper explores the use of temporal…

机器学习 · 计算机科学 2025-12-30 Zheng Xing , Weibing Zhao

Image-text matching is a key multimodal task that aims to model the semantic association between images and text as a matching relationship. With the advent of the multimedia information age, image, and text data show explosive growth, and…

机器学习 · 计算机科学 2024-06-24 Jinyin Wang , Haijing Zhang , Yihao Zhong , Yingbin Liang , Rongwei Ji , Yiru Cang

Text-driven infrared and visible image fusion has gained attention for enabling natural language to guide the fusion process. However, existing methods lack a goal-aligned task to supervise and evaluate how effectively the input text…

计算机视觉与模式识别 · 计算机科学 2026-01-21 Siju Ma , Changsiyu Gong , Xiaofeng Fan , Yong Ma , Chengjie Jiang

Zero-shot Composed Image Retrieval (ZS-CIR) aims to retrieve a target image given a reference image and a relative text, without relying on costly triplet annotations. Existing CLIP-based methods face two core challenges: (1) union-based…

计算机视觉与模式识别 · 计算机科学 2025-10-01 Yuqi Xiao , Yingying Zhu

We propose a visual-linguistic representation learning approach within a self-supervised learning framework by introducing a new operation, loss, and data augmentation strategy. First, we generate diverse features for the image-text…

计算机视觉与模式识别 · 计算机科学 2023-04-04 Jaeyoo Park , Bohyung Han

Scene text recognition has witnessed rapid development with the advance of convolutional neural networks. Nonetheless, most of the previous methods may not work well in recognizing text with low resolution which is often seen in natural…

计算机视觉与模式识别 · 计算机科学 2019-10-22 Wenjia Wang , Enze Xie , Peize Sun , Wenhai Wang , Lixun Tian , Chunhua Shen , Ping Luo

Given a language expression, referring remote sensing image segmentation (RRSIS) aims to identify ground objects and assign pixel-wise labels within the imagery. The one of key challenges for this task is to capture discriminative…

计算机视觉与模式识别 · 计算机科学 2024-12-30 Sen Lei , Xinyu Xiao , Tianlin Zhang , Heng-Chao Li , Zhenwei Shi , Qing Zhu

Multimodal fusion of remote sensing images serves as a core technology for overcoming the limitations of single-source data and improving the accuracy of surface information extraction, which exhibits significant application value in fields…

计算机视觉与模式识别 · 计算机科学 2026-01-12 Siyu Zhang , Lianlei Shan , Runhe Qiu

Skeleton-based Temporal Action Segmentation (STAS) aims to segment and recognize various actions from long, untrimmed sequences of human skeletal movements. Current STAS methods typically employ spatio-temporal modeling to establish…

计算机视觉与模式识别 · 计算机科学 2025-12-15 Haoyu Ji , Bowen Chen , Weihong Ren , Wenze Huang , Zhihao Yang , Zhiyong Wang , Honghai Liu

Audio carries richer information than text, including emotion, speaker traits, and environmental context, while also enabling lower-latency processing compared to speech-to-text pipelines. However, recent multimodal information retrieval…

声音 · 计算机科学 2026-04-23 Tong Zhao , Chenghao Zhang , Yutao Zhu , Zhicheng Dou