中文
相关论文

相关论文: Self-Supervised Cross-Modal Text-Image Time Series…

200 篇论文

Diffusion-based methods, endowed with a formidable generative prior, have received increasing attention in Image Super-Resolution (ISR) recently. However, as low-resolution (LR) images often undergo severe degradation, it is challenging for…

计算机视觉与模式识别 · 计算机科学 2024-07-22 Yunpeng Qu , Kun Yuan , Kai Zhao , Qizhi Xie , Jinhua Hao , Ming Sun , Chao Zhou

Developers need to perform adequate testing to ensure the quality of Automatic Speech Recognition (ASR) systems. However, manually collecting required test cases is tedious and time-consuming. Our recent work proposes CrossASR, a…

软件工程 · 计算机科学 2022-01-06 Muhammad Hilmi Asyrofi , Zhou Yang , David Lo

IRSTD (InfraRed Small Target Detection) detects small targets in infrared blurry backgrounds and is essential for various applications. The detection task is challenging due to the small size of the targets and their sparse distribution in…

计算机视觉与模式识别 · 计算机科学 2025-07-18 Pranav Singh , Pravendra Singh

Traditional cross-modal retrieval assumes explicit association of concepts across modalities, where there is no ambiguity in how the concepts are linked to each other, e.g., when we do the image search with a query "dogs", we expect to see…

计算机视觉与模式识别 · 计算机科学 2018-04-26 Yale Song , Mohammad Soleymani

Guided image super-resolution (GISR) aims to obtain a high-resolution (HR) target image by enhancing the spatial resolution of a low-resolution (LR) target image under the guidance of a HR image. However, previous model-based methods mainly…

图像与视频处理 · 电气工程与系统科学 2022-03-11 Man Zhou , Keyu Yan , Jinshan Pan , Wenqi Ren , Qi Xie , Xiangyong Cao

Image Super-Resolution (SR) provides a promising technique to enhance the image quality of low-resolution optical sensors, facilitating better-performing target detection and autonomous navigation in a wide range of robotics applications.…

计算机视觉与模式识别 · 计算机科学 2020-12-08 Fan Wang , Jiangxin Yang , Yanlong Cao , Yanpeng Cao , Michael Ying Yang

Visual-semantic embedding aims to find a shared latent space where related visual and textual instances are close to each other. Most current methods learn injective embedding functions that map an instance to a single point in the shared…

计算机视觉与模式识别 · 计算机科学 2019-07-18 Yale Song , Mohammad Soleymani

Referring Image Segmentation (RIS) is a fundamental vision-language task that outputs object masks based on text descriptions. Many works have achieved considerable progress for RIS, including different fusion method designs. In this work,…

计算机视觉与模式识别 · 计算机科学 2023-07-25 Jianzong Wu , Xiangtai Li , Xia Li , Henghui Ding , Yunhai Tong , Dacheng Tao

Medical vision-language pre-training methods mainly leverage the correspondence between paired medical images and radiological reports. Although multi-view spatial images and temporal sequences of image-report pairs are available in…

人工智能 · 计算机科学 2024-05-31 Jinxia Yang , Bing Su , Wayne Xin Zhao , Ji-Rong Wen

As posts on social media increase rapidly, analyzing the sentiments embedded in image-text pairs has become a popular research topic in recent years. Although existing works achieve impressive accomplishments in simultaneously harnessing…

计算与语言 · 计算机科学 2025-12-04 Daiqing Wu , Dongbao Yang , Yu Zhou , Can Ma

Diffusion models achieve remarkable success in processing images and text, and have been extended to special domains such as time series forecasting (TSF). Existing diffusion-based approaches for TSF primarily focus on modeling…

计算与语言 · 计算机科学 2025-04-29 Chen Su , Yuanhe Tian , Yan Song

AI-generated videos (AIGVs) have achieved unprecedented photorealism, posing severe threats to digital forensics. Existing AIGV detectors focus mainly on localized artifacts or short-term temporal inconsistencies, thus often fail to capture…

计算机视觉与模式识别 · 计算机科学 2026-04-07 Hang Wang , Chao Shen , Lei Zhang , Zhi-Qi Cheng

Zero-Shot Composed Image Retrieval (ZS-CIR) aims to retrieve target images given a multimodal query (comprising a reference image and a modification text), without training on annotated triplets. Existing methods typically convert the…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Tianyue Wang , Leigang Qu , Tianyu Yang , Xiangzhao Hao , Yifan Xu , Haiyun Guo , Jinqiao Wang

With the rapid progression of deep learning technologies, multi-modality image fusion has become increasingly prevalent in object detection tasks. Despite its popularity, the inherent disparities in how different sources depict scene…

计算机视觉与模式识别 · 计算机科学 2024-01-02 Xingyuan Li , Yang Zou , Jinyuan Liu , Zhiying Jiang , Long Ma , Xin Fan , Risheng Liu

Referring Remote Sensing Image Segmentation (RRSIS) is critical for ecological monitoring, urban planning, and disaster management, requiring precise segmentation of objects in remote sensing imagery guided by textual descriptions. This…

计算机视觉与模式识别 · 计算机科学 2025-02-13 Tianxiang Zhang , Zhaokun Wen , Bo Kong , Kecheng Liu , Yisi Zhang , Peixian Zhuang , Jiangyun Li

Visible-infrared cross-modality person re-identification is a challenging ReID task, which aims to retrieve and match the same identity's images between the heterogeneous visible and infrared modalities. Thus, the core of this task is to…

计算机视觉与模式识别 · 计算机科学 2023-05-24 Tengfei Liang , Yi Jin , Yajun Gao , Wu Liu , Songhe Feng , Tao Wang , Yidong Li

Continuous space-time video super-resolution (C-STVSR) aims to simultaneously enhance video resolution and frame rate at an arbitrary scale. Recently, implicit neural representation (INR) has been applied to video restoration, representing…

计算机视觉与模式识别 · 计算机科学 2025-10-02 Yunfan Lu , Yusheng Wang , Zipeng Wang , Pengteng Li , Bin Yang , Hui Xiong

Existing methods for Table Structure Recognition (TSR) from camera-captured or scanned documents perform poorly on complex tables consisting of nested rows / columns, multi-line texts and missing cell data. This is because current…

计算机视觉与模式识别 · 计算机科学 2022-03-15 Arushi Jain , Shubham Paliwal , Monika Sharma , Lovekesh Vig

Text-to-Image Retrieval (T2IR) is a highly valuable task that aims to match a given textual query to images in a gallery. Existing benchmarks primarily focus on textual queries describing overall image semantics or foreground salient…

计算机视觉与模式识别 · 计算机科学 2025-06-02 Chunxu Liu , Chi Xie , Xiaxu Chen , Wei Li , Feng Zhu , Rui Zhao , Limin Wang

When a very fast dynamic event is recorded with a low-framerate camera, the resulting video suffers from severe motion blur (due to exposure time) and motion aliasing (due to low sampling rate in time). True Temporal Super-Resolution (TSR)…

计算机视觉与模式识别 · 计算机科学 2020-10-16 Liad Pollak Zuckerman , Eyal Naor , George Pisha , Shai Bagon , Michal Irani
‹ 上一页 1 8 9 10 下一页 ›