中文
相关论文

相关论文: VOSR: A Vision-Only Generative Model for Image Sup…

200 篇论文

Diffusion-based models have been widely used in various visual generation tasks, showing promising results in image super-resolution (SR), while typically being limited by dozens or even hundreds of sampling steps. Although existing methods…

计算机视觉与模式识别 · 计算机科学 2025-06-04 Xue Wu , Jingwei Xin , Zhijun Tu , Jie Hu , Jie Li , Nannan Wang , Xinbo Gao

Recent advancements in large language models have demonstrated how chain-of-thought (CoT) and reinforcement learning (RL) can improve performance. However, applying such reasoning strategies to the visual generation domain remains largely…

计算机视觉与模式识别 · 计算机科学 2025-07-02 Dongzhi Jiang , Ziyu Guo , Renrui Zhang , Zhuofan Zong , Hao Li , Le Zhuo , Shilin Yan , Pheng-Ann Heng , Hongsheng Li

This study presents a new image super-resolution (SR) technique based on diffusion inversion, aiming at harnessing the rich image priors encapsulated in large pre-trained diffusion models to improve SR performance. We design a Partial noise…

计算机视觉与模式识别 · 计算机科学 2025-03-14 Zongsheng Yue , Kang Liao , Chen Change Loy

fMRI (functional Magnetic Resonance Imaging) visual decoding involves decoding the original image from brain signals elicited by visual stimuli. This often relies on manually labeled ROIs (Regions of Interest) to select brain voxels.…

计算机视觉与模式识别 · 计算机科学 2025-03-12 Ziyu Wang , Tengyu Pan , Zhenyu Li , Ji Wu , Xiuxing Li , Jianyong Wang

Infrared imaging is essential for autonomous driving and robotic operations as a supportive modality due to its reliable performance in challenging environments. Despite its popularity, the limitations of infrared cameras, such as low…

计算机视觉与模式识别 · 计算机科学 2025-03-04 Xingyuan Li , Zirui Wang , Yang Zou , Zhixin Chen , Jun Ma , Zhiying Jiang , Long Ma , Jinyuan Liu

Visual Document Retrieval (VDR) typically operates as text-to-image retrieval using specialized bi-encoders trained to directly embed document images. We revisit a zero-shot generate-and-encode pipeline: a vision-language model first…

信息检索 · 计算机科学 2025-09-22 Thong Nguyen , Yibin Lei , Jia-Huei Ju , Andrew Yates

Audio-visual speech recognition (AVSR) has gained remarkable success for ameliorating the noise-robustness of speech recognition. Mainstream methods focus on fusing audio and visual inputs to obtain modality-invariant representations.…

声音 · 计算机科学 2023-02-03 Chen Chen , Yuchen Hu , Qiang Zhang , Heqing Zou , Beier Zhu , Eng Siong Chng

Single-image-to-3D generative models can now produce high-quality geometry, yet conditioning on a single view inevitably introduces ambiguity about unseen regions. Multi-view conditioning can reduce this ambiguity, but existing methods…

计算机视觉与模式识别 · 计算机科学 2026-05-21 Hanxiao Sun , Mingxin Yang , Shuhui Yang , Zebin He , Xintong Han , Hongbo Fu , Chunchao Guo , Wenhan Luo

Recently, it has been demonstrated that deep neural networks can significantly improve the performance of single image super-resolution (SISR). Numerous studies have concentrated on raising the quantitative quality of super-resolved (SR)…

计算机视觉与模式识别 · 计算机科学 2020-09-14 Zheng Hui , Jie Li , Xinbo Gao , Xiumei Wang

Implicit Neural Representations (INR) have been successfully employed for Arbitrary-scale Super-Resolution (ASR). However, INR-based models need to query the multi-layer perceptron module numerous times and render a pixel in each query,…

图像与视频处理 · 电气工程与系统科学 2025-07-31 Du Chen , Liyi Chen , Zhengqiang Zhang , Lei Zhang

To replicate the success of text-to-image (T2I) generation, recent works employ large-scale video datasets to train a text-to-video (T2V) generator. Despite their promising results, such paradigm is computationally expensive. In this work,…

计算机视觉与模式识别 · 计算机科学 2023-03-20 Jay Zhangjie Wu , Yixiao Ge , Xintao Wang , Weixian Lei , Yuchao Gu , Yufei Shi , Wynne Hsu , Ying Shan , Xiaohu Qie , Mike Zheng Shou

Advances in unsupervised learning of object-representations have culminated in the development of a broad range of methods for unsupervised object segmentation and interpretable object-centric scene generation. These methods, however, are…

计算机视觉与模式识别 · 计算机科学 2022-01-26 Martin Engelcke , Oiwi Parker Jones , Ingmar Posner

This paper studies the problem of real-world video super-resolution (VSR) for animation videos, and reveals three key improvements for practical animation VSR. First, recent real-world super-resolution approaches typically rely on…

计算机视觉与模式识别 · 计算机科学 2023-01-18 Yanze Wu , Xintao Wang , Gen Li , Ying Shan

In most interactive image generation tasks, given regions of interest (ROI) by users, the generated results are expected to have adequate diversities in appearance while maintaining correct and reasonable structures in original images. Such…

计算机视觉与模式识别 · 计算机科学 2024-10-28 Jinshu Chen , Qihui Xu , Qi Kang , MengChu Zhou

We introduce a multi-scale Image Super Resolution (ISR) method building on recent advances in Visual Auto-Regressive (VAR) modeling. VAR models break image tokenization into additive, gradually increasing scales, using Residual Quantization…

计算机视觉与模式识别 · 计算机科学 2026-05-15 Isma Hadji , Enrique Sanchez , Adrian Bulat , Brais Martinez , Georgios Tzimiropoulos

Typical large vision-language models (LVLMs) apply autoregressive supervision solely to textual sequences, without fully incorporating the visual modality into the learning process. This results in three key limitations: (1) an inability to…

计算机视觉与模式识别 · 计算机科学 2026-01-06 Dianyi Wang , Wei Song , Yikun Wang , Siyuan Wang , Kaicheng Yu , Zhongyu Wei , Jiaqi Wang

A light-weight super-resolution (LSR) method from a single image targeting mobile applications is proposed in this work. LSR predicts the residual image between the interpolated low-resolution (ILR) and high-resolution (HR) images using a…

图像与视频处理 · 电气工程与系统科学 2023-02-28 Wei Wang , Xuejing Lei , Yueru Chen , Ming-Sui Lee , C. -C. Jay Kuo

This paper explores sentence-level multilingual Visual Speech Recognition (VSR) that can recognize different languages with a single trained model. As the massive multilingual modeling of visual data requires huge computational costs, we…

音频与语音处理 · 电气工程与系统科学 2024-07-19 Minsu Kim , Jeong Hun Yeo , Se Jin Park , Hyeongseop Rha , Yong Man Ro

Despite their wide-spread success, Text-to-Image models (T2I) still struggle to produce images that are both aesthetically pleasing and faithful to the user's input text. We introduce DreamSync, a model-agnostic training algorithm by design…

计算机视觉与模式识别 · 计算机科学 2023-12-01 Jiao Sun , Deqing Fu , Yushi Hu , Su Wang , Royi Rassin , Da-Cheng Juan , Dana Alon , Charles Herrmann , Sjoerd van Steenkiste , Ranjay Krishna , Cyrus Rashtchian

Visual servoing enables robotic systems to perform accurate closed-loop control, which is required in many applications. However, existing methods either require precise calibration of the robot kinematic model and cameras or use neural…