English
Related papers

Related papers: PaAgent: Portrait-Aware Image Restoration Agent vi…

200 papers

Large Audio Language Models (LALMs) excel at perception but struggle with complex reasoning requiring precise acoustic measurements. While external tools can extract fine-grained features like exact tempo or pitch, effective integration…

Sound · Computer Science 2026-02-17 Siqian Tong , Xuan Li , Yiwei Wang , Baolong Bi , Yujun Cai , Shenghua Liu , Yuchen He , Chengpeng Hao

In dynamic scenes, images often suffer from dynamic blur due to superposition of motions or low signal-noise ratio resulted from quick shutter speed when avoiding motions. Recovering sharp and clean results from the captured images heavily…

Computer Vision and Pattern Recognition · Computer Science 2022-03-24 Cheng Zhang , Shaolin Su , Yu Zhu , Qingsen Yan , Jinqiu Sun , Yanning Zhang

We propose Intra and Inter Parser-Prompted Transformers (PPTformer) that explore useful features from visual foundation models for image restoration. Specifically, PPTformer contains two parts: an Image Restoration Network (IRNet) for…

Computer Vision and Pattern Recognition · Computer Science 2025-03-19 Cong Wang , Jinshan Pan , Liyan Wang , Wei Wang

Many super-resolution (SR) algorithms have been proposed to increase image resolution. However, full-reference (FR) image quality assessment (IQA) metrics for comparing and evaluating different SR algorithms are limited. In this work, we…

Computer Vision and Pattern Recognition · Computer Science 2025-09-09 Yixiao Li , Xiaoyuan Yang , Guanghui Yue , Jun Fu , Qiuping Jiang , Xu Jia , Paul L. Rosin , Hantao Liu , Wei Zhou

Group activities usually involve spatiotemporal dynamics among many interactive individuals, while only a few participants at several key frames essentially define the activity. Therefore, effectively modeling the group-relevant and…

Computer Vision and Pattern Recognition · Computer Science 2020-03-04 Guyue Hu , Bo Cui , Yuan He , Shan Yu

Photoacoustic computed tomography (PACT) is a promising imaging modality that combines the advantages of optical contrast with ultrasound detection. Utilizing ultrasound transducers with larger surface areas can improve detection…

Image retargeting aims at altering an image size while preserving important content and minimizing noticeable distortions. However, previous image retargeting methods create outputs that suffer from artifacts and distortions. Besides, most…

Computer Vision and Pattern Recognition · Computer Science 2022-09-13 Mohammad Reza Naderi , Mohammad Hossein Givkashi , Nader Karimi , Shahram Shirani , Shadrokh Samavi

Diffusion models enable high-quality and diverse visual content synthesis. However, they struggle to generate rare or unseen concepts. To address this challenge, we explore the usage of Retrieval-Augmented Generation (RAG) with image…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Rotem Shalev-Arkushin , Rinon Gal , Amit H. Bermano , Ohad Fried

Multi-label image classification demands adaptive training strategies to navigate complex, evolving visual-semantic landscapes, yet conventional methods rely on static configurations that falter in dynamic settings. We propose MAT-Agent, a…

Computer Vision and Pattern Recognition · Computer Science 2025-10-22 Jusheng Zhang , Kaitong Cai , Yijia Fan , Ningyuan Liu , Keze Wang

In remote sensing imagery analysis, patch-based methods have limitations in capturing information beyond the sliding window. This shortcoming poses a significant challenge in processing complex and variable geo-objects, which results in…

Computer Vision and Pattern Recognition · Computer Science 2023-09-28 Yinhe Liu , Sunan Shi , Junjue Wang , Yanfei Zhong

Embodied agents for creative tasks like photography must bridge the semantic gap between high-level language commands and geometric control. We introduce PhotoAgent, an agent that achieves this by integrating Large Multimodal Models (LMMs)…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Lirong Che , Zhenfeng Gan , Yanbo Chen , Junbo Tan , Xueqian Wang

In-context learning has recently been linked to implicit gradient descent in linear self-attention models, suggesting that context can induce a forward-pass update. Retrieval-augmented generation (RAG) also relies on context, but retrieved…

Computation and Language · Computer Science 2026-05-27 Mingchen Li , Jiatan Huang , Chuxu Zhang , Liang Zhao , Hong Yu

Model fusion is a key strategy for robust recognition in unconstrained scenarios, as different models provide complementary strengths. This is especially important for whole-body human recognition, where biometric cues such as face, gait,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Jie Zhu , Xiao Guo , Yiyang Su , Anil Jain , Xiaoming Liu

IRGAN is an information retrieval (IR) modeling approach that uses a theoretical minimax game between a generative and a discriminative model to iteratively optimize both of them, hence unifying the generative and discriminative approaches.…

Information Retrieval · Computer Science 2019-10-02 Moksh Jain , Sowmya Kamath S

Great successes have been achieved using deep learning techniques for image super-resolution (SR) with fixed scales. To increase its real world applicability, numerous models have also been proposed to restore SR images with arbitrary scale…

Image and Video Processing · Electrical Eng. & Systems 2022-09-28 Zhihong Pan , Baopu Li , Dongliang He , Wenhao Wu , Errui Ding

Face aging, which aims at aesthetically rendering a given face to predict its future appearance, has received significant research attention in recent years. Although great progress has been achieved with the success of Generative…

Computer Vision and Pattern Recognition · Computer Science 2019-11-18 Yunfan Liu , Qi Li , Zhenan Sun , Tieniu Tan

High-quality imaging in photoacoustic computed tomography (PACT) usually requires a high-channel count system for dense spatial sampling around the object to avoid aliasing-related artefacts. To reduce system complexity, various image…

Image and Video Processing · Electrical Eng. & Systems 2024-09-24 Bowei Yao , Shilong Cui , Haizhao Dai , Qing Wu , Youshen Xiao , Fei Gao , Jingyi Yu , Yuyao Zhang , Xiran Cai

Nowadays, a huge number of images are available. However, retrieving a required image for an ordinary user is a challenging task in computer vision systems. During the past two decades, many types of research have been introduced to improve…

Multimedia · Computer Science 2020-01-30 Amir Vatani , Milad Taleby Ahvanooey , Mostafa Rahimi

Deep artificial neural networks, trained with labeled data sets are widely used in numerous vision and robotics applications today. In terms of AI, these are called reflex models, referring to the fact that they do not self-evolve or…

Computer Vision and Pattern Recognition · Computer Science 2020-02-20 Hai Xiao , Jin Shang , Mengyuan Huang

Existing free-energy guided No-Reference Image Quality Assessment (NR-IQA) methods still suffer from finding a balance between learning feature information at the pixel level of the image and capturing high-level feature information and the…

Computer Vision and Pattern Recognition · Computer Science 2024-12-30 Zhaoyang Wang , Bo Hu , Mingyang Zhang , Jie Li , Leida Li , Maoguo Gong , Xinbo Gao
‹ Prev 1 3 4 5 6 7 10 Next ›