English
Related papers

Related papers: JarvisIR: Elevating Autonomous Driving Perception …

200 papers

Photo retouching has become integral to contemporary visual storytelling, enabling users to capture aesthetics and express creativity. While professional tools such as Adobe Lightroom offer powerful capabilities, they demand substantial…

Computer Vision and Pattern Recognition · Computer Science 2025-06-24 Yunlong Lin , Zixu Lin , Kunjie Lin , Jinbin Bai , Panwang Pan , Chenxin Li , Haoyu Chen , Zhongdao Wang , Xinghao Ding , Wenbo Li , Shuicheng Yan

Real-world image restoration (IR) is inherently complex and often requires combining multiple specialized models to address diverse degradations. Inspired by human problem-solving, we propose AgenticIR, an agentic system that mimics the…

Computer Vision and Pattern Recognition · Computer Science 2025-02-18 Kaiwen Zhu , Jinjin Gu , Zhiyuan You , Yu Qiao , Chao Dong

Many everyday tasks rely on external tutorials such as manuals and videos, requiring users to constantly switch between reading instructions and performing actions, which disrupts workflow and increases cognitive load. Augmented reality…

Human-Computer Interaction · Computer Science 2026-05-19 Yusi Sun , Ying Jiang , Jiayin Lu , Yin yang , Yong-Hong Kuo , Chenfanfu Jiang

Video restoration in real-world scenarios is challenged by heterogeneous degradations, where static architectures and fixed inference pipelines often fail to generalize. Recent agent-based approaches offer dynamic decision making, yet…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Xuanyu Zhang , Weiqi Li , Qunliang Xing , Jingfen Xie , Bin Chen , Junlin Li , Li Zhang , Jian Zhang , Shijie Zhao

Predictive applications of machine learning often rely on small (sub 1 Bn parameter) specialized models tuned to particular domains or modalities. Such models often achieve excellent performance, but lack flexibility. LLMs and VLMs offer…

Machine Learning · Computer Science 2026-04-30 Benjamin Feuer , Lennart Purucker , Oussama Elachqar , Chinmay Hegde

Reliable visual perception under adverse weather conditions, such as rain, haze, snow, or a mixture of them, is desirable yet challenging for autonomous driving and outdoor robots. In this paper, we propose a unified Memory-Enhanced…

Computer Vision and Pattern Recognition · Computer Science 2025-11-24 Qianyi Shao , Yuanfan Zhang , Renxiang Xiao , Liang Hu

Multimodal Large Language Models (MLLMs) have recently demonstrated impressive capabilities in connecting vision and language, yet their proficiency in fundamental visual reasoning tasks remains limited. This limitation can be attributed to…

Computer Vision and Pattern Recognition · Computer Science 2025-12-19 Davide Caffagni , Sara Sarto , Marcella Cornia , Lorenzo Baraldi , Pier Luigi Dovesi , Shaghayegh Roohi , Mark Granroth-Wilding , Rita Cucchiara

Image restoration is critical for improving the quality of degraded images, which is vital for applications like autonomous driving, security surveillance, and digital content enhancement. However, existing methods are often tailored to…

Computer Vision and Pattern Recognition · Computer Science 2025-04-14 Ziyan Liu , Yuxu Lu , Huashan Yu , Dong yang

Autonomous vehicles rely heavily upon their perception subsystems to see the environment in which they operate. Unfortunately, the effect of variable weather conditions presents a significant challenge to object detection algorithms, and…

Multimodal Large Language Model (MLLM)-driven image restoration agent demonstrates effectiveness in degradation coupling scenarios by flexibly selecting tools and determining removal orders. However, their zero-shot planning often fails…

Computer Vision and Pattern Recognition · Computer Science 2026-05-22 Kailin Zhuang , Jiawei Wu , Zhi Jin

Vision-language models (VLMs) have shown strong performance on text-to-image retrieval benchmarks. However, bridging this success to real-world applications remains a challenge. In practice, human search behavior is rarely a one-shot…

Computer Vision and Pattern Recognition · Computer Science 2025-10-31 Diji Yang , Minghao Liu , Chung-Hsiang Lo , Yi Zhang , James Davis

Adverse weather severely impairs real-world visual perception, while existing vision models trained on synthetic data with fixed parameters struggle to generalize to complex degradations. To address this, we first construct HFLS-Weather, a…

Computer Vision and Pattern Recognition · Computer Science 2025-11-10 Fuyang Liu , Jiaqi Xu , Xiaowei Hu

Recent advances in diffusion-based Large Restoration Models (LRMs) have significantly improved photo-realistic image restoration by leveraging the internal knowledge embedded within model weights. However, existing LRMs often suffer from…

Computer Vision and Pattern Recognition · Computer Science 2024-10-10 Hang Guo , Tao Dai , Zhihao Ouyang , Taolin Zhang , Yaohua Zha , Bin Chen , Shu-tao Xia

This paper addresses the limitations of adverse weather image restoration approaches trained on synthetic data when applied to real-world scenarios. We formulate a semi-supervised learning framework employing vision-language models to…

Computer Vision and Pattern Recognition · Computer Science 2024-09-04 Jiaqi Xu , Mengyang Wu , Xiaowei Hu , Chi-Wing Fu , Qi Dou , Pheng-Ann Heng

Deceptive reviews, refer to fabricated feedback designed to artificially manipulate the perceived quality of products. Within modern e-commerce ecosystems, these reviews remain a critical governance challenge. Despite advances in…

Information Retrieval · Computer Science 2026-05-08 Nan Lu , Leyang Li , Yurong Hu , Rui Lin , Shaoyi Xu

Question-answering (QA) interfaces powered by large language models (LLMs) present a promising direction for improving interactivity with HVAC system insights, particularly for non-expert users. However, enabling accurate, real-time, and…

Artificial Intelligence · Computer Science 2025-07-08 Sungmin Lee , Minju Kang , Joonhee Lee , Seungyong Lee , Dongju Kim , Jingi Hong , Jun Shin , Pei Zhang , JeongGil Ko

Agent-based editing models have substantially advanced interactive experiences, processing quality, and creative flexibility. However, two critical challenges persist: (1) instruction hallucination, text-only chain-of-thought (CoT)…

Computer Vision and Pattern Recognition · Computer Science 2025-12-05 Yunlong Lin , Linqing Wang , Kunjie Lin , Zixu Lin , Kaixiong Gong , Wenbo Li , Bin Lin , Zhenxi Li , Shiyi Zhang , Yuyang Peng , Wenxun Dai , Xinghao Ding , Chunyu Wang , Qinglin Lu

Real-time scene comprehension is a key advance in artificial intelligence, enhancing robotics, surveillance, and assistive tools. However, hallucination remains a challenge. AI systems often misinterpret visual inputs, detecting nonexistent…

Machine Learning · Computer Science 2025-04-08 Zahir Alsulaimawi

In reality, images often exhibit multiple degradations, such as rain and fog at night (triple degradations). However, in many cases, individuals may not want to remove all degradations, for instance, a blurry lens revealing a beautiful…

Computer Vision and Pattern Recognition · Computer Science 2024-04-17 Runwei Guan , Rongsheng Hu , Zhuhao Zhou , Tianlang Xue , Ka Lok Man , Jeremy Smith , Eng Gee Lim , Weiping Ding , Yutao Yue

Large vision-language models (LVLMs) suffer from hallucination a lot, generating responses that apparently contradict to the image content occasionally. The key problem lies in its weak ability to comprehend detailed content in a…

Computer Vision and Pattern Recognition · Computer Science 2023-11-29 Zhiyang Chen , Yousong Zhu , Yufei Zhan , Zhaowen Li , Chaoyang Zhao , Jinqiao Wang , Ming Tang
‹ Prev 1 2 3 10 Next ›