English
Related papers

Related papers: PEARL: Geometry Aligns Semantics for Training-Free…

200 papers

Prompt learning has emerged as a promising method for adapting pre-trained visual-language models (VLMs) to a range of downstream tasks. While optimizing the context can be effective for improving performance on specific tasks, it can often…

Computation and Language · Computer Science 2025-06-04 Fangming Cui , Jan Fong , Rongfei Zeng , Xinmei Tian , Jun Yu

While Contrastive Language-Image Pre-training (CLIP) has advanced open-vocabulary predictions, its performance on semantic segmentation remains suboptimal. This shortfall primarily stems from its spatial-invariant semantic features and…

Computer Vision and Pattern Recognition · Computer Science 2024-11-15 Yuheng Shi , Minjing Dong , Chang Xu

Open-vocabulary scene understanding using 3D Gaussian (3DGS) representations has garnered considerable attention. However, existing methods mostly lift knowledge from large 2D vision models into 3DGS on a scene-by-scene basis, restricting…

Computer Vision and Pattern Recognition · Computer Science 2025-01-03 Runnan Chen , Xiangyu Sun , Zhaoqing Wang , Youquan Liu , Jiepeng Wang , Lingdong Kong , Jiankang Deng , Mingming Gong , Liang Pan , Wenping Wang , Tongliang Liu

3D object segmentation with Large Language Models (LLMs) has become a prevailing paradigm due to its broad semantics, task flexibility, and strong generalization. However, this paradigm is hindered by representation misalignment: LLMs…

Computer Vision and Pattern Recognition · Computer Science 2026-02-20 Zhuoxu Huang , Mingqi Gao , Jungong Han

Referring image segmentation (RIS) aims to segment an object mentioned in natural language from an image. The main challenge is text-to-pixel fine-grained correlation. In the previous methods, the final results are obtained by convolutions…

Computer Vision and Pattern Recognition · Computer Science 2024-10-15 Yichen Yan , Xingjian He , Wenxuan Wang , Sihan Chen , Jing Liu

Foundation models have exhibited unprecedented capabilities in tackling many domains and tasks. Models such as CLIP are currently widely used to bridge cross-modal representations, and text-to-image diffusion models are arguably the leading…

Computer Vision and Pattern Recognition · Computer Science 2025-11-19 Barbara Toniella Corradini , Mustafa Shukor , Paul Couairon , Guillaume Couairon , Franco Scarselli , Matthieu Cord

Text-to-video (T2V) generation aims to synthesize videos with high visual quality and temporal consistency that are semantically aligned with input text. Reward-based post-training has emerged as a promising direction to improve the quality…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Minh-Quan Le , Gaurav Mittal , Cheng Zhao , David Gu , Dimitris Samaras , Mei Chen

Recently, open-vocabulary image classification by vision language pre-training has demonstrated incredible achievements, that the model can classify arbitrary categories without seeing additional annotated images of that category. However,…

Computer Vision and Pattern Recognition · Computer Science 2023-01-02 Mengde Xu , Zheng Zhang , Fangyun Wei , Yutong Lin , Yue Cao , Han Hu , Xiang Bai

With the rapid advances in diffusion models, generating decent images from text prompts is no longer challenging. The key to text-to-image generation is how to optimize the results of a text-to-image generation model so that they can be…

Computer Vision and Pattern Recognition · Computer Science 2024-10-15 Xiwen Wang , Jizhe Zhou , Xuekang Zhu , Cheng Li , Mao Li

This paper presents a new method for the zero-shot open-vocabulary semantic segmentation (OVSS) of 3D automotive lidar data. To circumvent the recognized image-text modality gap that is intrinsic to approaches based on Vision Language…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Nermin Samet , Gilles Puy , Renaud Marlet

Leveraging recent diffusion models, LiDAR-based large-scale 3D scene generation has achieved great success. While recent voxel-based approaches can generate both geometric structures and semantic labels, existing range-view methods are…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Dekai Zhu , Yixuan Hu , Youquan Liu , Dongyue Lu , Lingdong Kong , Slobodan Ilic

Conversational semantic role labeling (CSRL) is a newly proposed task that uncovers the shallow semantic structures in a dialogue text. Unfortunately several important characteristics of the CSRL task have been overlooked by the existing…

Computation and Language · Computer Science 2022-10-07 Hao Fei , Shengqiong Wu , Meishan Zhang , Yafeng Ren , Donghong Ji

Open-vocabulary segmentation aims to achieve segmentation of arbitrary categories given unlimited text inputs as guidance. To achieve this, recent works have focused on developing various technical routes to exploit the potential of…

Computer Vision and Pattern Recognition · Computer Science 2025-06-25 Yong Liu , SongLi Wu , Sule Bai , Jiahao Wang , Yitong Wang , Yansong Tang

Recent work uses reinforcement learning (RL) to fine-tune text-to-image diffusion models, improving text-image alignment and sample quality. However, existing approaches introduce unnecessary complexity: they cache the full sampling…

Machine Learning · Computer Science 2025-07-02 Yanting Miao , William Loh , Pacal Poupart , Suraj Kothawade

We introduce the first zero-shot approach for Video Semantic Segmentation (VSS) based on pre-trained diffusion models. A growing research direction attempts to employ diffusion models to perform downstream vision tasks by exploiting their…

Computer Vision and Pattern Recognition · Computer Science 2024-05-28 Qian Wang , Abdelrahman Eldesokey , Mohit Mendiratta , Fangneng Zhan , Adam Kortylewski , Christian Theobalt , Peter Wonka

Multimodal large language models (MLLMs) have recently achieved impressive general-purpose vision-language capabilities through visual instruction tuning. However, current MLLMs primarily focus on image-level or box-level understanding,…

Computer Vision and Pattern Recognition · Computer Science 2025-09-09 Yuqian Yuan , Wentong Li , Jian Liu , Dongqi Tang , Xinjie Luo , Chi Qin , Lei Zhang , Jianke Zhu

Deep learning has thrived by training on large-scale datasets. However, in robotics applications sample efficiency is critical. We propose a novel adaptive masked proxies method that constructs the final segmentation layer weights from few…

Computer Vision and Pattern Recognition · Computer Science 2019-10-16 Mennatullah Siam , Boris Oreshkin , Martin Jagersand

Recent advances in robust semi-supervised learning (SSL) typically filter out-of-distribution (OOD) information at the sample level. We argue that an overlooked problem of robust SSL is its corrupted information on semantic level,…

Computer Vision and Pattern Recognition · Computer Science 2023-05-31 Yu Wang , Pengchong Qiao , Chang Liu , Guoli Song , Xiawu Zheng , Jie Chen

Open-vocabulary segmentation models such as SAM3 perform well across broad categories via text prompting, yet degrade when target classes are visually underrepresented in pretraining or depart from canonical depictions-limitations text…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Abderrahmene Boudiaf , Irfan Hussain , Sajid Javed

High-resolution remote sensing images contain densely distributed objects with pronounced scale variations and complex boundaries, which impose higher demands on both the geometric localization and semantic prediction capabilities of…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Jianzheng Wang , Huan Ni