English
Related papers

Related papers: Semantic Alignment in Hyperbolic Space for Open-Vo…

200 papers

Visual and linguistic concepts naturally organize themselves in a hierarchy, where a textual concept "dog" entails all images that contain dogs. Despite being intuitive, current large-scale vision and language models such as CLIP do not…

Computer Vision and Pattern Recognition · Computer Science 2024-01-19 Karan Desai , Maximilian Nickel , Tanmay Rajpurohit , Justin Johnson , Ramakrishna Vedantam

For natural language understanding and generation, embedding concepts using an order-based representation is an essential task. Unlike traditional point vector based representation, an order-based representation imposes geometric…

Computation and Language · Computer Science 2024-04-18 Croix Gyurek , Niloy Talukder , Mohammad Al Hasan

Hyperbolic spaces, which have the capacity to embed tree structures without distortion owing to their exponential volume growth, have recently been applied to machine learning to better capture the hierarchical nature of data. In this…

Machine Learning · Computer Science 2021-03-18 Ryohei Shimizu , Yusuke Mukuta , Tatsuya Harada

Recent progress in artificial intelligence has encouraged numerous attempts to understand and decode human visual system from brain signals. These prior works typically align neural activity independently with semantic and perceptual…

Artificial Intelligence · Computer Science 2026-03-25 Sangmin Jo , Wootaek Jeong , Da-Woon Heo , Yoohwan Hwang , Heung-Il Suk

Recently, the strong generalization ability of CLIP has facilitated open-vocabulary semantic segmentation, which labels pixels using arbitrary text. However, existing methods that fine-tune CLIP for segmentation on limited seen categories…

Computer Vision and Pattern Recognition · Computer Science 2025-11-21 Muyao Yuan , Yuanhong Zhang , Weizhan Zhang , Lan Ma , Yuan Gao , Jiangyong Ying , Yudeng Xin

This paper presents a novel hierarchical alignment model (HAM) that learns multi-granularity visual and linguistic representations in an end-to-end manner. We extract key points and proposal points to model 3D contexts and instances, and…

Computer Vision and Pattern Recognition · Computer Science 2023-06-12 Jiaming Chen , Weixin Luo , Ran Song , Xiaolin Wei , Lin Ma , Wei Zhang

Vision-language models have achieved remarkable success in multi-modal representation learning from large-scale pairs of visual scenes and linguistic descriptions. However, they still struggle to simultaneously express two distinct types of…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Daiki Yoshikawa , Takashi Matsubara

Retrieval-augmented generation (RAG) for biomedical knowledge faces a hierarchy-aware ontology grounding challenge: resources like HPO, DO, and MeSH use deep ``is-a" taxonomies, yet production stacks rely on Euclidean embeddings and ANN…

Information Retrieval · Computer Science 2026-04-14 Ou Deng , Shoji Nishimura , Atsushi Ogihara , Qun Jin

The global rise in the number of people with physical disabilities, in part due to improvements in post-trauma survivorship and longevity, has amplified the demand for advanced assistive technologies to improve mobility and independence.…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Yifan Xu , Vineet Kamat , Carol Menassa

Current breakthroughs in natural language processing have benefited dramatically from neural language models, through which distributional semantics can leverage neural data representations to facilitate downstream applications. Since…

Computation and Language · Computer Science 2022-10-04 Dongqiang Yang , Ning Li , Li Zou , Hongwei Ma

Addressing the retrieval of unsafe content from vision-language models such as CLIP is an important step towards real-world integration. Current efforts have relied on unlearning techniques that try to erase the model's knowledge of unsafe…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Tobia Poppi , Tejaswi Kasarla , Pascal Mettes , Lorenzo Baraldi , Rita Cucchiara

Open-vocabulary semantic segmentation in the remote sensing (RS) field requires both language-aligned recognition and fine-grained spatial delineation. Although CLIP offers robust semantic generalization, its global-aligned visual…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Jie Feng , Fengze Li , Junpeng Zhang , Siyu Chen , Yuping Liang , Junying Chen , Ronghua Shang

Representing data in hyperbolic space can effectively capture latent hierarchical relationships. With the goal of enabling accurate classification of points in hyperbolic space while respecting their hyperbolic geometry, we introduce…

Machine Learning · Computer Science 2018-06-04 Hyunghoon Cho , Benjamin DeMeo , Jian Peng , Bonnie Berger

With the rapid development of text-to-image generation technology, accurately assessing the alignment between generated images and text prompts has become a critical challenge. Existing methods rely on Euclidean space metrics, neglecting…

Computer Vision and Pattern Recognition · Computer Science 2026-03-20 Wenzhi Chen , Bo Hu , Leida Li , Lihuo He , Wen Lu , Xinbo Gao

The organization of latent token representations plays a crucial role in determining the stability, generalization, and contextual consistency of language models, yet conventional approaches to embedding refinement often rely on parameter…

Computation and Language · Computer Science 2025-03-26 Meiquan Dong , Haoran Liu , Yan Huang , Zixuan Feng , Jianhong Tang , Ruoxi Wang

Recent research in Vision-Language Models (VLMs) has significantly advanced our capabilities in cross-modal reasoning. However, existing methods suffer from performance degradation with domain changes or require substantial computational…

Computer Vision and Pattern Recognition · Computer Science 2025-12-10 Yi Zhang , Chun-Wun Cheng , Junyi He , Ke Yu , Yushun Tang , Carola-Bibiane Schönlieb , Zhihai He , Angelica I. Aviles-Rivero

Open-vocabulary segmentation aims to achieve segmentation of arbitrary categories given unlimited text inputs as guidance. To achieve this, recent works have focused on developing various technical routes to exploit the potential of…

Computer Vision and Pattern Recognition · Computer Science 2025-06-25 Yong Liu , SongLi Wu , Sule Bai , Jiahao Wang , Yitong Wang , Yansong Tang

As bird's-eye-view (BEV) semantic segmentation is simple-to-visualize and easy-to-handle, it has been applied in autonomous driving to provide the surrounding information to downstream tasks. Inferring BEV semantic segmentation conditioned…

Computer Vision and Pattern Recognition · Computer Science 2023-08-21 Naiyu Fang , Lemiao Qiu , Shuyou Zhang , Zili Wang , Kerui Hu , Kang Wang

Electronic health records (EHR) contain narrative notes that provide extensive details on the medical condition and management of patients. Natural language processing (NLP) of clinical notes can use observed frequencies of clinical terms…

Computation and Language · Computer Science 2023-07-04 Bryan Cai , Sihang Zeng , Yucong Lin , Zheng Yuan , Doudou Zhou , Lu Tian

Semantic segmentation (SS) aims to classify each pixel into one of the pre-defined classes. This task plays an important role in self-driving cars and autonomous drones. In SS, many works have shown that most misclassified pixels are…

Computer Vision and Pattern Recognition · Computer Science 2023-05-29 Bike Chen , Wei Peng , Xiaofeng Cao , Juha Röning