English
Related papers

Related papers: PositionIC: Unified Position and Identity Consiste…

200 papers

In unsupervised domain adaptation (UDA), a model trained on source data (e.g. synthetic) is adapted to target data (e.g. real-world) without access to target annotation. Most previous UDA methods struggle with classes that have a similar…

Computer Vision and Pattern Recognition · Computer Science 2023-03-27 Lukas Hoyer , Dengxin Dai , Haoran Wang , Luc Van Gool

A holistic understanding of object properties across diverse sensory modalities (e.g., visual, audio, and haptic) is essential for tasks ranging from object categorization to complex manipulation. Drawing inspiration from cognitive science…

Robotics · Computer Science 2024-02-26 Gyan Tatiya , Jonathan Francis , Ho-Hsiang Wu , Yonatan Bisk , Jivko Sinapov

In recent years, attention mechanisms have significantly enhanced the performance of object detection by focusing on key feature information. However, prevalent methods still encounter difficulties in effectively balancing local and global…

Computer Vision and Pattern Recognition · Computer Science 2024-11-15 Yifan Shao

There is a rapidly growing interest in controlling consistency across multiple generated images using diffusion models. Among various methods, recent works have found that simply manipulating attention modules by concatenating features from…

Computer Vision and Pattern Recognition · Computer Science 2024-05-29 Jiaojiao Fan , Haotian Xue , Qinsheng Zhang , Yongxin Chen

Text-to-image (T2I) models excel on single-entity prompts but struggle with multi-entity scenes, often exhibiting attribute leakage, identity entanglement, and subject omissions. We present a principled theoretical framework that steers…

Computer Vision and Pattern Recognition · Computer Science 2026-03-26 Eric Tillmann Bill , Enis Simsar , Thomas Hofmann

Conditional image generation methods are increasingly used in human-centric applications, yet existing human amodal completion (HAC) models offer users limited control over the completed content. Given an occluded person image, they…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Seung Young Noh , Ju Yong Chang

Medical image classification plays a crucial role in computer-aided clinical diagnosis. While deep learning techniques have significantly enhanced efficiency and reduced costs, the privacy-sensitive nature of medical imaging data…

Computer Vision and Pattern Recognition · Computer Science 2024-07-04 Sufen Ren , Yule Hu , Shengchao Chen , Guanjun Wang

Image to point cloud global localization is crucial for robot navigation in GNSS-denied environments and has become increasingly important for multi-robot map fusion and urban asset management. The modality gap between images and point…

Computer Vision and Pattern Recognition · Computer Science 2024-12-23 Yuhao Li , Jianping Li , Zhen Dong , Yuan Wang , Bisheng Yang

Multi-modal medical image segmentation plays an essential role in clinical diagnosis. It remains challenging as the input modalities are often not well-aligned spatially. Existing learning-based methods mainly consider sharing trainable…

Computer Vision and Pattern Recognition · Computer Science 2021-01-06 Jingkun Chen , Wenqi Li , Hongwei Li , Jianguo Zhang

LiDAR relocalization aims to estimate the global 6-DoF pose of a sensor in the environment. However, existing regression-based approaches are prone to dynamic or ambiguous scenarios, as they either solely rely on single-frame inference or…

Computer Vision and Pattern Recognition · Computer Science 2026-02-04 Minghang Zhu , Zhijing Wang , Yuxin Guo , Wen Li , Sheng Ao , Cheng Wang

Multimodal image fusion and semantic segmentation are critical for autonomous driving. Despite advancements, current models often struggle with segmenting densely packed elements due to a lack of comprehensive fusion features for guidance…

Computer Vision and Pattern Recognition · Computer Science 2025-06-25 Daixun Li , Weiying Xie , Mingxiang Cao , Yunke Wang , Yusi Zhang , Leyuan Fang , Yunsong Li , Chang Xu

Multi-Instance Generation has advanced significantly in spatial placement and attribute binding. However, existing approaches still face challenges in fine-grained semantic understanding, particularly when dealing with complex textual…

Computer Vision and Pattern Recognition · Computer Science 2026-02-23 Shiyan Du , Conghan Yue , Xinyu Cheng , Dongyu Zhang

Maintaining an up-to-date map to reflect recent changes in the scene is very important, particularly in situations involving repeated traversals by a robot operating in an environment over an extended period. Undetected changes may cause a…

Recent advances in high-throughput sequencing technologies have enabled the extraction of multiple features that depict patient samples at diverse and complementary molecular levels. The generation of such data has led to new challenges in…

Genomics · Quantitative Biology 2022-09-14 Hakim Benkirane , Yoann Pradat , Stefan Michiels , Paul-Henry Cournède

Perceptual aliasing and weak textures pose significant challenges to the task of place recognition, hindering the performance of Simultaneous Localization and Mapping (SLAM) systems. This paper presents a novel model, called UMF (standing…

Computer Vision and Pattern Recognition · Computer Science 2024-10-10 Alberto García-Hernández , Riccardo Giubilato , Klaus H. Strobl , Javier Civera , Rudolph Triebel

Mainstream image caption models are usually two-stage captioners, i.e., calculating object features by pre-trained detector, and feeding them into a language model to generate text descriptions. However, such an operation will cause a…

Computer Vision and Pattern Recognition · Computer Science 2022-11-07 Bo Wang , Zhao Zhang , Mingbo Zhao , Xiaojie Jin , Mingliang Xu , Meng Wang

The performance of multi-modal 3D occupancy prediction is limited by ineffective fusion, mainly due to geometry-semantics mismatch from fixed fusion strategies and surface detail loss caused by sparse, noisy annotations. The mismatch stems…

Computer Vision and Pattern Recognition · Computer Science 2025-05-20 Luyao Lei , Shuo Xu , Yifan Bai , Xing Wei

The increasing demand for controllable outputs in text-to-image generation has spurred advancements in multi-instance generation (MIG), allowing users to define both instance layouts and attributes. However, unlike image-conditional…

Computer Vision and Pattern Recognition · Computer Science 2025-12-03 Dewei Zhou , Ji Xie , Zongxin Yang , Yi Yang

We introduce MOSAIC (Masked Objective with Selective Adaptation for In-domain Contrastive learning), a multi-stage framework for domain adaptation of text embedding models that incorporates joint domain-specific masked supervision. Our…

Computation and Language · Computer Science 2026-01-30 Vera Pavlova , Mohammed Makhlouf

This paper aims to model 3D human motion across domains, where a single model is expected to handle multiple modalities, tasks, and datasets. Existing cross-domain models often rely on domain-specific components and multi-stage training,…

Computer Vision and Pattern Recognition · Computer Science 2025-08-15 Mengyuan Liu , Xinshun Wang , Zhongbin Fang , Deheng Ye , Xia Li , Tao Tang , Songtao Wu , Xiangtai Li , Ming-Hsuan Yang