中文
相关论文

相关论文: PositionIC: Unified Position and Identity Consiste…

200 篇论文

In unsupervised domain adaptation (UDA), a model trained on source data (e.g. synthetic) is adapted to target data (e.g. real-world) without access to target annotation. Most previous UDA methods struggle with classes that have a similar…

计算机视觉与模式识别 · 计算机科学 2023-03-27 Lukas Hoyer , Dengxin Dai , Haoran Wang , Luc Van Gool

A holistic understanding of object properties across diverse sensory modalities (e.g., visual, audio, and haptic) is essential for tasks ranging from object categorization to complex manipulation. Drawing inspiration from cognitive science…

机器人学 · 计算机科学 2024-02-26 Gyan Tatiya , Jonathan Francis , Ho-Hsiang Wu , Yonatan Bisk , Jivko Sinapov

In recent years, attention mechanisms have significantly enhanced the performance of object detection by focusing on key feature information. However, prevalent methods still encounter difficulties in effectively balancing local and global…

计算机视觉与模式识别 · 计算机科学 2024-11-15 Yifan Shao

There is a rapidly growing interest in controlling consistency across multiple generated images using diffusion models. Among various methods, recent works have found that simply manipulating attention modules by concatenating features from…

计算机视觉与模式识别 · 计算机科学 2024-05-29 Jiaojiao Fan , Haotian Xue , Qinsheng Zhang , Yongxin Chen

Text-to-image (T2I) models excel on single-entity prompts but struggle with multi-entity scenes, often exhibiting attribute leakage, identity entanglement, and subject omissions. We present a principled theoretical framework that steers…

计算机视觉与模式识别 · 计算机科学 2026-03-26 Eric Tillmann Bill , Enis Simsar , Thomas Hofmann

Conditional image generation methods are increasingly used in human-centric applications, yet existing human amodal completion (HAC) models offer users limited control over the completed content. Given an occluded person image, they…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Seung Young Noh , Ju Yong Chang

Medical image classification plays a crucial role in computer-aided clinical diagnosis. While deep learning techniques have significantly enhanced efficiency and reduced costs, the privacy-sensitive nature of medical imaging data…

计算机视觉与模式识别 · 计算机科学 2024-07-04 Sufen Ren , Yule Hu , Shengchao Chen , Guanjun Wang

Image to point cloud global localization is crucial for robot navigation in GNSS-denied environments and has become increasingly important for multi-robot map fusion and urban asset management. The modality gap between images and point…

计算机视觉与模式识别 · 计算机科学 2024-12-23 Yuhao Li , Jianping Li , Zhen Dong , Yuan Wang , Bisheng Yang

Multi-modal medical image segmentation plays an essential role in clinical diagnosis. It remains challenging as the input modalities are often not well-aligned spatially. Existing learning-based methods mainly consider sharing trainable…

计算机视觉与模式识别 · 计算机科学 2021-01-06 Jingkun Chen , Wenqi Li , Hongwei Li , Jianguo Zhang

LiDAR relocalization aims to estimate the global 6-DoF pose of a sensor in the environment. However, existing regression-based approaches are prone to dynamic or ambiguous scenarios, as they either solely rely on single-frame inference or…

计算机视觉与模式识别 · 计算机科学 2026-02-04 Minghang Zhu , Zhijing Wang , Yuxin Guo , Wen Li , Sheng Ao , Cheng Wang

Multimodal image fusion and semantic segmentation are critical for autonomous driving. Despite advancements, current models often struggle with segmenting densely packed elements due to a lack of comprehensive fusion features for guidance…

计算机视觉与模式识别 · 计算机科学 2025-06-25 Daixun Li , Weiying Xie , Mingxiang Cao , Yunke Wang , Yusi Zhang , Leyuan Fang , Yunsong Li , Chang Xu

Multi-Instance Generation has advanced significantly in spatial placement and attribute binding. However, existing approaches still face challenges in fine-grained semantic understanding, particularly when dealing with complex textual…

计算机视觉与模式识别 · 计算机科学 2026-02-23 Shiyan Du , Conghan Yue , Xinyu Cheng , Dongyu Zhang

Maintaining an up-to-date map to reflect recent changes in the scene is very important, particularly in situations involving repeated traversals by a robot operating in an environment over an extended period. Undetected changes may cause a…

机器人学 · 计算机科学 2022-07-18 Jingxing Qian , Veronica Chatrath , Jun Yang , James Servos , Angela P. Schoellig , Steven L. Waslander

Recent advances in high-throughput sequencing technologies have enabled the extraction of multiple features that depict patient samples at diverse and complementary molecular levels. The generation of such data has led to new challenges in…

基因组学 · 定量生物学 2022-09-14 Hakim Benkirane , Yoann Pradat , Stefan Michiels , Paul-Henry Cournède

Perceptual aliasing and weak textures pose significant challenges to the task of place recognition, hindering the performance of Simultaneous Localization and Mapping (SLAM) systems. This paper presents a novel model, called UMF (standing…

计算机视觉与模式识别 · 计算机科学 2024-10-10 Alberto García-Hernández , Riccardo Giubilato , Klaus H. Strobl , Javier Civera , Rudolph Triebel

Mainstream image caption models are usually two-stage captioners, i.e., calculating object features by pre-trained detector, and feeding them into a language model to generate text descriptions. However, such an operation will cause a…

计算机视觉与模式识别 · 计算机科学 2022-11-07 Bo Wang , Zhao Zhang , Mingbo Zhao , Xiaojie Jin , Mingliang Xu , Meng Wang

The performance of multi-modal 3D occupancy prediction is limited by ineffective fusion, mainly due to geometry-semantics mismatch from fixed fusion strategies and surface detail loss caused by sparse, noisy annotations. The mismatch stems…

计算机视觉与模式识别 · 计算机科学 2025-05-20 Luyao Lei , Shuo Xu , Yifan Bai , Xing Wei

The increasing demand for controllable outputs in text-to-image generation has spurred advancements in multi-instance generation (MIG), allowing users to define both instance layouts and attributes. However, unlike image-conditional…

计算机视觉与模式识别 · 计算机科学 2025-12-03 Dewei Zhou , Ji Xie , Zongxin Yang , Yi Yang

We introduce MOSAIC (Masked Objective with Selective Adaptation for In-domain Contrastive learning), a multi-stage framework for domain adaptation of text embedding models that incorporates joint domain-specific masked supervision. Our…

计算与语言 · 计算机科学 2026-01-30 Vera Pavlova , Mohammed Makhlouf

This paper aims to model 3D human motion across domains, where a single model is expected to handle multiple modalities, tasks, and datasets. Existing cross-domain models often rely on domain-specific components and multi-stage training,…

计算机视觉与模式识别 · 计算机科学 2025-08-15 Mengyuan Liu , Xinshun Wang , Zhongbin Fang , Deheng Ye , Xia Li , Tao Tang , Songtao Wu , Xiangtai Li , Ming-Hsuan Yang