中文
相关论文

相关论文: UMO: Scaling Multi-Identity Consistency for Image …

200 篇论文

In recent years, "U-shaped" neural networks featuring encoder and decoder structures have gained popularity in the field of medical image segmentation. Various variants of this model have been developed. Nevertheless, the evaluation of…

图像与视频处理 · 电气工程与系统科学 2023-06-02 Qi Ye , Lihua Guo

Multimodal representation learning poses significant challenges in capturing informative and distinct features from multiple modalities. Existing methods often struggle to exploit the unique characteristics of each modality due to unified…

计算机视觉与模式识别 · 计算机科学 2023-11-08 Cam-Van Thi Nguyen , Ngoc-Hoa Thi Nguyen , Duc-Trong Le , Quang-Thuy Ha

The boosting on the need of security notably increased the amount of possible facial recognition applications, especially due to the success of the Internet of Things (IoT) paradigm. However, although handcrafted and deep learning-inspired…

计算机视觉与模式识别 · 计算机科学 2019-12-02 Giulia Orrù , Gian Luca Marcialis , Fabio Roli

Although subject-driven generation has been extensively explored in image generation due to its wide applications, it still has challenges in data scalability and subject expansibility. For the first challenge, moving from curating…

计算机视觉与模式识别 · 计算机科学 2025-04-04 Shaojin Wu , Mengqi Huang , Wenxu Wu , Yufeng Cheng , Fei Ding , Qian He

The RGB-infrared cross-modality person re-identification (ReID) task aims to recognize the images of the same identity between the visible modality and the infrared modality. Existing methods mainly use a two-stream architecture to…

计算机视觉与模式识别 · 计算机科学 2021-10-22 Yajun Gao , Tengfei Liang , Yi Jin , Xiaoyan Gu , Wu Liu , Yidong Li , Congyan Lang

Unified multimodal models (UMMs) were designed to combine the reasoning ability of large language models (LLMs) with the generation capability of vision models. In practice, however, this synergy remains elusive: UMMs fail to transfer…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Songlin Yang , Xianghao Kong , Anyi Rao

Feature visualization has gained substantial popularity, particularly after the influential work by Olah et al. in 2017, which established it as a crucial tool for explainability. However, its widespread adoption has been limited due to a…

We present UniRef-Image-Edit, a high-performance multi-modal generation system that unifies single-image editing and multi-image composition within a single framework. Existing diffusion-based editing methods often struggle to maintain…

Face recognition technology has dramatically transformed the landscape of security, surveillance, and authentication systems, offering a user-friendly and non-invasive biometric solution. However, despite its significant advantages, face…

计算机视觉与模式识别 · 计算机科学 2025-07-04 Arun Kunwar , Ajita Rattani

In this paper, we propose a novel framework, Combo, for harmonious co-speech holistic 3D human motion generation and efficient customizable adaption. In particular, we identify that one fundamental challenge as the…

计算机视觉与模式识别 · 计算机科学 2025-09-22 Chao Xu , Mingze Sun , Zhi-Qi Cheng , Fei Wang , Yang Liu , Baigui Sun , Ruqi Huang , Alexander Hauptmann

Alignment of Large Language Models (LLMs) aims to align outputs with human preferences, and personalized alignment further adapts models to individual users. This relies on personalized reward models that capture user-specific preferences…

计算与语言 · 计算机科学 2026-04-21 Hongru Cai , Yongqi Li , Tiezheng Yu , Fengbin Zhu , Wenjie Wang , Fuli Feng , Wenjie Li

Self-supervised learning on images seeks to extract meaningful visual representations from unlabeled data. When scaled to large datasets, this paradigm has achieved state-of-the-art performance and the resulting trained models such as…

计算机视觉与模式识别 · 计算机科学 2025-11-24 David Nordström , Johan Edstedt , Fredrik Kahl , Georg Bökman

Deep image embedding provides a way to measure the semantic similarity of two images. It plays a central role in many applications such as image search, face verification, and zero-shot learning. It is desirable to have a universal deep…

计算机视觉与模式识别 · 计算机科学 2020-03-10 Yang Feng , Futang Peng , Xu Zhang , Wei Zhu , Shanfeng Zhang , Howard Zhou , Zhen Li , Tom Duerig , Shih-Fu Chang , Jiebo Luo

Multimodal Recommendation (MMR) systems are crucial for modern platforms but are often hampered by inherent noise and uncertainty in modal features, such as blurry images, diverse visual appearances, or ambiguous text. Existing methods…

信息检索 · 计算机科学 2026-01-28 Xinzhuo Wu , Hongbo Wang , Yuan Lin , Kan Xu , Liang Yang , Hongfei Lin

Recent advancements in Unified Multimodal Models (UMMs) have enabled remarkable image understanding and generation capabilities. However, while models like Gemini-2.5-Flash-Image show emerging abilities to reason over multiple related…

计算机视觉与模式识别 · 计算机科学 2026-02-24 Mingrui Wu , Hang Liu , Jiayi Ji , Xiaoshuai Sun , Rongrong Ji

We propose a self-supervised shared encoder model that achieves strong results on several visual, language and multimodal benchmarks while being data, memory and run-time efficient. We make three key contributions. First, in contrast to…

计算机视觉与模式识别 · 计算机科学 2023-04-13 Rakesh Chada , Zhaoheng Zheng , Pradeep Natarajan

Relighting is a crucial task with both practical demand and artistic value, and recent diffusion models have shown strong potential by enabling rich and controllable lighting effects. However, as they are typically optimized in semantic…

计算机视觉与模式识别 · 计算机科学 2025-11-04 Ropeway Liu , Hangjie Yuan , Bo Dong , Jiazheng Xing , Jinwang Wang , Rui Zhao , Yan Xing , Weihua Chen , Fan Wang

The current text-to-video (T2V) generation has made significant progress in synthesizing realistic general videos, but it is still under-explored in identity-specific human video generation with customized ID images. The key challenge lies…

计算机视觉与模式识别 · 计算机科学 2025-03-18 Hengjia Li , Haonan Qiu , Shiwei Zhang , Xiang Wang , Yujie Wei , Zekun Li , Yingya Zhang , Boxi Wu , Deng Cai

Utilizing multi-modal data enhances scene understanding by providing complementary semantic and geometric information. Existing methods fuse features or distill knowledge from multiple modalities into a unified representation, improving…

计算机视觉与模式识别 · 计算机科学 2025-06-05 Jialei Chen , Xu Zheng , Danda Pani Paudel , Luc Van Gool , Hiroshi Murase , Daisuke Deguchi

The goal of creating intelligent, human-centered wearable systems for continuous activity understanding faces a fundamental trade-off: Egocentric video-based models capture rich semantic information and have demonstrated strong performance…

计算机视觉与模式识别 · 计算机科学 2026-04-22 Baiyu Chen , Wilson Wongso , Zechen Li , Yonchanok Khaokaew , Hao Xue , Flora Salim