中文
相关论文

相关论文: UMO: Scaling Multi-Identity Consistency for Image …

200 篇论文

Multi-subject video generation aims to synthesize videos from textual prompts and multiple reference images, ensuring that each subject preserves natural scale and visual fidelity. However, current methods face two challenges: scale…

计算机视觉与模式识别 · 计算机科学 2025-12-30 Run Ling , Ke Cao , Jian Lu , Ao Ma , Haowei Liu , Runze He , Changwei Wang , Rongtao Xu , Yihua Shao , Zhanjie Zhang , Peng Wu , Guibing Guo , Wei Feng , Zheng Zhang , Jingjing Lv , Junjie Shen , Ching Law , Xingwei Wang

Person re-identification (Re-ID) is a challenging task that involves identifying the same person across different camera views in surveillance systems. Current methods usually rely on features from single-camera views, which can be limiting…

计算机视觉与模式识别 · 计算机科学 2025-05-06 Quang-Huy Che , Le-Chuong Nguyen , Duc-Tuan Luu , Vinh-Tiep Nguyen

Multi-subject personalized image generation aims to synthesize customized images containing multiple specified subjects without requiring test-time optimization. However, achieving fine-grained independent control over multiple subjects…

计算机视觉与模式识别 · 计算机科学 2025-09-03 Qiaoqiao Jin , Siming Fu , Dong She , Weinan Jia , Hualiang Wang , Mu Liu , Jidong Jiang

Software configuration tuning is essential for optimizing a given performance objective (e.g., minimizing latency). Yet, due to the software's intrinsically complex configuration landscape and expensive measurement, there has been a rather…

软件工程 · 计算机科学 2024-03-18 Pengzhou Chen , Tao Chen , Miqing Li

Video identity customization seeks to synthesize realistic, temporally coherent videos of a specific subject, given a single reference image and a text prompt. This task presents two core challenges: (1) maintaining identity consistency…

计算机视觉与模式识别 · 计算机科学 2026-02-04 Guiyu Zhang , Chen Shi , Zijian Jiang , Xunzhi Xiang , Jingjing Qian , Shaoshuai Shi , Li Jiang

Diffusion-based text-to-image generation has advanced significantly, yet customizing scenes with multiple distinct subjects while maintaining fine-grained control over their interactions remains challenging. Existing methods often struggle…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Pengxiang Cai , Mengyang Li

Visual captioning aims to generate textual descriptions given images or videos. Traditionally, image captioning models are trained on human annotated datasets such as Flickr30k and MS-COCO, which are limited in size and diversity. This…

计算机视觉与模式识别 · 计算机科学 2021-03-01 Marimuthu Kalimuthu , Aditya Mogadala , Marius Mosbach , Dietrich Klakow

3D Mixed Reality interfaces have nearly unlimited space for layout placement, making automatic UI adaptation crucial for enhancing the user experience. Such adaptation is often formulated as a multi-objective optimization (MOO) problem,…

人机交互 · 计算机科学 2025-09-24 Yao Song , Christoph Gebhardt , Yi-Chi Liao , Christian Holz

The Quality-Diversity (QD) optimization aims to discover a collection of high-performing solutions that simultaneously exhibit diverse behaviors within a user-defined behavior space. This paradigm has stimulated significant research…

机器学习 · 计算机科学 2026-02-03 Xi Lin , Ping Guo , Yilu Liu , Qingfu Zhang , Jianyong Sun

Recent advancements in Multimodal Large Language Models (MLLMs) underscore the significance of scalable models and data to boost performance, yet this often incurs substantial computational costs. Although the Mixture of Experts (MoE)…

人工智能 · 计算机科学 2024-05-21 Yunxin Li , Shenyuan Jiang , Baotian Hu , Longyue Wang , Wanqi Zhong , Wenhan Luo , Lin Ma , Min Zhang

Blind face restoration is a highly ill-posed problem due to the lack of necessary context. Although existing methods produce high-quality outputs, they often fail to faithfully preserve the individual's identity. In this paper, we propose a…

计算机视觉与模式识别 · 计算机科学 2025-01-13 Siyu Liu , Zheng-Peng Duan , Jia OuYang , Jiayi Fu , Hyunhee Park , Zikun Liu , Chun-Le Guo , Chongyi Li

Image customization, a crucial technique for industrial media production, aims to generate content that is consistent with reference images. However, current approaches conventionally separate image customization into position-aware and…

计算机视觉与模式识别 · 计算机科学 2025-10-02 Yaowei Li , Xiaoyu Li , Zhaoyang Zhang , Yuxuan Bian , Gan Liu , Xinyuan Li , Jiale Xu , Wenbo Hu , Yating Liu , Lingen Li , Jing Cai , Yuexian Zou , Yancheng He , Ying Shan

The fusion of multiple sensor modalities, especially through deep learning architectures, has been an active area of study. However, an under-explored aspect of such work is whether the methods can be robust to degradations across their…

计算机视觉与模式识别 · 计算机科学 2020-03-05 Junjiao Tian , Wesley Cheung , Nathan Glaser , Yen-Cheng Liu , Zsolt Kira

Identity-consistent generation has become an important focus in text-to-image research, with recent models achieving notable success in producing images aligned with a reference identity. Yet, the scarcity of large-scale paired datasets…

计算机视觉与模式识别 · 计算机科学 2025-10-17 Hengyuan Xu , Wei Cheng , Peng Xing , Yixiao Fang , Shuhan Wu , Rui Wang , Xianfang Zeng , Daxin Jiang , Gang Yu , Xingjun Ma , Yu-Gang Jiang

Multi-Objective Optimization (MOO) techniques have become increasingly popular in recent years due to their potential for solving real-world problems in various fields, such as logistics, finance, environmental management, and engineering.…

神经与进化计算 · 计算机科学 2024-07-15 Noor A. Rashed , Yossra H. Ali , Tarik A. Rashid , A. Salih

Identifying user's identity is a key problem in many data mining applications, such as product recommendation, customized content delivery and criminal identification. Given a set of accounts from the same or different social network…

计算机视觉与模式识别 · 计算机科学 2016-10-26 Xiang Jiang , Shikui Wei , Ruizhen Zhao , Yao Zhao , Xindong Wu

Unified multimodal models (UMMs) aim to integrate multimodal understanding and generation within a unified architecture, yet it remains unclear to what extent their representations are truly aligned across modalities. To investigate this…

计算与语言 · 计算机科学 2026-04-08 Cheng Yang , Chufan Shi , Bo Shui , Yaokang Wu , Muzi Tao , Huijuan Wang , Ivan Yee Lee , Yong Liu , Xuezhe Ma , Taylor Berg-Kirkpatrick

Large language models (LLMs) are shifting from answer providers to intelligent tutors in educational settings, yet current supervised fine-tuning methods only learn surface teaching patterns without dynamic adaptation capabilities. Recent…

人工智能 · 计算机科学 2026-01-06 Shouang Wei , Min Zhang , Xin Lin , Bo Jiang , Kun Kuang , Zhongxiang Dai

When compared to unimodal systems, multimodal biometric systems have several advantages, including lower error rate, higher accuracy, and larger population coverage. However, multimodal systems have an increased demand for integrity and…

计算机视觉与模式识别 · 计算机科学 2021-01-01 Veeru Talreja , Matthew Valenti , Nasser Nasrabadi

Recent advancements in Multimodal Large Language Models (LLMs) have focused primarily on scaling by increasing text-image pair data and enhancing LLMs to improve performance on multimodal tasks. However, these scaling approaches are…

计算机视觉与模式识别 · 计算机科学 2024-05-10 Jiachen Li , Xinyao Wang , Sijie Zhu , Chia-Wen Kuo , Lu Xu , Fan Chen , Jitesh Jain , Humphrey Shi , Longyin Wen