中文
相关论文

相关论文: Omni-ID: Holistic Identity Representation Designed…

200 篇论文

We introduce Omni-MMSI, a new task that requires comprehensive social interaction understanding from raw audio, vision, and speech input. The task involves perceiving identity-attributed social cues (e.g., who is speaking what) and…

计算机视觉与模式识别 · 计算机科学 2026-04-02 Xinpeng Li , Bolin Lai , Hardy Chen , Shijian Deng , Cihang Xie , Yuyin Zhou , James Matthew Rehg , Yapeng Tian

This paper presents OmniDataComposer, an innovative approach for multimodal data fusion and unlimited data generation with an intent to refine and uncomplicate interplay among diverse data modalities. Coming to the core breakthrough, it…

计算机视觉与模式识别 · 计算机科学 2023-08-21 Dongyang Yu , Shihao Wang , Yuan Fang , Wangpeng An

Person re-identification (re-id) remains challenging due to significant intra-class variations across different cameras. Recently, there has been a growing interest in using generative models to augment training data and enhance the…

计算机视觉与模式识别 · 计算机科学 2021-05-20 Zhedong Zheng , Xiaodong Yang , Zhiding Yu , Liang Zheng , Yi Yang , Jan Kautz

Existing Image Manipulation Localization (IML) methods mostly rely heavily on task-specific designs, making them perform well only on the target IML task, while joint training on multiple IML tasks causes significant performance…

计算机视觉与模式识别 · 计算机科学 2025-04-30 Chenfan Qu , Yiwu Zhong , Fengjun Guo , Lianwen Jin

This paper presents UniPortrait, an innovative human image personalization framework that unifies single- and multi-ID customization with high face fidelity, extensive facial editability, free-form input description, and diverse layout…

计算机视觉与模式识别 · 计算机科学 2024-09-09 Junjie He , Yifeng Geng , Liefeng Bo

Recent face generation methods have tried to synthesize faces based on the given contour condition, like a low-resolution image or sketch. However, the problem of identity ambiguity remains unsolved, which usually occurs when the contour is…

计算机视觉与模式识别 · 计算机科学 2022-08-03 Qingyan Bai , Weihao Xia , Fei Yin , Yujiu Yang

Drawing on recent advancements in diffusion models for text-to-image generation, identity-preserved personalization has made significant progress in accurately capturing specific identities with just a single reference image. However,…

计算机视觉与模式识别 · 计算机科学 2024-03-19 Yi Wu , Ziqiang Li , Heliang Zheng , Chaoyue Wang , Bin Li

Multimodal Large Language Models (MLLMs) are making significant progress in multimodal reasoning. Early approaches focus on pure text-based reasoning. More recent studies have incorporated multimodal information into the reasoning steps;…

人工智能 · 计算机科学 2026-04-21 Dongjie Cheng , Yongqi Li , Zhixin Ma , Hongru Cai , Yupeng Hu , Wenjie Wang , Liqiang Nie , Wenjie Li

Tuning-free face personalization methods have developed along two distinct paradigms: text embedding approaches that map facial features into the text embedding space, and adapter-based methods that inject features through auxiliary…

计算机视觉与模式识别 · 计算机科学 2025-12-04 Lianyu Pang , Ji Zhou , Qiping Wang , Baoquan Zhao , Zhenguo Yang , Qing Li , Xudong Mao

Diffusion-based technologies have made significant strides, particularly in personalized and customized facialgeneration. However, existing methods face challenges in achieving high-fidelity and detailed identity (ID)consistency, primarily…

计算机视觉与模式识别 · 计算机科学 2024-12-31 Jiehui Huang , Xiao Dong , Wenhui Song , Zheng Chong , Zhenchao Tang , Jun Zhou , Yuhao Cheng , Long Chen , Hanhui Li , Yiqiang Yan , Shengcai Liao , Xiaodan Liang

In the field of personalized image generation, the ability to create images preserving concepts has significantly improved. Creating an image that naturally integrates multiple concepts in a cohesive and visually appealing composition can…

计算机视觉与模式识别 · 计算机科学 2024-05-01 Chanran Kim , Jeongin Lee , Shichang Joung , Bongmo Kim , Yeul-Min Baek

The paper introduces AniTalker, an innovative framework designed to generate lifelike talking faces from a single portrait. Unlike existing models that primarily focus on verbal cues such as lip synchronization and fail to capture the…

计算机视觉与模式识别 · 计算机科学 2024-05-07 Tao Liu , Feilong Chen , Shuai Fan , Chenpeng Du , Qi Chen , Xie Chen , Kai Yu

$360^{\circ}$ omnidirectional images (ODIs) have gained considerable attention recently, and are widely used in various virtual reality (VR) and augmented reality (AR) applications. However, capturing such images is expensive and requires…

计算机视觉与模式识别 · 计算机科学 2025-08-22 Liu Yang , Huiyu Duan , Yucheng Zhu , Xiaohong Liu , Lu Liu , Zitong Xu , Guangji Ma , Xiongkuo Min , Guangtao Zhai , Patrick Le Callet

We present Concat-ID, a unified framework for identity-preserving video generation. Concat-ID employs variational autoencoders to extract image features, which are then concatenated with video latents along the sequence dimension. It relies…

计算机视觉与模式识别 · 计算机科学 2025-07-03 Yong Zhong , Zhuoyi Yang , Jiayan Teng , Xiaotao Gu , Chongxuan Li

The creation of 3D human face avatars from a single unconstrained image is a fundamental task that underlies numerous real-world vision and graphics applications. Despite the significant progress made in generative models, existing methods…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Wenqing Wang , Haosen Yang , Josef Kittler , Xiatian Zhu

The visuomotor policy can easily overfit to its training datasets, such as fixed camera positions and backgrounds. This overfitting makes the policy perform well in the in-distribution scenarios but underperform in the out-of-distribution…

机器人学 · 计算机科学 2025-08-19 Jilei Mao , Jiarui Guan , Yingjuan Tang , Qirui Hu , Zhihang Li , Junjie Yu , Yongjie Mao , Yunzhe Sun , Shuang Liu , Xiaozhu Ju

Recent advances have demonstrated compelling capabilities in synthesizing real individuals into generated videos, reflecting the growing demand for identity-aware content creation. Nevertheless, an openly accessible framework enabling…

计算机视觉与模式识别 · 计算机科学 2026-03-26 Yingjie Chen , Shilun Lin , Cai Xing , Binxin Yang , Long Zhou , Qixin Yan , Wenjing Wang , Dingming Liu , Hao Liu , Chen Li , Jing Lyu

Many vision applications require identity consistency beyond strict biometric recognition, especially under non-frontal views or when facial cues are missing. However, conventional face recognition models enforce intra-identity invariance,…

计算机视觉与模式识别 · 计算机科学 2026-05-11 Yingfeng Wang , Yuxuan Xiao , Shengcai Liao

Humans inherently possess generalizable visual representations that empower them to efficiently explore and interact with the environments in manipulation tasks. We advocate that such a representation automatically arises from…

机器人学 · 计算机科学 2023-10-05 Mingxiao Huo , Mingyu Ding , Chenfeng Xu , Thomas Tian , Xinghao Zhu , Yao Mu , Lingfeng Sun , Masayoshi Tomizuka , Wei Zhan

While convenient in daily life, face recognition technologies also raise privacy concerns for regular users on the social media since they could be used to analyze face images and videos, efficiently and surreptitiously without any security…

计算机视觉与模式识别 · 计算机科学 2022-05-25 Yaoyao Zhong , Weihong Deng