中文
相关论文

相关论文: ID-LoRA: Identity-Driven Audio-Video Personalizati…

200 篇论文

Recent years have witnessed great progress in creating vivid audio-driven portraits from monocular videos. However, how to seamlessly adapt the created video avatars to other scenarios with different backgrounds and lighting conditions…

计算机视觉与模式识别 · 计算机科学 2023-09-06 Haonan Qiu , Zhaoxi Chen , Yuming Jiang , Hang Zhou , Xiangyu Fan , Lei Yang , Wayne Wu , Ziwei Liu

Diffusion models have become a powerful backbone for text-to-image generation, producing high-quality visuals from natural language prompts. However, when prompts involve multiple objects alongside global or local style instructions, the…

计算机视觉与模式识别 · 计算机科学 2025-08-19 Ankit Sanjyal

Low-rank adaptation (LoRA) is widely used for parameter-efficient fine-tuning, but its standard all-token, all-head design ignores the heterogeneous structure of vision language model (VLM) inputs. We introduce \emph{Image-LoRA}, a…

计算机视觉与模式识别 · 计算机科学 2026-05-12 Tiange Luo , Lajanugen Logeswaran , Jaekyeom Kim , Justin Johnson , Honglak Lee

Low-Rank Adaptation (LoRA) has emerged as a powerful and popular technique for personalization, enabling efficient adaptation of pre-trained image generation models for specific tasks without comprehensive retraining. While employing…

计算机视觉与模式识别 · 计算机科学 2025-07-31 Tuna Han Salih Meral , Enis Simsar , Federico Tombari , Pinar Yanardag

This paper introduces V2A-DPO, a novel Direct Preference Optimization (DPO) framework tailored for flow-based video-to-audio generation (V2A) models, incorporating key adaptations to effectively align generated audio with human preferences.…

声音 · 计算机科学 2026-03-13 Nolan Chan , Timmy Gang , Yongqian Wang , Yuzhe Liang , Dingdong Wang

Recent advancements in image generation models have enabled personalized image creation with both user-defined subjects (content) and styles. Prior works achieved personalization by merging corresponding low-rank adapters (LoRAs) through…

计算机视觉与模式识别 · 计算机科学 2025-08-12 Donald Shenaj , Ondrej Bohdal , Mete Ozay , Pietro Zanuttigh , Umberto Michieli

Human listeners readily adjust to unfamiliar speakers and language varieties through exposure, but do these adaptation benefits extend to state-of-the-art spoken language models? We introduce a scalable framework that allows for in-context…

计算与语言 · 计算机科学 2025-05-22 Nathan Roll , Calbert Graham , Yuka Tatsumi , Kim Tien Nguyen , Meghan Sumner , Dan Jurafsky

Vision-guided speech generation aims to produce authentic speech from facial appearance or lip motions without relying on auditory signals, offering significant potential for applications such as dubbing in filmmaking and assisting…

计算机视觉与模式识别 · 计算机科学 2025-03-20 Jiaxin Ye , Hongming Shan

Recent advancements in video-audio joint generation have achieved remarkable success in semantic correspondence. However, achieving precise temporal synchronization, which requires fine-grained alignment between audio events and their…

计算机视觉与模式识别 · 计算机科学 2026-05-13 Xin Cheng , Xihua Wang , Ying Ba , Yuyue Wang , Kaisi Guan , Yinbo Wang , Wenpu Li , Ruihua Song

This work introduces a new task, text-conditioned selective video-to-audio (V2A) generation, which produces only the user-intended sound from a multi-object video. This capability is especially crucial in multimedia production, where audio…

计算机视觉与模式识别 · 计算机科学 2026-03-30 Junwon Lee , Juhan Nam , Jiyoung Lee

Text-to-audio (TTA) generation is a recent popular problem that aims to synthesize general audio given text descriptions. Previous methods utilized latent diffusion models to learn audio embedding in a latent space with text embedding as…

计算机视觉与模式识别 · 计算机科学 2023-05-23 Shentong Mo , Jing Shi , Yapeng Tian

Conversational AI has made significant progress, yet generating expressive and controllable text-to-speech (TTS) remains challenging. Specifically, controlling fine-grained voice styles and emotions is notoriously difficult and typically…

音频与语音处理 · 电气工程与系统科学 2026-04-13 Zhicheng Ouyang , Seong-Gyun Leem , Bach Viet Do , Haibin Wu , Ariya Rastrow , Yuzong Liu , Florian Metze

Large language models are increasingly adopted as semantic backbones for neural text-to-speech systems. However, frozen LLM representations are insufficient for modeling speaker specific acoustic and perceptual characteristics. Our…

声音 · 计算机科学 2026-03-12 Anupam Purwar , Aditya Choudhary

Identity-preserving text-to-video (IPT2V) generation creates videos faithful to both a reference subject image and a text prompt. While fine-tuning large pretrained video diffusion models on ID-matched data achieves state-of-the-art results…

计算机视觉与模式识别 · 计算机科学 2025-09-03 Jiayi Gao , Changcheng Hua , Qingchao Chen , Yuxin Peng , Yang Liu

The rapid development of diffusion models has triggered diverse applications. Identity-preserving text-to-image generation (ID-T2I) particularly has received significant attention due to its wide range of application scenarios like AI…

计算机视觉与模式识别 · 计算机科学 2024-04-25 Weifeng Chen , Jiacheng Zhang , Jie Wu , Hefeng Wu , Xuefeng Xiao , Liang Lin

Story visualization requires generating sequential imagery that aligns semantically with evolving narratives while maintaining rigorous consistency in character identity and visual style. However, existing methodologies often struggle with…

计算机视觉与模式识别 · 计算机科学 2026-03-19 Jianzhang Zhang , Yijing Tian , Jiwang Qu , Chuang Liu

While state-of-the-art audio-video generation models like Veo3 and Sora2 demonstrate remarkable capabilities, their closed-source nature makes their architectures and training paradigms inaccessible. To bridge this gap in accessibility and…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Hebeizi Li , Zihao Liang , Benyuan Sun , Zihao Yin , Xiao Sha , Chenliang Wang , Yi Yang

We present Lynx, a high-fidelity model for personalized video synthesis from a single input image. Built on an open-source Diffusion Transformer (DiT) foundation model, Lynx introduces two lightweight adapters to ensure identity fidelity.…

计算机视觉与模式识别 · 计算机科学 2025-09-22 Shen Sang , Tiancheng Zhi , Tianpei Gu , Jing Liu , Linjie Luo

Single-view reference-to-video methods often struggle to preserve identity consistency under large facial-angle variations. This limitation naturally motivates the incorporation of multi-view facial references. However, simply introducing…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Bin Hu , Zipeng Qi , Guoxi Huang , Zunnan Xu , Ruicheng Zhang , Chongjie Ye , Jun Zhou , Xiu Li , Jingdong Wang

We primarily focus on the field of large language models (LLMs) for recommendation, which has been actively explored recently and poses a significant challenge in effectively enhancing recommender systems with logical reasoning abilities…

信息检索 · 计算机科学 2024-08-13 Jiachen Zhu , Jianghao Lin , Xinyi Dai , Bo Chen , Rong Shan , Jieming Zhu , Ruiming Tang , Yong Yu , Weinan Zhang