中文
相关论文

相关论文: Aligning Anime Video Generation with Human Feedbac…

200 篇论文

We present Distributional RewArds for Generative OptimizatioN (DRAGON), a versatile framework for fine-tuning media generation models towards a desired outcome. Compared with traditional reinforcement learning with human feedback (RLHF) or…

声音 · 计算机科学 2025-11-18 Yatong Bai , Jonah Casebeer , Somayeh Sojoudi , Nicholas J. Bryan

We present a versatile model, FaceAnime, for various video generation tasks from still images. Video generation from a single face image is an interesting problem and usually tackled by utilizing Generative Adversarial Networks (GANs) to…

计算机视觉与模式识别 · 计算机科学 2021-06-01 Xiaoguang Tu , Yingtian Zou , Jian Zhao , Wenjie Ai , Jian Dong , Yuan Yao , Zhikang Wang , Guodong Guo , Zhifeng Li , Wei Liu , Jiashi Feng

Human-motion video generation has been a challenging task, primarily due to the difficulty inherent in learning human body movements. While some approaches have attempted to drive human-centric video generation explicitly through pose…

计算机视觉与模式识别 · 计算机科学 2025-04-02 Boyuan Wang , Xiaofeng Wang , Chaojun Ni , Guosheng Zhao , Zhiqin Yang , Zheng Zhu , Muyang Zhang , Yukun Zhou , Xinze Chen , Guan Huang , Lihong Liu , Xingang Wang

High-quality preference data is essential for aligning foundation models with human values through preference learning. However, manual annotation of such data is often time-consuming and costly. Recent methods often adopt a self-rewarding…

In generative modeling, we often wish to produce samples that maximize a user-specified reward such as aesthetic quality or alignment with human preferences, a problem known as \textit{guidance}. Despite their widespread use, existing…

机器学习 · 计算机科学 2026-05-21 Jerry Y. Huang , Justin Lin , Sheel Shah , Kartik Nair , Nicholas M. Boffi

Recent advances in video diffusion models have significantly enhanced text-to-video generation, particularly through alignment tuning using reward models trained on human preferences. While these methods improve visual quality, they can…

计算与语言 · 计算机科学 2026-02-12 Zefan Cai , Haoyi Qiu , Haozhe Zhao , Ke Wan , Jiachen Li , Jiuxiang Gu , Wen Xiao , Nanyun Peng , Junjie Hu

The development of trustworthy conversational information-seeking systems relies on dialogue models that can generate faithful and accurate responses based on relevant knowledge texts. However, two main challenges hinder this task. Firstly,…

计算与语言 · 计算机科学 2023-11-03 Wanyu Du , Yangfeng Ji

Recent advances in video-large language models (Video-LLMs) have led to significant progress in video understanding. Current preference optimization methods often rely on proprietary APIs or human-annotated captions to generate preference…

计算机视觉与模式识别 · 计算机科学 2025-08-12 Yogesh Kulkarni , Pooyan Fazli

AI agents are commonly aligned with "human values" through reinforcement learning from human feedback (RLHF), where a single reward model is learned from aggregated human feedback and used to align an agent's behavior. However, human values…

人工智能 · 计算机科学 2025-06-24 Carter Blair , Kate Larson , Edith Law

We consider the task of generating diverse and realistic videos guided by natural audio samples from a wide variety of semantic classes. For this task, the videos are required to be aligned both globally and temporally with the input audio:…

机器学习 · 计算机科学 2023-09-29 Guy Yariv , Itai Gat , Sagie Benaim , Lior Wolf , Idan Schwartz , Yossi Adi

Generating realistic talking-head videos remains challenging due to persistent issues such as imperfect lip synchronization, unnatural motion, and evaluation metrics that correlate poorly with human perception. We propose FlowPortrait, a…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Weiting Tan , Andy T. Liu , Ming Tu , Xinghua Qu , Philipp Koehn , Lu Lu

Recent studies have demonstrated the efficacy of integrating Group Relative Policy Optimization (GRPO) into flow matching models, particularly for text-to-image and text-to-video generation. However, we find that directly applying these…

计算机视觉与模式识别 · 计算机科学 2026-05-27 Jin Wang , Jianxiang Lu , Guangzheng Xu , Comi Chen , Haoyu Yang , Linqing Wang , Peng Chen , Mingtao Chen , Zhichao Hu , Longhuang Wu , Shuai Shao , Qinglin Lu , Ping Luo

Video generation models are rapidly improving in their ability to synthesize human actions in novel contexts, holding the potential to serve as high-level planners for contextual robot control. To realize this potential, a key research…

机器人学 · 计算机科学 2025-12-12 James Ni , Zekai Wang , Wei Lin , Amir Bar , Yann LeCun , Trevor Darrell , Jitendra Malik , Roei Herzig

Learning from human feedback has gained traction in fields like robotics and natural language processing in recent years. While prior works mostly rely on human feedback in the form of comparisons, language is a preferable modality that…

机器人学 · 计算机科学 2024-10-10 Zhaojing Yang , Miru Jun , Jeremy Tien , Stuart J. Russell , Anca Dragan , Erdem Bıyık

With the rapid development of AI-generated content (AIGC), the creation of high-quality AI-generated videos has become faster and easier, resulting in the Internet being flooded with all kinds of video content. However, the impact of these…

信息检索 · 计算机科学 2025-07-30 Haowen Gao , Liang Pang , Shicheng Xu , Leigang Qu , Tat-Seng Chua , Huawei Shen , Xueqi Cheng

Designing dense rewards is crucial for reinforcement learning (RL), yet in robotics it often demands extensive manual effort and lacks scalability. One promising solution is to view task progress as a dense reward signal, as it quantifies…

人工智能 · 计算机科学 2026-05-21 Yuyang Liu , Chuan Wen , Yihang Hu , Dinesh Jayaraman , Yang Gao

We present CHARM, a novel parametric representation and generative framework for anime hairstyle modeling. While traditional hair modeling methods focus on realistic hair using strand-based or volumetric representations, anime hairstyle…

图形学 · 计算机科学 2025-09-26 Yuze He , Yanning Zhou , Wang Zhao , Jingwen Ye , Yushi Bai , Kaiwen Xiao , Yong-Jin Liu , Zhongqian Sun , Wei Yang

Human motion copy is an intriguing yet challenging task in artificial intelligence and computer vision, which strives to generate a fake video of a target person performing the motion of a source person. The problem is inherently…

计算机视觉与模式识别 · 计算机科学 2024-06-25 Sifan Wu , Zhenguang Liu , Beibei Zhang , Roger Zimmermann , Zhongjie Ba , Xiaosong Zhang , Kui Ren

Recent advances in video reward models and post-training strategies have improved text-to-video (T2V) generation. While these models typically assess visual quality, motion quality, and text alignment, they often overlook key structural…

计算机视觉与模式识别 · 计算机科学 2026-03-27 Yuan Wang , Borui Liao , Huijuan Huang , Jinda Lu , Ouxiang Li , Kuien Liu , Meng Wang , Xiang Wang

Text-to-video models have made remarkable advancements through optimization on high-quality text-video pairs, where the textual prompts play a pivotal role in determining quality of output videos. However, achieving the desired output often…

计算机视觉与模式识别 · 计算机科学 2024-12-20 Yatai Ji , Jiacheng Zhang , Jie Wu , Shilong Zhang , Shoufa Chen , Chongjian GE , Peize Sun , Weifeng Chen , Wenqi Shao , Xuefeng Xiao , Weilin Huang , Ping Luo
‹ 上一页 1 8 9 10 下一页 ›