中文
相关论文

相关论文: Gemini2: Generating Keyframe-Oriented Animated Tra…

200 篇论文

Cross-modal transformers have demonstrated superiority in various vision tasks by effectively integrating different modalities. This paper first critiques prior token exchange methods which replace less informative tokens with inter-modal…

计算机视觉与模式识别 · 计算机科学 2024-06-05 Ding Jia , Jianyuan Guo , Kai Han , Han Wu , Chao Zhang , Chang Xu , Xinghao Chen

Generative AI assistants have been widely used in front-end programming. However, besides code writing, developers often encounter the need to generate animation effects. As novices in creative design without the assistance of professional…

人机交互 · 计算机科学 2025-06-30 Tianrun Qiu , Yuxin Ma

In this pilot project, we teamed up with artists to develop new workflows for 2D animation while producing a short educational cartoon. We identified several workflows to streamline the animation process, bringing the artists' vision to the…

图形学 · 计算机科学 2024-05-24 Jaime Guajardo , Ozgun Bursalioglu , Dan B Goldman

As robots increasingly enter human-centered environments, they must not only be able to navigate safely around humans, but also adhere to complex social norms. Humans often rely on non-verbal communication through gestures and facial…

We present Gradient Gating (G$^2$), a novel framework for improving the performance of Graph Neural Networks (GNNs). Our framework is based on gating the output of GNN layers with a mechanism for multi-rate flow of message passing…

Generation of scientific visualization from analytical natural language text is a challenging task. In this paper, we propose Text2Chart, a multi-staged chart generator method. Text2Chart takes natural language text as input and produce…

Generative inbetweening (GI) seeks to synthesize realistic intermediate frames between the first and last keyframes beyond mere interpolation. As sequences become sparser and motions larger, previous GI models struggle with inconsistent…

计算机视觉与模式识别 · 计算机科学 2026-03-19 Tae Eun Choi , Sumin Shim , Junhyeok Kim , Seong Jae Hwang

Animation is ubiquitous in visualization systems, and a common technique for creating these animations is the transition. In the transition approach, animations are created by smoothly interpolating a visual attribute between a start and…

图形学 · 计算机科学 2017-03-03 Andrew McCaleb Reach , Chris North

Accurate future video prediction requires both high visual fidelity and consistent scene semantics, particularly in complex dynamic environments such as autonomous driving. We present Re2Pix, a hierarchical video prediction framework that…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Efstathios Karypidis , Spyros Gidaris , Nikos Komodakis

Instant-messaging human social chat typically progresses through a sequence of short messages. Existing step-by-step AI chatting systems typically split a one-shot generation into multiple messages and send them sequentially, but they lack…

计算与语言 · 计算机科学 2026-01-12 Hao Yang , Hongyuan Lu , Dingkang Yang , Wenliang Yang , Peng Sun , Xiaochuan Zhang , Jun Xiao , Kefan He , Wai Lam , Yang Liu , Xinhua Zeng

Recent multimodal image generators such as GPT-4o, Gemini 2.0 Flash, and Gemini 2.5 Pro excel at following complex instructions, editing images and maintaining concept consistency. However, they are still evaluated by disjoint toolkits:…

计算机视觉与模式识别 · 计算机科学 2025-05-29 Hang Hua , Ziyun Zeng , Yizhi Song , Yunlong Tang , Liu He , Daniel Aliaga , Wei Xiong , Jiebo Luo

By generating plausible and smooth transitions between two image frames, video inbetweening is an essential tool for video editing and long video synthesis. Traditional works lack the capability to generate complex large motions. While…

计算机视觉与模式识别 · 计算机科学 2025-01-09 Maham Tanveer , Yang Zhou , Simon Niklaus , Ali Mahdavi Amiri , Hao Zhang , Krishna Kumar Singh , Nanxuan Zhao

We introduce the GANformer2 model, an iterative object-oriented transformer, explored for the task of generative modeling. The network incorporates strong and explicit structural priors, to reflect the compositional nature of visual scenes,…

计算机视觉与模式识别 · 计算机科学 2021-11-18 Drew A. Hudson , C. Lawrence Zitnick

Current personalized recommender systems predominantly rely on static offline data for algorithm design and evaluation, significantly limiting their ability to capture long-term user preference evolution and social influence dynamics in…

多智能体系统 · 计算机科学 2025-05-28 Hailin Zhong , Hanlin Wang , Yujun Ye , Meiyi Zhang , Shengxin Zhu

Data analysts often need to iterate between data transformations and chart designs to create rich visualizations for exploratory data analysis. Although many AI-powered systems have been introduced to reduce the effort of visualization…

人机交互 · 计算机科学 2025-02-24 Chenglong Wang , Bongshin Lee , Steven Drucker , Dan Marshall , Jianfeng Gao

Layout design, such as user interface or graphical layout in general, is fundamentally an iterative revision process. Through revising a design repeatedly, the designer converges on an ideal layout. In this paper, we investigate how…

人机交互 · 计算机科学 2024-06-28 Tao Li , Chin-Yi Cheng , Amber Xie , Gang Li , Yang Li

In this paper we present a new deep learning-driven approach to image-based synthesis of animations involving humanoid characters. Unlike previous deep approaches to image-based animation our method makes no assumptions on the type of…

图形学 · 计算机科学 2019-08-14 John Kanji , David I. W. Levin

Game Design Pillars are natural language artifacts commonly used in game development to communicate a project's core vision and ensure a coherent player experience. Their linguistic nature aligns well with the strengths of Large Language…

人机交互 · 计算机科学 2026-05-12 Julian Geheeb , Marvin Julian Schwarz , Daniel Dyrda , Georg Groh

While generative AI enables high-fidelity UI generation from text prompts, users struggle to articulate design intent and evaluate or refine results-creating gulfs of execution and evaluation. To understand the information needed for UI…

人机交互 · 计算机科学 2026-02-10 Seokhyeon Park , Soohyun Lee , Eugene Choi , Hyunwoo Kim , Minkyu Kweon , Yumin Song , Jinwook Seo

It is essential to help drivers have appropriate understandings of level 2 automated driving systems for keeping driving safety. A human machine interface (HMI) was proposed to present real time results of image recognition by the automated…

人机交互 · 计算机科学 2021-06-28 Bo Yang , Koichiro Inoue , Satoshi Kitazaki , Kimihiko Nakano