中文
相关论文

相关论文: Shape2Animal: Creative Animal Generation from Natu…

200 篇论文

We present a new implicit warping framework for image animation using sets of source images through the transfer of the motion of a driving video. A single cross- modal attention layer is used to find correspondences between the source…

计算机视觉与模式识别 · 计算机科学 2022-10-05 Arun Mallya , Ting-Chun Wang , Ming-Yu Liu

As robots increasingly enter human-centered environments, they must not only be able to navigate safely around humans, but also adhere to complex social norms. Humans often rely on non-verbal communication through gestures and facial…

With the rise of AI-generated content (AIGC), generating perceptually natural and feeling-aligned music from multimodal inputs has become a central challenge. Existing approaches often rely on explicit emotion labels that require costly…

声音 · 计算机科学 2025-12-02 Jiaying Hong , Ting Zhu , Thanet Markchom , Huizhi Liang

We propose a novel deep learning framework for animation video resequencing. Our system produces new video sequences by minimizing a perceptual distance of images from an existing animation video clip. To measure perceptual distance, we…

图形学 · 计算机科学 2021-11-03 Charles C. Morace , Thi-Ngoc-Hanh Le , Sheng-Yi Yao , Shang-Wei Zhang , Tong-Yee Lee

Humans interact with the environment using a combination of perception - transforming sensory inputs from their environment into symbols, and cognition - mapping symbols to knowledge about the environment for supporting abstraction,…

人工智能 · 计算机科学 2023-11-07 Amit Sheth , Kaushik Roy , Manas Gaur

Recent breakthroughs in generative AI have opened the door to new research perspectives in the domain of art and cultural heritage, where a large number of artifacts have been digitized. There is a need for innovation to ease the access and…

计算机视觉与模式识别 · 计算机科学 2025-06-09 Valentine Bernasconi , Gustavo Marfia

Sketch2Prototype is an AI-based framework that transforms a hand-drawn sketch into a diverse set of 2D images and 3D prototypes through sketch-to-text, text-to-image, and image-to-3D stages. This framework, shown across various sketches,…

人机交互 · 计算机科学 2024-05-24 Kristen M. Edwards , Brandon Man , Faez Ahmed

Generating coherent and useful image/video scenes from a free-form textual description is technically a very difficult problem to handle. Textual description of the same scene can vary greatly from person to person, or sometimes even for…

计算机视觉与模式识别 · 计算机科学 2020-12-01 Faria Huq , Nafees Ahmed , Anindya Iqbal

We propose a new paradigm to automatically generate training data with accurate labels at scale using the text-to-image synthesis frameworks (e.g., DALL-E, Stable Diffusion, etc.). The proposed approach1 decouples training data generation…

计算机视觉与模式识别 · 计算机科学 2023-09-13 Yunhao Ge , Jiashu Xu , Brian Nlong Zhao , Neel Joshi , Laurent Itti , Vibhav Vineet

Personalized content filtering, such as recommender systems, has become a critical infrastructure to alleviate information overload. However, these systems merely filter existing content and are constrained by its limited diversity, making…

信息检索 · 计算机科学 2025-02-05 Yiyan Xu , Wenjie Wang , Yang Zhang , Biao Tang , Peng Yan , Fuli Feng , Xiangnan He

Animal stereotypes are deeply embedded in human culture and language. They often shape our perceptions and expectations of various species. Our study investigates how animal stereotypes manifest in vision-language models during the task of…

计算机视觉与模式识别 · 计算机科学 2025-04-30 Tabinda Aman , Mohammad Nadeem , Shahab Saquib Sohail , Mohammad Anas , Erik Cambria

Recent character image animation methods based on diffusion models, such as Animate Anyone, have made significant progress in generating consistent and generalizable character animations. However, these approaches fail to produce reasonable…

计算机视觉与模式识别 · 计算机科学 2025-02-11 Li Hu , Guangyuan Wang , Zhen Shen , Xin Gao , Dechao Meng , Lian Zhuo , Peng Zhang , Bang Zhang , Liefeng Bo

Object recognition has seen significant progress in the image domain, with focus primarily on 2D perception. We propose to leverage existing large-scale datasets of 3D models to understand the underlying 3D structure of objects seen in an…

计算机视觉与模式识别 · 计算机科学 2020-07-28 Weicheng Kuo , Anelia Angelova , Tsung-Yi Lin , Angela Dai

Text-to-image synthesis (T2I) aims to generate photo-realistic images which are semantically consistent with the text descriptions. Existing methods are usually built upon conditional generative adversarial networks (GANs) and initialize an…

计算机视觉与模式识别 · 计算机科学 2022-03-25 Kai Hu , Wentong Liao , Michael Ying Yang , Bodo Rosenhahn

We introduce ShadowDraw, a framework that transforms ordinary 3D objects into shadow-drawing compositional art. Given a 3D object, our system predicts scene parameters, including object pose and lighting, together with a partial line…

计算机视觉与模式识别 · 计算机科学 2025-12-05 Rundong Luo , Noah Snavely , Wei-Chiu Ma

In this work, we develop intuitive controls for editing the style of 3D objects. Our framework, Text2Mesh, stylizes a 3D mesh by predicting color and local geometric details which conform to a target text prompt. We consider a disentangled…

计算机视觉与模式识别 · 计算机科学 2021-12-07 Oscar Michel , Roi Bar-On , Richard Liu , Sagie Benaim , Rana Hanocka

There has been significant work on learning realistic, articulated, 3D models of the human body. In contrast, there are few such models of animals, despite many applications. The main challenge is that animals are much less cooperative than…

计算机视觉与模式识别 · 计算机科学 2017-04-13 Silvia Zuffi , Angjoo Kanazawa , David Jacobs , Michael J. Black

We present CLIPSwarm, an algorithm to generate robot swarm formations from natural language descriptions. CLIPSwarm receives an input text and finds the position of the robots to form a shape that corresponds to the given text. To do so, we…

机器人学 · 计算机科学 2023-11-21 Pablo Pueyo , Eduardo Montijano , Ana C. Murillo , Mac Schwager

Since the generative neural networks have made a breakthrough in the image generation problem, lots of researches on their applications have been studied such as image restoration, style transfer and image completion. However, there has…

计算机视觉与模式识别 · 计算机科学 2021-10-08 Jeesoo Kim , Jangho Kim , Jaeyoung Yoo , Daesik Kim , Nojun Kwak

Large text-guided diffusion models, such as DALLE-2, are able to generate stunning photorealistic images given natural language descriptions. While such models are highly flexible, they struggle to understand the composition of certain…

计算机视觉与模式识别 · 计算机科学 2023-01-18 Nan Liu , Shuang Li , Yilun Du , Antonio Torralba , Joshua B. Tenenbaum