English
Related papers

Related papers: MagicCraft: Natural Language-Driven Generation of …

200 papers

We present MagicMirror, a framework for generating identity-preserved videos with cinematic-level quality and dynamic motion. While recent advances in video diffusion models have shown impressive capabilities in text-to-video generation,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Yuechen Zhang , Yaoyang Liu , Bin Xia , Bohao Peng , Zexin Yan , Eric Lo , Jiaya Jia

The financial services industry perpetually processes an overwhelming amount of complex data. Digital reports are often created based on tedious manual analysis as well as visualization of the underlying trends and characteristics of data.…

Computation and Language · Computer Science 2021-02-03 Vineeth Ravi , Selim Amrouni , Andrea Stefanucci , Armineh Nourbakhsh , Prashant Reddy , Manuela Veloso

Creating 2D animations is a complex, iterative process requiring continuous adjustments to movement, timing, and coordination of multiple elements within a scene. To support designers of varying levels of experience with animation design…

Human-Computer Interaction · Computer Science 2025-08-14 Tiffany Tseng , Ruijia Cheng , Jeffrey Nichols

Generating videos with realistic and physically plausible motion is one of the main recent challenges in computer vision. While diffusion models are achieving compelling results in image generation, video diffusion models are limited by…

Machine Learning · Computer Science 2024-10-28 Luca Savant Aira , Antonio Montanaro , Emanuele Aiello , Diego Valsesia , Enrico Magli

Generating realistic 3D cities is fundamental to world models, virtual reality, and game development, where an ideal urban scene must satisfy both stylistic diversity, fine-grained, and controllability. However, existing methods struggle to…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Zilong Huang , Jun He , Xiaobin Huang , Ziyi Xiong , Yang Luo , Junyan Ye , Weijia Li , Yiping Chen , Ting Han

Generative AI promises to allow people to create high-quality personalized media. Although powerful, we identify three fundamental design problems with existing tooling through a literature review. We introduce a multimodal generative AI…

Human-Computer Interaction · Computer Science 2025-06-23 Gregory Croisdale , Emily Huang , John Joon Young Chung , Anhong Guo , Xu Wang , Austin Z. Henley , Cyrus Omar

Effective prompting of generative AI is challenging for many users, particularly in expressing context for comprehension tasks such as explaining spreadsheet formulas, Python code, and text passages. Prompt middleware aims to address this…

Human-Computer Interaction · Computer Science 2024-12-04 Ian Drosos , Jack Williams , Advait Sarkar , Nicholas Wilson

Socially interactive agents are gaining prominence in domains like healthcare, education, and service contexts, particularly virtual agents due to their inherent scalability. To facilitate authentic interactions, these systems require…

Human-Computer Interaction · Computer Science 2025-01-22 Oliver Chojnowski , Alexander Eberhard , Michael Schiffmann , Ana Müller , Anja Richert

Realistic object interactions are crucial for creating immersive virtual experiences, yet synthesizing realistic 3D object dynamics in response to novel interactions remains a significant challenge. Unlike unconditional or text-conditioned…

Computer Vision and Pattern Recognition · Computer Science 2024-10-08 Tianyuan Zhang , Hong-Xing Yu , Rundi Wu , Brandon Y. Feng , Changxi Zheng , Noah Snavely , Jiajun Wu , William T. Freeman

Today's video-conferencing tools support a rich range of professional and social activities, but their generic meeting environments cannot be dynamically adapted to align with distributed collaborators' needs. To enable end-user…

Human-Computer Interaction · Computer Science 2024-10-02 Shwetha Rajaram , Nels Numan , Balasaravanan Thoravi Kumaravel , Nicolai Marquardt , Andrew D. Wilson

Cinematographers adeptly capture the essence of the world, crafting compelling visual narratives through intricate camera movements. Witnessing the strides made by large language models in perceiving and interacting with the 3D world, this…

Computer Vision and Pattern Recognition · Computer Science 2024-09-27 Xinhang Liu , Yu-Wing Tai , Chi-Keung Tang

Humans are able to perceive, understand and reason about causal events. Developing models with similar physical and causal understanding capabilities is a long-standing goal of artificial intelligence. As a step towards this direction, we…

Artificial Intelligence · Computer Science 2022-03-02 Tayfun Ates , M. Samil Atesoglu , Cagatay Yigit , Ilker Kesen , Mert Kobas , Erkut Erdem , Aykut Erdem , Tilbe Goksun , Deniz Yuret

Generating and editing a 3D scene guided by natural language poses a challenge, primarily due to the complexity of specifying the positional relations and volumetric changes within the 3D space. Recent advancements in Large Language Models…

Computer Vision and Pattern Recognition · Computer Science 2023-05-26 Yiqi Lin , Hao Wu , Ruichen Wang , Haonan Lu , Xiaodong Lin , Hui Xiong , Lin Wang

We present ShapeCrafter, a neural network for recursive text-conditioned 3D shape generation. Existing methods to generate text-conditioned 3D shapes consume an entire text prompt to generate a 3D shape in a single step. However, humans…

Computer Vision and Pattern Recognition · Computer Science 2023-04-11 Rao Fu , Xiao Zhan , Yiwen Chen , Daniel Ritchie , Srinath Sridhar

Interactive documents help readers engage with complex ideas through dynamic visualization, interactive animations, and exploratory interfaces. However, creating such documents remains costly, as it requires both domain expertise and web…

Human-Computer Interaction · Computer Science 2026-03-31 Yinghao Tang , Yupeng Xie , Yingchaojie Feng , Tingfeng Lan , Jiale Lao , Yue Cheng , Wei Chen

Generative AI in Virtual Reality offers the potential for collaborative object-building, yet challenges remain in aligning AI contributions with user expectations. In particular, users often struggle to understand and collaborate with AI…

Human-Computer Interaction · Computer Science 2025-02-24 Julian Rasch , Julia Töws , Teresa Hirzle , Florian Müller , Martin Schmitz

Designers often encounter friction when animating static SVG graphics, especially when the visual structure does not match the desired level of motion detail. Existing tools typically depend on predefined groupings or require technical…

Human-Computer Interaction · Computer Science 2025-11-11 Jihyeon Park , Jiyoon Myung , Seone Shin , Jungki Son , Joohyung Han

Recent years have witnessed remarkable advances in artificial intelligence generated content(AIGC), with diverse input modalities, e.g., text, image, video, audio and 3D. The 3D is the most close visual modality to real-world 3D environment…

Computer Vision and Pattern Recognition · Computer Science 2024-03-20 Jian Liu , Xiaoshui Huang , Tianyu Huang , Lu Chen , Yuenan Hou , Shixiang Tang , Ziwei Liu , Wanli Ouyang , Wangmeng Zuo , Junjun Jiang , Xianming Liu

Customized video generation aims to generate high-quality videos guided by text prompts and subject's reference images. However, since it is only trained on static images, the fine-tuning process of subject learning disrupts abilities of…

Computer Vision and Pattern Recognition · Computer Science 2024-12-30 Tao Wu , Yong Zhang , Xintao Wang , Xianpan Zhou , Guangcong Zheng , Zhongang Qi , Ying Shan , Xi Li

We introduce \textit{WonderVerse}, a simple but effective framework for generating extendable 3D scenes. Unlike existing methods that rely on iterative depth estimation and image inpainting, often leading to geometric distortions and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Hao Feng , Zhi Zuo , Jia-Hui Pan , Ka-Hei Hui , Qi Dou , Jingyu Hu , Zhengzhe Liu