中文
相关论文

相关论文: Idea23D: Collaborative LMM Agents Enable 3D Model …

200 篇论文

While large language models (LLMs) have been thoroughly evaluated for deductive and inductive reasoning, their proficiency in holistic rule learning in interactive environments remains less explored. We introduce RULEARN, a novel benchmark…

计算与语言 · 计算机科学 2025-10-31 Kaiyu He , Mian Zhang , Shuo Yan , Peilin Wu , Zhiyu Zoey Chen

Significant progress has been made in text-to-video generation through the use of powerful generative models and large-scale internet data. However, substantial challenges remain in precisely controlling individual concepts within the…

计算机视觉与模式识别 · 计算机科学 2024-09-30 Hanxin Zhu , Tianyu He , Anni Tang , Junliang Guo , Zhibo Chen , Jiang Bian

Recent advances in AI-generated content (AIGC) have significantly accelerated image editing techniques, driving increasing demand for diverse and fine-grained edits. Despite these advances, existing image editing methods still face…

计算机视觉与模式识别 · 计算机科学 2025-05-27 Shuyu Wang , Weiqi Li , Qian Wang , Shijie Zhao , Jian Zhang

Generative artificial intelligence (GenAI) can rapidly produce large and diverse volumes of content. This lends to it a quality of creativity which can be empowering in the early stages of design. In seeking to understand how creative ways…

人机交互 · 计算机科学 2024-03-20 Gionnieve Lim , Simon T. Perrault

Artificial Intelligence Generated Content (AIGC) has shown remarkable progress in generating realistic images. However, in this paper, we take a step "backward" and address AIGC for the most rudimentary visual modality of human sketches.…

计算机视觉与模式识别 · 计算机科学 2023-08-29 Zhiyu Qu , Tao Xiang , Yi-Zhe Song

Artificial Intelligence-Generated Content (AIGC) has the potential to transform how people build and interact with virtual environments. Within this paper, we discuss potential benefits but also challenges that AIGC has for the creation of…

人机交互 · 计算机科学 2024-11-01 Jens Grubert , Junlong Chen , Per Ola Kristensson

As Artificial Intelligence Generated Content (AIGC) advances, a variety of methods have been developed to generate text, images, videos, and 3D objects from single or multimodal inputs, contributing efforts to emulate human-like cognitive…

计算机视觉与模式识别 · 计算机科学 2024-08-21 Yiying Yang , Fukun Yin , Jiayuan Fan , Xin Chen , Wanzhang Li , Gang Yu

Recent advances in multimodal large language models (MLLMs) and diffusion models (DMs) have opened new possibilities for AI-generated content. Yet, personalized cover image generation remains underexplored, despite its critical role in…

计算与语言 · 计算机科学 2026-05-28 Zhipeng Bian , Jieming Zhu , Qijiong Liu , Wang Lin , Guohao Cai , Zhaocheng Du , Jiacheng Sun , Zhou Zhao , Zhenhua Dong

We propose a method to fuse frozen text-only large language models (LLMs) with pre-trained image encoder and decoder models, by mapping between their embedding spaces. Our model demonstrates a wide suite of multimodal capabilities: image…

计算与语言 · 计算机科学 2023-10-16 Jing Yu Koh , Daniel Fried , Ruslan Salakhutdinov

Driven by advances in generative artificial intelligence (AI) techniques and algorithms, the widespread adoption of AI-generated content (AIGC) has emerged, allowing for the generation of diverse and high-quality content. Especially, the…

分布式、并行与集群计算 · 计算机科学 2023-12-27 Hongyang Du , Ruichen Zhang , Dusit Niyato , Jiawen Kang , Zehui Xiong , Dong In Kim , Xuemin , Shen , H. Vincent Poor

With the rapid development of artificial intelligence (AI), digital humans have attracted more and more attention and are expected to achieve a wide range of applications in several industries. Then, most of the existing digital humans…

多媒体 · 计算机科学 2023-11-01 Yingjie Zhou , Yaodong Chen , Kaiyue Bi , Lian Xiong , Hui Liu

The evolution of text to visual components facilitates people's daily lives, such as generating image, videos from text and identifying the desired elements within the images. Computer vision models involving the multimodal abilities in the…

计算机视觉与模式识别 · 计算机科学 2024-03-25 Chris Kelly , Luhui Hu , Jiayin Hu , Yu Tian , Deshun Yang , Bang Yang , Cindy Yang , Zihao Li , Zaoshan Huang , Yuexian Zou

Recent strides in Text-to-3D techniques have been propelled by distilling knowledge from powerful large text-to-image diffusion models (LDMs). Nonetheless, existing Text-to-3D approaches often grapple with challenges such as…

计算机视觉与模式识别 · 计算机科学 2023-08-23 Yiwen Chen , Chi Zhang , Xiaofeng Yang , Zhongang Cai , Gang Yu , Lei Yang , Guosheng Lin

This study presents a theory-inspired visual narrative generative system that integrates conceptual principles-comic authoring idioms-with generative and language models to enhance the comic creation process. Our system combines human…

人工智能 · 计算机科学 2024-09-27 Yi-Chun Chen , Arnav Jhala

Recent Multi-Modal Large Language Models (MLLMs) have demonstrated strong capabilities in learning joint representations from text and images. However, their spatial reasoning remains limited. We introduce 3DFroMLLM, a novel framework that…

计算机视觉与模式识别 · 计算机科学 2025-08-13 Noor Ahmed , Cameron Braunstein , Steffen Eger , Eddy Ilg

While 3D content generation has advanced significantly, existing methods still face challenges with input formats, latent space design, and output representations. This paper introduces a novel 3D generation framework that addresses these…

计算机视觉与模式识别 · 计算机科学 2025-04-11 Yushi Lan , Shangchen Zhou , Zhaoyang Lyu , Fangzhou Hong , Shuai Yang , Bo Dai , Xingang Pan , Chen Change Loy

We present Thinking with Generated Images, a novel paradigm that fundamentally transforms how large multimodal models (LMMs) engage with visual reasoning by enabling them to natively think across text and vision modalities through…

计算机视觉与模式识别 · 计算机科学 2025-05-29 Ethan Chern , Zhulin Hu , Steffi Chern , Siqi Kou , Jiadi Su , Yan Ma , Zhijie Deng , Pengfei Liu

Text-to-3D generation has attracted much attention from the computer vision community. Existing methods mainly optimize a neural field from scratch for each text prompt, relying on heavy and repetitive training cost which impedes their…

计算机视觉与模式识别 · 计算机科学 2024-04-30 Ming Li , Pan Zhou , Jia-Wei Liu , Jussi Keppo , Min Lin , Shuicheng Yan , Xiangyu Xu

With the remarkable success of generative models like ChatGPT, Artificial Intelligence Generated Content (AIGC) is undergoing explosive development. Not limited to text and images, generative models can generate industrial time series data,…

机器学习 · 计算机科学 2025-02-18 Lei Ren , Haiteng Wang , Jinwang Li , Yang Tang , Chunhua Yang

Diffusion models (DMs) have emerged as powerful tools for high-quality content generation, yet their intensive computational requirements for inference pose challenges for resource-constrained edge devices. Cloud-based solutions aid in…

机器学习 · 计算机科学 2025-08-08 Nan Li , Wanting Yang , Marie Siew , Zehui Xiong , Binbin Chen , Shiwen Mao , Kwok-Yan Lam