中文
相关论文

相关论文: Idea23D: Collaborative LMM Agents Enable 3D Model …

200 篇论文

This work investigates a challenging task named open-domain interleaved image-text generation, which generates interleaved texts and images following an input query. We propose a new interleaved generation framework based on prompting…

计算机视觉与模式识别 · 计算机科学 2023-11-07 Jie An , Zhengyuan Yang , Linjie Li , Jianfeng Wang , Kevin Lin , Zicheng Liu , Lijuan Wang , Jiebo Luo

3D-consistent image generation from a single 2D semantic label is an important and challenging research topic in computer graphics and computer vision. Although some related works have made great progress in this field, most of the existing…

计算机视觉与模式识别 · 计算机科学 2024-03-12 Bo Li , Yi-ke Li , Zhi-fen He , Bin Liu , Yun-Kun Lai

The great success of Large Language Models (LLMs) has expanded the potential of multimodality, contributing to the gradual evolution of General Artificial Intelligence (AGI). A true AGI agent should not only possess the capability to…

计算机视觉与模式识别 · 计算机科学 2023-10-03 Yuying Ge , Sijie Zhao , Ziyun Zeng , Yixiao Ge , Chen Li , Xintao Wang , Ying Shan

3D content creation has long been a complex and time-consuming process, often requiring specialized skills and resources. While recent advancements have allowed for text-guided 3D object and scene generation, they still fall short of…

计算机视觉与模式识别 · 计算机科学 2024-08-06 Xingyi Li , Yizheng Wu , Jun Cen , Juewen Peng , Kewei Wang , Ke Xian , Zhe Wang , Zhiguo Cao , Guosheng Lin

This paper presents Matrix, an advanced AI-powered framework designed for real-time 3D object generation in Augmented Reality (AR) environments. By integrating a cutting-edge text-to-3D generative AI model, multilingual speech-to-text…

人机交互 · 计算机科学 2025-03-24 Majid Behravan , Denis Gracanin

Interleaved multimodal generation enables capabilities beyond unimodal generation models, such as step-by-step instructional guides, visual planning, and generating visual drafts for reasoning. However, the quality of existing interleaved…

Effective ideation requires both broad exploration of diverse ideas and deep evaluation of their potential. Generative AI can support such processes, but current tools typically emphasize either generating many ideas or supporting in-depth…

人机交互 · 计算机科学 2025-10-01 Yaqing Yang , Vikram Mohanty , Yan-Ying Chen , Matthew K. Hong , Nikolas Martelaro , Aniket Kittur

Computer-Aided Design (CAD) is widely used for conceptual design and parametric 3D modeling, but typically requires a high level of expertise from designers. To lower the entry barrier and facilitate early-stage CAD modeling, we present…

人工智能 · 计算机科学 2026-05-20 Fengxiao Fan , Jingzhe Ni , Xiaolong Yin , Sirui Wang , Xingyu Lu , Qiang Zou , Ruofeng Tong , Min Tang , Peng Du

While recent works have achieved great success on image-to-3D object generation, high quality and fidelity 3D head generation from a single image remains a great challenge. Previous text-based methods for generating 3D heads were limited by…

计算机视觉与模式识别 · 计算机科学 2024-12-24 Jinkun Hao , Junshu Tang , Jiangning Zhang , Ran Yi , Yijia Hong , Moran Li , Weijian Cao , Yating Wang , Chengjie Wang , Lizhuang Ma

The multifaceted nature of human perception and comprehension indicates that, when we think, our body can naturally take any combination of senses, a.k.a., modalities and form a beautiful picture in our brain. For example, when we see a…

计算机视觉与模式识别 · 计算机科学 2024-02-01 Yuanhuiyi Lyu , Xu Zheng , Lin Wang

As the Metaverse continues to grow, the need for efficient communication and intelligent content generation becomes increasingly important. Semantic communication focuses on conveying meaning and understanding from user inputs, while…

人机交互 · 计算机科学 2023-07-25 Yijing Lin , Zhipeng Gao , Hongyang Du , Dusit Niyato , Jiawen Kang , Abbas Jamalipour , Xuemin Sherman Shen

Recent advances in artificial intelligence (AI), coupled with a surge in training data, have led to the widespread use of AI for digital content generation, with ChatGPT serving as a representative example. Despite the increased efficiency…

人工智能 · 计算机科学 2023-03-28 Jiacheng Wang , Hongyang Du , Dusit Niyato , Zehui Xiong , Jiawen Kang , Shiwen Mao , Xuemin , Shen

Design inspiration is crucial for establishing the direction of a design as well as evoking feelings and conveying meanings during the conceptual design process. Many practice designers use text-based searches on platforms like Pinterest to…

人机交互 · 计算机科学 2024-07-18 Ye Wang , Nicole B. Damen , Thomas Gale , Voho Seo , Hooman Shayani

Text- or image-to-3D generators and 3D scanners can now produce 3D assets with high-quality shapes and textures. These assets typically consist of a single, fused representation, like an implicit neural field, a Gaussian mixture, or a mesh,…

计算机视觉与模式识别 · 计算机科学 2024-12-31 Minghao Chen , Roman Shapovalov , Iro Laina , Tom Monnier , Jianyuan Wang , David Novotny , Andrea Vedaldi

Recent advances in 4D generation mainly focus on generating 4D content by distilling pre-trained text or single-view image-conditioned models. It is inconvenient for them to take advantage of various off-the-shelf 3D assets with multi-view…

计算机视觉与模式识别 · 计算机科学 2024-09-10 Yanqin Jiang , Chaohui Yu , Chenjie Cao , Fan Wang , Weiming Hu , Jin Gao

Currently, dialogue systems have achieved high performance in processing text-based communication. However, they have not yet effectively incorporated visual information, which poses a significant challenge. Furthermore, existing models…

计算与语言 · 计算机科学 2023-12-19 Viktor Moskvoretskii , Anton Frolov , Denis Kuznetsov

Automatic furniture layout is long desired for convenient interior design. Leveraging the remarkable visual reasoning capabilities of multimodal large language models (MLLMs), recent methods address layout generation in a static manner,…

计算机视觉与模式识别 · 计算机科学 2024-08-01 Can Wang , Hongliang Zhong , Menglei Chai , Mingming He , Dongdong Chen , Jing Liao

Generative models have achieved success in producing semantically plausible 2D images, but it remains challenging in 3D generation due to the absence of spatial geometry constraints. Typically, existing methods utilize geometric features as…

计算机视觉与模式识别 · 计算机科学 2026-03-09 Haonan Wang , Hanyu Zhou , Haoyue Liu , Tao Gu , Luxin Yan

Large Language Models are increasingly deployed for decision-making, yet their adoption in high-stakes domains remains limited by miscalibrated probabilities, unfaithful explanations, and inability to incorporate expert knowledge precisely.…

人工智能 · 计算机科学 2026-04-15 Yanji He , Yuxin Jiang , Yiwen Wu , Bo Huang , Jiaheng Wei , Wei Wang

Large Language Model Multi-Agent Systems (LLM-MAS) have achieved great progress in solving complex tasks. It performs communication among agents within the system to collaboratively solve tasks, under the premise of shared information.…

人工智能 · 计算机科学 2024-10-18 Wei Liu , Chenxi Wang , Yifei Wang , Zihao Xie , Rennai Qiu , Yufan Dang , Zhuoyun Du , Weize Chen , Cheng Yang , Chen Qian