中文
相关论文

相关论文: Make-it-Real: Unleashing Large Multimodal Model fo…

200 篇论文

Automatically evaluating vision-language tasks is challenging, especially when it comes to reflecting human judgments due to limitations in accounting for fine-grained details. Although GPT-4V has shown promising results in various…

计算机视觉与模式识别 · 计算机科学 2023-11-03 Xinlu Zhang , Yujie Lu , Weizhi Wang , An Yan , Jun Yan , Lianke Qin , Heng Wang , Xifeng Yan , William Yang Wang , Linda Ruth Petzold

Generative AI tools are becoming more prevalent in 3D modeling, enabling users to manipulate or create new models with text or images as inputs. This makes it easier for users to rapidly customize and iterate on their 3D designs and explore…

人机交互 · 计算机科学 2024-04-18 Faraz Faruqi , Yingtao Tian , Vrushank Phadnis , Varun Jampani , Stefanie Mueller

We are witnessing a proliferation of textured 3D models captured from the real world with automatic photo-reconstruction tools. Digital 3D models of this class come with a unique set of characteristics and defects -- especially concerning…

图形学 · 计算机科学 2020-12-29 Andrea Maggiordomo , Federico Ponchio , Paolo Cignoni , Marco Tarini

Image-based 3D object modeling refers to the process of converting raw optical images to 3D digital representations of the objects. Very often, such models are desired to be dimensionally true, semantically labeled with photorealistic…

计算机视觉与模式识别 · 计算机科学 2021-06-29 Rongjun Qin , Xu Huang

BRDF models are ubiquitous tools for the representation of material appearance. However, there is now an astonishingly large number of different models in practical use. Both a lack of BRDF model standardisation across implementations found…

图形学 · 计算机科学 2018-08-22 Alejandro Sztrajman , Jaroslav Krivanek , Alexander Wilkie , Tim Weyrich

We introduce the novel task of Language-Guided Object Placement in Real 3D Scenes. Our model is given a 3D scene's point cloud, a 3D asset, and a textual prompt broadly describing where the 3D asset should be placed. The task here is to…

Advancements in text-to-image diffusion models have led to significant progress in fast 3D content creation. One common approach is to generate a set of multi-view images of an object, and then reconstruct it into a 3D model. However, this…

计算机视觉与模式识别 · 计算机科学 2024-12-04 Yiftach Edelstein , Or Patashnik , Dana Cohen-Bar , Lihi Zelnik-Manor

Transformer neural networks show promising capabilities, in particular for uses in materials analysis, design and manufacturing, including their capacity to work effectively with both human language, symbols, code, and numerical data. Here…

计算与语言 · 计算机科学 2023-11-01 Markus J. Buehler

Recent advances in deep learning, such as neural radiance fields and implicit neural representations, have significantly advanced 3D reconstruction. However, accurately reconstructing objects with complex optical properties, such as metals,…

计算机视觉与模式识别 · 计算机科学 2025-11-04 Zheng Dang , Jialu Huang , Fei Wang , Mathieu Salzmann

Recent advancements in 3D generation models have opened new possibilities for simulating dynamic 3D object movements and customizing behaviors, yet creating this content remains challenging. Current methods often require manual assignment…

计算机视觉与模式识别 · 计算机科学 2025-07-29 Haoyu Zhao , Hao Wang , Xingyue Zhao , Hao Fei , Hongqiu Wang , Chengjiang Long , Hua Zou

The remarkable advancements in Multimodal Large Language Models (MLLMs) have not rendered them immune to challenges, particularly in the context of handling deceptive information in prompts, thus producing hallucinated responses under such…

计算机视觉与模式识别 · 计算机科学 2024-07-24 Yusu Qian , Haotian Zhang , Yinfei Yang , Zhe Gan

In this paper, we investigate the use of multimodal large language models (MLLMs) for generating virtual activities, leveraging the integration of vision-language modalities to enable the interpretation of virtual environments. Our approach…

人机交互 · 计算机科学 2025-11-13 Changyang Li , Qingan Yan , Minyoung Kim , Zhan Li , Yi Xu , Lap-Fai Yu

Creating photorealistic materials for light transport algorithms requires carefully fine-tuning a set of material properties to achieve a desired artistic effect. This is typically a lengthy process that involves a trained artist with…

图形学 · 计算机科学 2019-09-26 Károly Zsolnai-Fehér , Peter Wonka , Michael Wimmer

Engineering educational curriculum and standards cover many material and manufacturing options. However, engineers and designers are often unfamiliar with certain composite materials or manufacturing techniques. Large language models (LLMs)…

The generation of high-quality 3D car assets is essential for various applications, including video games, autonomous driving, and virtual reality. Current 3D generation methods utilizing NeRF or 3D-GS as representations for 3D objects,…

计算机视觉与模式识别 · 计算机科学 2024-10-11 Xiaoxue Chen , Jv Zheng , Hao Huang , Haoran Xu , Weihao Gu , Kangliang Chen , He xiang , Huan-ang Gao , Hao Zhao , Guyue Zhou , Yaqin Zhang

The advancement of Large Language Models (LLMs), including GPT-4, provides exciting new opportunities for generative design. We investigate the application of this tool across the entire design and manufacturing workflow. Specifically, we…

Two-dimensional (2D) materials have been a central focus of recent research because they host a variety of properties, making them attractive both for fundamental science and for applications. It is thus crucial to be able to identify…

材料科学 · 物理学 2022-11-18 Mohammad Tohidi Vahdat , Kumar Agrawal Varoon , Giovanni Pizzi

High-throughput data generation methods and machine learning (ML) algorithms have given rise to a new era of computational materials science by learning relationships among composition, structure, and properties and by exploiting such…

The success of large language models (LLMs) has inspired an emerging research field of multimodal learning. However, a grand challenge of exploiting LLMs for multimodal learning is the size of pre-trained LLMs which are always with billions…

计算与语言 · 计算机科学 2024-04-08 Zhengqing Yuan , Yunhong He , Kun Wang , Yanfang Ye , Lichao Sun

Multimodal large language models (MLLMs) have advanced the capabilities to interpret and act on visual input in 3D environments, empowering diverse applications such as robotics and situated conversational agents. When MLLMs reason over…

计算机视觉与模式识别 · 计算机科学 2026-02-10 Zhuoheng Li , Ying Chen