English
Related papers

Related papers: Make-it-Real: Unleashing Large Multimodal Model fo…

200 papers

Creating and animating 3D biped cartoon characters is crucial and valuable in various applications. Compared with geometry, the diverse texture design plays an important role in making 3D biped cartoon characters vivid and charming.…

Computer Vision and Pattern Recognition · Computer Science 2024-03-26 Junshu Tang , Yanhong Zeng , Ke Fan , Xuheng Wang , Bo Dai , Kai Chen , Lizhuang Ma

Multimodal Large Language Models (MLLMs) have excelled in 2D image-text comprehension and image generation, but their understanding of the 3D world is notably deficient, limiting progress in 3D language understanding and generation. To…

Computer Vision and Pattern Recognition · Computer Science 2025-06-02 Zhangyang Qi , Ye Fang , Zeyi Sun , Xiaoyang Wu , Tong Wu , Jiaqi Wang , Dahua Lin , Hengshuang Zhao

Accurate material retrieval is critical for creating realistic 3D assets. Existing methods rely on datasets that capture shape-invariant and lighting-varied representations of materials, which are scarce and face challenges due to limited…

Computer Vision and Pattern Recognition · Computer Science 2025-04-09 Jianhui Wang , Zhifei Yang , Yangfan He , Huixiong Zhang , Yuxuan Chen , Jingwei Huang

State-of-the-art video generation models produce remarkable photorealism, but they lack the precise control required to align generated content with specific scene requirements. Furthermore, without an underlying explicit geometry, these…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Dana Cohen-Bar , Ido Sobol , Raphael Bensadoun , Shelly Sheynin , Oran Gafni , Or Patashnik , Daniel Cohen-Or , Amit Zohar

The remote sensing image intelligence understanding model is undergoing a new profound paradigm shift which has been promoted by multi-modal large language model (MLLM), i.e. from the paradigm learning a domain model (LaDM) shifts to…

Computer Vision and Pattern Recognition · Computer Science 2024-06-19 Linrui Xu , Ling Zhao , Wang Guo , Qiujun Li , Kewang Long , Kaiqi Zou , Yuhan Wang , Haifeng Li

Large language models (LLMs) have demonstrated a powerful ability to answer various queries as a general-purpose assistant. The continuous multi-modal large language models (MLLM) empower LLMs with the ability to perceive visual signals.…

Computation and Language · Computer Science 2024-01-05 Ziqiang Zheng , Yiwei Chen , Jipeng Zhang , Tuan-Anh Vu , Huimin Zeng , Yue Him Wong Tim , Sai-Kit Yeung

Current text-to-image generation models often struggle to follow textual instructions, especially the ones requiring spatial reasoning. On the other hand, Large Language Models (LLMs), such as GPT-4, have shown remarkable precision in…

Computer Vision and Pattern Recognition · Computer Science 2023-05-31 Tianjun Zhang , Yi Zhang , Vibhav Vineet , Neel Joshi , Xin Wang

Multimodal large language models (MLLMs) are designed to process and integrate information from multiple sources, such as text, speech, images, and videos. Despite its success in language understanding, it is critical to evaluate the…

Computer Vision and Pattern Recognition · Computer Science 2024-04-11 Hao Lu , Xuesong Niu , Jiyao Wang , Yin Wang , Qingyong Hu , Jiaqi Tang , Yuting Zhang , Kaishen Yuan , Bin Huang , Zitong Yu , Dengbo He , Shuiguang Deng , Hao Chen , Yingcong Chen , Shiguang Shan

Multimodal Large Language Models (MLLMs) exhibit impressive capabilities across a variety of tasks, especially when equipped with carefully designed visual prompts. However, existing studies primarily focus on logical reasoning and visual…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Dingning Liu , Cheng Wang , Peng Gao , Renrui Zhang , Xinzhu Ma , Yuan Meng , Zhihui Wang

Multimodal large language models (MLLMs) have achieved impressive performance across various tasks such as image captioning and visual question answer(VQA); however, they often struggle to accurately interpret depth information inherent in…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Hao Yang , Hongbo Zhang , Yanyan Zhao , Bing Qin

Large Multimodal Models (LMMs) have demonstrated impressive performance across various vision and language tasks, yet their potential applications in recommendation tasks with visual assistance remain unexplored. To bridge this gap, we…

Information Retrieval · Computer Science 2023-11-08 Peilin Zhou , Meng Cao , You-Liang Huang , Qichen Ye , Peiyan Zhang , Junling Liu , Yueqi Xie , Yining Hua , Jaeboum Kim

Materials characterization is fundamental to acquiring materials information, revealing the processing-microstructure-property relationships that guide material design and optimization. While multimodal large language models (MLLMs) have…

Computer Vision and Pattern Recognition · Computer Science 2025-09-12 Zhengzhao Lai , Youbin Zheng , Zhenyang Cai , Haonan Lyu , Jinpu Yang , Hongqing Liang , Yan Hu , Benyou Wang

Multimodal large language models (MLLMs) represent an evolutionary expansion in the capabilities of traditional large language models, enabling them to tackle challenges that surpass the scope of purely text-based applications. It leverages…

Computation and Language · Computer Science 2025-01-17 Jinlong He , Pengfei Li , Gang Liu , Genrong He , Zhaolin Chen , Shenjun Zhong

Retouching is an essential task in post-manipulation of raw photographs. Generative editing, guided by text or strokes, provides a new tool accessible to users but can easily change the identity of the original objects in unacceptable and…

Graphics · Computer Science 2025-05-12 Niladri Shekhar Dutt , Duygu Ceylan , Niloy J. Mitra

Determining material properties from camera images can expand the ability to identify complex objects in indoor environments, which is valuable for consumer robotics applications. To support this, we introduce MatPredict, a dataset that…

Computer Vision and Pattern Recognition · Computer Science 2025-05-20 Yuzhen Chen , Hojun Son , Arpan Kusari

Scene generation with 3D assets presents a complex challenge, requiring both high-level semantic understanding and low-level geometric reasoning. While Multimodal Large Language Models (MLLMs) excel at semantic tasks, their application to…

Computer Vision and Pattern Recognition · Computer Science 2025-03-10 Ian Huang , Yanan Bao , Karen Truong , Howard Zhou , Cordelia Schmid , Leonidas Guibas , Alireza Fathi

In this work, we investigate the problem of creating high-fidelity 3D content from only a single image. This is inherently challenging: it essentially involves estimating the underlying 3D geometry while simultaneously hallucinating unseen…

Computer Vision and Pattern Recognition · Computer Science 2023-04-04 Junshu Tang , Tengfei Wang , Bo Zhang , Ting Zhang , Ran Yi , Lizhuang Ma , Dong Chen

Selection of occluded objects is a challenging problem in virtual reality, even more so if multiple objects are involved. With the advent of new artificial intelligence technologies, we explore the possibility of leveraging large language…

Human-Computer Interaction · Computer Science 2024-10-29 Junlong Chen , Jens Grubert , Per Ola Kristensson

Large multimodal language models (MLLMs) such as GPT-4V and GPT-4o have achieved remarkable advancements in understanding and generating multimodal content, showcasing superior quality and capabilities across diverse tasks. However, their…

Computer Vision and Pattern Recognition · Computer Science 2025-01-09 Xuelu Feng , Yunsheng Li , Dongdong Chen , Mei Gao , Mengchen Liu , Junsong Yuan , Chunming Qiao

In this work, we explore how multimodal large language models can support real-time context- and value-aware decision-making. To do so, we combine the GPT-4o language model with a TurtleBot 4 platform simulating a smart vacuum cleaning…

Robotics · Computer Science 2026-02-03 Giulio Antonio Abbo , Senne Lenaerts , Tony Belpaeme