English
Related papers

Related papers: A Vision Language Model for Generating Procedural …

200 papers

Reasoning about spatial relationships between objects is essential for many real-world robotic tasks, such as fetch-and-delivery, object rearrangement, and object search. The ability to detect and disambiguate different objects and identify…

Computer Vision and Pattern Recognition · Computer Science 2024-10-11 Negar Nejatishahidin , Madhukar Reddy Vongala , Jana Kosecka

Three-dimensional geospatial analysis is critical for applications in urban planning, climate adaptation, and environmental assessment. However, current methodologies depend on costly, specialized sensors, such as LiDAR and multispectral…

Computer Vision and Pattern Recognition · Computer Science 2025-12-23 Mai Tsujimoto , Junjue Wang , Weihao Xuan , Naoto Yokoya

We tackle the challenging problem of creating full and accurate three dimensional reconstructions of botanical trees with the topological and geometric accuracy required for subsequent physical simulation, e.g. in response to wind forces.…

Computer Vision and Pattern Recognition · Computer Science 2018-12-24 Ed Quigley , Winnie Lin , Yilin Zhu , Ronald Fedkiw

Vision-language models (VLMs) struggle with 3D-related tasks such as spatial cognition and physical understanding, which are crucial for real-world applications like robotics and embodied agents. We attribute this to a modality gap between…

Computer Vision and Pattern Recognition · Computer Science 2026-04-16 Yifan Liu , Fangneng Zhan , Kaichen Zhou , Yilun Du , Paul Pu Liang , Hanspeter Pfister

Combining large language models with evolutionary computation algorithms represents a promising research direction leveraging the remarkable generative and in-context learning capabilities of LLMs with the strengths of evolutionary…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Tobias Preintner , Weixuan Yuan , Adrian König , Thomas Bäck , Elena Raponi , Niki van Stein

Semantic correspondence made tremendous progress through the recent advancements of large vision models (LVM). While these LVMs have been shown to reliably capture local semantics, the same can currently not be said for capturing global…

Computer Vision and Pattern Recognition · Computer Science 2025-03-31 Krispin Wandel , Hesheng Wang

Grounding natural language in 3D environments is a critical step toward achieving robust 3D vision-language alignment. Current datasets and models for 3D visual grounding predominantly focus on identifying and localizing objects from…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Zhuofan Zhang , Ziyu Zhu , Junhao Li , Pengxiang Li , Tianxu Wang , Tengyu Liu , Xiaojian Ma , Yixin Chen , Baoxiong Jia , Siyuan Huang , Qing Li

Vision-language models have demonstrated impressive capabilities in generating 2D images under various conditions; however, the success of these models is largely enabled by extensive, readily available pretrained foundation models.…

Image and Video Processing · Electrical Eng. & Systems 2025-10-02 Mohamed Mohamed , Brennan Nichyporuk , Douglas L. Arnold , Tal Arbel

3D Visual Grounding (3DVG) aims to localize objects in 3D scenes using natural language descriptions. Although supervised methods achieve higher accuracy in constrained settings, zero-shot 3DVG holds greater promise for real-world…

Computer Vision and Pattern Recognition · Computer Science 2025-08-29 Jiawen Lin , Shiran Bian , Yihang Zhu , Wenbin Tan , Yachao Zhang , Yuan Xie , Yanyun Qu

In the manufacturing industry, computer vision systems based on artificial intelligence (AI) are widely used to reduce costs and increase production. Training these AI models requires a large amount of training data that is costly to…

Computer Vision and Pattern Recognition · Computer Science 2026-01-08 Steven Moonen , Rob Salaets , Kenneth Batstone , Abdellatif Bey-Temsamani , Nick Michiels

Despite their ability to understand chemical knowledge, large language models (LLMs) remain limited in their capacity to propose novel molecules with desired functions (e.g., drug-like properties). In addition, the molecules that LLMs…

Semantic reconstruction of agricultural scenes plays a vital role in tasks such as phenotyping and yield estimation. However, traditional approaches that rely on manual scanning or fixed camera setups remain a major bottleneck in this…

Robotics · Computer Science 2026-01-21 Jose Cuaran , Naveen K. Upalapati , Girish Chowdhary

Current text-to-image models struggle to provide precise camera control using natural language alone. In this work, we present a framework for precise camera control with global scene understanding in text-to-image generation by learning…

Computer Vision and Pattern Recognition · Computer Science 2026-04-23 Xinxuan Lu , Charless Fowlkes , Alexander C. Berg

Reliable crop disease detection requires models that perform consistently across diverse acquisition conditions, yet existing evaluations often focus on single architectural families or lab-generated datasets. This work presents a…

Computer Vision and Pattern Recognition · Computer Science 2026-04-09 Hamza Mooraj , George Pantazopoulos , Alessandro Suglia

Automation in agriculture plays a vital role in addressing challenges related to crop monitoring and disease management, particularly through early detection systems. This study investigates the effectiveness of combining multimodal Large…

Computer Vision and Pattern Recognition · Computer Science 2025-04-30 Konstantinos I. Roumeliotis , Ranjan Sapkota , Manoj Karkee , Nikolaos D. Tselikas , Dimitrios K. Nasiopoulos

Much scientific enquiry across disciplines is founded upon a mechanistic treatment of dynamic systems that ties form to function. A highly visible instance of this is in molecular biology, where an important goal is to determine…

Biomolecules · Quantitative Biology 2021-06-17 Xiaojie Guo , Yuanqi Du , Sivani Tadepalli , Liang Zhao , Amarda Shehu

Traditional Automatic License Plate Recognition (ALPR) systems employ multi-stage pipelines consisting of object detection networks followed by separate Optical Character Recognition (OCR) modules, introducing compounding errors, increased…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Karthik Sivakoti

As robotics become increasingly integrated into construction workflows, their ability to interpret and respond to human behavior will be essential for enabling safe and effective collaboration. Vision-Language Models (VLMs) have emerged as…

Computer Vision and Pattern Recognition · Computer Science 2026-01-19 Hieu Bui , Nathaniel E. Chodosh , Arash Tavakoli

Large language models (LLMs) and Vision-Language Models (VLMs) have been proven to excel at multiple tasks, such as commonsense reasoning. Powerful as these models can be, they are not grounded in the 3D physical world, which involves…

Computer Vision and Pattern Recognition · Computer Science 2023-07-25 Yining Hong , Haoyu Zhen , Peihao Chen , Shuhong Zheng , Yilun Du , Zhenfang Chen , Chuang Gan

Synthetic generation of three-dimensional cell models from histopathological images aims to enhance understanding of cell mutation, and progression of cancer, necessary for clinical assessment and optimal treatment. Classical reconstruction…

Image and Video Processing · Electrical Eng. & Systems 2021-02-09 Yoav Alon , Xiang Yu , Huiyu Zhou