English
Related papers

Related papers: WAFFLE: Multimodal Floorplan Understanding in the …

200 papers

Wireless foundation models (WFMs) have recently demonstrated promising capabilities, jointly performing multiple wireless functions and adapting effectively to new environments. However, while current WFMs process only one modality,…

Signal Processing · Electrical Eng. & Systems 2026-02-20 Ahmed Aboulfotouh , Hatem Abou-Zeid

Recent advancements in Large Vision-Language Models (VLMs) have demonstrated exceptional semantic understanding, yet these models consistently struggle with spatial reasoning, often failing at fundamental geometric tasks such as depth…

Computer Vision and Pattern Recognition · Computer Science 2026-05-13 Zishan Liu , Ruoxi Zang , Yanglin Zhang , Wei Liu , Yin Zhang , Jian Yao , Jiayin Zheng , Zhengzhe Liu

Inter-object relations underpin spatial intelligence, yet existing representations -- linguistic prepositions or object-level scene graphs -- are too coarse to specify which regions actually support, contain, or contact one another, leading…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Yinuo Bai , Peijun Xu , Kuixiang Shao , Yuyang Jiao , Jingxuan Zhang , Kaixin Yao , Jiayuan Gu , Jingyi Yu

In this paper, we observe that semi-structured tabulated text is ubiquitous; understanding them requires not only comprehending the meaning of text fragments, but also implicit relationships between them. We argue that such data can prove…

Computation and Language · Computer Science 2020-05-14 Vivek Gupta , Maitrey Mehta , Pegah Nokhiz , Vivek Srikumar

Due to the rapid growth of the World Wide Web, resource discovery becomes an increasing problem. As an answer to the demand for information management, a third generation of World-Wide Web tools will evolve: information gathering and…

Software Engineering · Computer Science 2018-10-31 Robert E. Kent , Christian Neuss

Vision systems to see and reason about the compositional nature of visual scenes are fundamental to understanding our world. The complex relations between objects and their locations, ambiguities, and variations in the real-world…

Computer Vision and Pattern Recognition · Computer Science 2023-07-27 Muhammad Awais , Muzammal Naseer , Salman Khan , Rao Muhammad Anwer , Hisham Cholakkal , Mubarak Shah , Ming-Hsuan Yang , Fahad Shahbaz Khan

The present study explores the interpretability of latent spaces produced by time series foundation models, focusing on their potential for visual analysis tasks. Specifically, we evaluate the MOMENT family of models, a set of…

Recognition of floor plans has been a challenging and popular task. Despite that many recent approaches have been proposed for this task, they typically fail to make the room-level unified prediction. Specifically, multiple semantic…

Computer Vision and Pattern Recognition · Computer Science 2022-11-01 Zhangyu Wang , Ningyuan Sun

Building Information Modeling (BIM) technology is a key component of modern construction engineering and project management workflows. As-is BIM models that represent the spatial reality of a project site can offer crucial information to…

Computer Vision and Pattern Recognition · Computer Science 2023-05-26 Seongyong Kim , Yosuke Yajima , Jisoo Park , Jingdao Chen , Yong K. Cho

We introduce TableVista, a comprehensive benchmark for evaluating foundation models in multimodal table reasoning under visual and structural complexity. TableVista consists of 3,000 high-quality table reasoning problems, where each…

Computation and Language · Computer Science 2026-05-08 Zheyuan Yang , Liqiang Shang , Junjie Chen , Xun Yang , Chenglong Xu , Bo Yuan , Chenyuan Jiao , Yaoru Sun , Yilun Zhao

This paper proposes LayoutLLM, a more flexible document analysis method for understanding imaged documents. Visually Rich Document Understanding tasks, such as document image classification and information extraction, have gained…

Computation and Language · Computer Science 2024-03-22 Masato Fujitake

Multi-modal Large Language Models (MLLMs) have demonstrated remarkable capabilities in understanding and generating content across various modalities, such as images and text. However, their interpretability remains a challenge, hindering…

Computer Vision and Pattern Recognition · Computer Science 2024-05-29 Loris Giulivi , Giacomo Boracchi

We are perceiving and communicating with the world in a multisensory manner, where different information sources are sophisticatedly processed and interpreted by separate parts of the human brain to constitute a complex, yet harmonious and…

Computer Vision and Pattern Recognition · Computer Science 2024-06-12 Ye Zhu , Yu Wu , Nicu Sebe , Yan Yan

Foundation models have demonstrated remarkable success across diverse domains and tasks, primarily due to the thrive of large-scale, diverse, and high-quality datasets. However, in the field of medical imaging, the curation and assembling…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Zhongying Deng , Cheng Tang , Ziyan Huang , Jiashi Lin , Ying Chen , Junzhi Ning , Chenglong Ma , Jiyao Liu , Wei Li , Yinghao Zhu , Shujian Gao , Yanyan Huang , Sibo Ju , Yanzhou Su , Pengcheng Chen , Wenhao Tang , Tianbin Li , Haoyu Wang , Yuanfeng Ji , Hui Sun , Shaobo Min , Liang Peng , Feilong Tang , Haochen Xue , Rulin Zhou , Chaoyang Zhang , Wenjie Li , Shaohao Rui , Weijie Ma , Xingyue Zhao , Yibin Wang , Kun Yuan , Zhaohui Lu , Shujun Wang , Jinjie Wei , Lihao Liu , Dingkang Yang , Lin Wang , Yulong Li , Haolin Yang , Yiqing Shen , Lequan Yu , Xiaowei Hu , Yun Gu , Yicheng Wu , Benyou Wang , Minghui Zhang , Angelica I. Aviles-Rivero , Qi Gao , Hongming Shan , Xiaoyu Ren , Fang Yan , Hongyu Zhou , Haodong Duan , Maosong Cao , Shanshan Wang , Bin Fu , Xiaomeng Li , Zhi Hou , Chunfeng Song , Lei Bai , Yuan Cheng , Yuandong Pu , Xiang Li , Wenhai Wang , Hao Chen , Jiaxin Zhuang , Songyang Zhang , Huiguang He , Mengzhang Li , Bohan Zhuang , Zhian Bai , Rongshan Yu , Liansheng Wang , Yukun Zhou , Xiaosong Wang , Xin Guo , Guanbin Li , Xiangru Lin , Dakai Jin , Mianxin Liu , Wenlong Zhang , Qi Qin , Conghui He , Yuqiang Li , Ye Luo , Nanqing Dong , Jie Xu , Wenqi Shao , Bo Zhang , Qiujuan Yan , Yihao Liu , Jun Ma , Zhi Lu , Yuewen Cao , Zongwei Zhou , Jianming Liang , Shixiang Tang , Qi Duan , Dongzhan Zhou , Chen Jiang , Yuyin Zhou , Yanwu Xu , Jiancheng Yang , Shaoting Zhang , Xiaohong Liu , Siqi Luo , Yi Xin , Chaoyu Liu , Haochen Wen , Xin Chen , Alejandro Lozano , Min Woo Sun , Yuhui Zhang , Yue Yao , Xiaoxiao Sun , Serena Yeung-Levy , Xia Li , Jing Ke , Chunhui Zhang , Zongyuan Ge , Ming Hu , Jin Ye , Zhifeng Li , Yirong Chen , Yu Qiao , Junjun He

Enterprise documents such as forms, invoices, receipts, reports, contracts, and other similar records, often carry rich semantics at the intersection of textual and spatial modalities. The visual cues offered by their complex layouts play a…

Computation and Language · Computer Science 2024-01-03 Dongsheng Wang , Natraj Raman , Mathieu Sibue , Zhiqiang Ma , Petr Babkin , Simerjot Kaur , Yulong Pei , Armineh Nourbakhsh , Xiaomo Liu

Recent advancements in foundation models have improved autonomous tool usage and reasoning, but their capabilities in map-based reasoning remain underexplored. To address this, we introduce MapEval, a benchmark designed to assess foundation…

Multimodal Large Language Models (MLLM) have made significant progress in the field of document analysis. Despite this, existing benchmarks typically focus only on extracting text and simple layout information, neglecting the complex…

Computer Vision and Pattern Recognition · Computer Science 2024-07-04 Lei Chen , Feng Yan , Yujie Zhong , Shaoxiang Chen , Zequn Jie , Lin Ma

In this paper, we provide two case studies to demonstrate how artificial intelligence can empower civil engineering. In the first case, a machine learning-assisted framework, BRAILS, is proposed for city-scale building information modeling.…

Computer Vision and Pattern Recognition · Computer Science 2020-08-25 Qian Yu , Chaofeng Wang , Barbaros Cetiner , Stella X. Yu , Frank Mckenna , Ertugrul Taciroglu , Kincho H. Law

We introduce WordScape, a novel pipeline for the creation of cross-disciplinary, multilingual corpora comprising millions of pages with annotations for document layout detection. Relating visual and textual items on document pages has…

Multimodal foundation models that can holistically process text alongside images, video, audio, and other sensory modalities are increasingly used in a variety of real-world applications. However, it is challenging to characterize and study…

‹ Prev 1 3 4 5 6 7 10 Next ›