English
Related papers

Related papers: From Pixels to BFS: High Maze Accuracy Does Not Im…

200 papers

To validate the safety of automated vehicles (AV), scenario-based testing aims to systematically describe driving scenarios an AV might encounter. In this process, continuous inputs such as velocities result in an infinite number of…

Machine Learning · Computer Science 2022-12-08 Max Winkelmann , Mike Kohlhoff , Hadj Hamma Tadjine , Steffen Müller

Text-to-image synthesis refers to generating visual-realistic and semantically consistent images from given textual descriptions. Previous approaches generate an initial low-resolution image and then refine it to be high-resolution. Despite…

Computer Vision and Pattern Recognition · Computer Science 2023-02-21 Haoran Sun , Yang Wang , Haipeng Liu , Biao Qian

The image geolocalization task aims to predict the location where an image was taken anywhere on Earth using visual clues. Existing large vision-language model (LVLM) approaches leverage world knowledge, chain-of-thought reasoning, and…

Computer Vision and Pattern Recognition · Computer Science 2026-01-12 Yuxiang Ji , Yong Wang , Ziyu Ma , Yiming Hu , Hailang Huang , Xuecai Hu , Guanhua Chen , Liaoni Wu , Xiangxiang Chu

Utilizing transformer architectures for semantic segmentation of high-resolution images is hindered by the attention's quadratic computational complexity in the number of tokens. A solution to this challenge involves decreasing the number…

Computer Vision and Pattern Recognition · Computer Science 2024-05-24 Daniel Kienzle , Marco Kantonis , Robin Schön , Rainer Lienhart

Challenges have become the state-of-the-art approach to benchmark image analysis algorithms in a comparative manner. While the validation on identical data sets was a great step forward, results analysis is often restricted to pure ranking…

In this paper, we identify an important reproducibility challenge in the image-to-set prediction literature that impedes proper comparisons among published methods, namely, researchers use different evaluation protocols to assess their…

Computer Vision and Pattern Recognition · Computer Science 2020-05-28 Luis Pineda , Amaia Salvador , Michal Drozdzal , Adriana Romero

Multimodal large language models (MLLMs) extend the success of language models to visual understanding, and recent efforts have sought to build unified MLLMs that support both understanding and generation. However, constructing such models…

Computer Vision and Pattern Recognition · Computer Science 2025-10-03 Hanyu Wang , Jiaming Han , Ziyan Yang , Qi Zhao , Shanchuan Lin , Xiangyu Yue , Abhinav Shrivastava , Zhenheng Yang , Hao Chen

Recently, researchers have explored ML-based Traffic Engineering (TE), leveraging neural networks to solve TE problems traditionally addressed by optimization. However, existing ML-based TE schemes remain impractical: they either fail to…

Networking and Internet Architecture · Computer Science 2026-04-16 Ximeng Liu , Zhuoran Liu , Yingming Mao , Yatao Li , Shizhen Zhao , Xinbing Wang

Autoregressive models have shown remarkable success in image generation by adapting sequential prediction techniques from language modeling. However, applying these approaches to images requires discretizing continuous pixel data through…

Computer Vision and Pattern Recognition · Computer Science 2025-03-21 Ziyao Guo , Kaipeng Zhang , Michael Qizhe Shieh

The fast evolution of generative models has heightened the demand for reliable detection of AI-generated images. To tackle this challenge, we introduce FUSE, a hybrid system that combines spectral features extracted through Fast Fourier…

Computer Vision and Pattern Recognition · Computer Science 2025-12-29 Md. Zahid Hossain , Most. Sharmin Sultana Samu , Md. Kamrozzaman Bhuiyan , Farhad Uz Zaman , Md. Rakibul Islam

Two questions regarding practitioners' use of patent embeddings arise: (i) Does one fine-tuning recipe suffice for all downstream applications? (ii) Is fine-tuning on one patent landscape sufficient for downstream application on other…

Information Retrieval · Computer Science 2026-05-27 Amirhossein Yousefiramandi , Ciaran Cooney

Despite recent progress in multimodal agentic systems, existing approaches often treat image manipulation and web search as disjoint capabilities, rely heavily on costly reinforcement learning, and lack planning grounded in real…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Yifan Zhang , Liang Hu , Haofeng Sun , Peiyu Wang , Yichen Wei , Shukang Yin , Jiangbo Pei , Wei Shen , Peng Xia , Yi Peng , Tianyidan Xie , Eric Li , Yang Liu , Xuchen Song , Yahui Zhou

Jointing visual-semantic embeddings (VSE) have become a research hotpot for the task of image annotation, which suffers from the issue of semantic gap, i.e., the gap between images' visual features (low-level) and labels' semantic features…

Computer Vision and Pattern Recognition · Computer Science 2018-08-14 Guibing Guo , Songlin Zhai , Fajie Yuan , Yuan Liu , Xingwei Wang

Large Multimodal Models (LMMs) have made significant strides in visual question-answering for single images. Recent advancements like long-context LMMs have allowed them to ingest larger, or even multiple, images. However, the ability to…

Computer Vision and Pattern Recognition · Computer Science 2025-03-12 Tsung-Han Wu , Giscard Biamby , Jerome Quenum , Ritwik Gupta , Joseph E. Gonzalez , Trevor Darrell , David M. Chan

Cartographic reasoning is the skill of interpreting geographic relationships by aligning legends, map scales, compass directions, map texts, and geometries across one or more map images. Although essential as a concrete cognitive capability…

Computer Vision and Pattern Recognition · Computer Science 2026-03-02 Jiyoon Pyo , Yuankun Jiao , Dongwon Jung , Zekun Li , Leeje Jang , Sofia Kirsanova , Jina Kim , Yijun Lin , Qin Liu , Junyi Xie , Hadi Askari , Nan Xu , Muhao Chen , Yao-Yi Chiang

Multimodal Large Language Models (MLLMs) demonstrate a strong understanding of the real world and can even handle complex tasks. However, they still fail on some straightforward visual question-answering (VQA) problems. This paper dives…

Computation and Language · Computer Science 2024-10-16 Sihang Zhao , Youliang Yuan , Xiaoying Tang , Pinjia He

Image clustering aims to partition unlabeled image datasets into distinct groups. A core aspect of this task is constructing and leveraging prior knowledge to guide the clustering process. Recent approaches introduce semantic descriptions…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Feijiang Li , Zhenxiong Li , Jieting Wang , Zizheng Jiu , Saixiong Liu , Liang Du

Image classifiers are information-discarding machines, by design. Yet, how these models discard information remains mysterious. We hypothesize that one way for image classifiers to reach high accuracy is to first zoom to the most…

Computer Vision and Pattern Recognition · Computer Science 2023-10-10 Mohammad Reza Taesiri , Giang Nguyen , Sarra Habchi , Cor-Paul Bezemer , Anh Nguyen

Multimodal foundation models have shown compelling but conflicting performance in medical image interpretation. However, the mechanisms by which these models integrate and prioritize different data modalities, including images and text,…

Computer Vision and Pattern Recognition · Computer Science 2024-11-26 Thomas Buckley , James A. Diao , Pranav Rajpurkar , Adam Rodman , Arjun K. Manrai

Large Language Models (LLMs) show potential for enhancing robotic path planning. This paper assesses visual input's utility for multimodal LLMs in such tasks via a comprehensive benchmark. We evaluated 15 multimodal LLMs on generating valid…

Robotics · Computer Science 2025-07-17 Jacinto Colan , Ana Davila , Yasuhisa Hasegawa