English
Related papers

Related papers: The Telephone Game: Evaluating Semantic Drift in U…

200 papers

Unified Multimodal Models (uMMs) aim to support both visual understanding and visual generation within a shared representation. However, existing evaluation protocols assess these two capabilities independently and do not examine whether…

Computer Vision and Pattern Recognition · Computer Science 2026-04-29 Weixing Wang , Liudvikas Zekas , Anton Hackl , Constantin Alexander Auga , Parisa Shahabinejad , Jona Otholt , Antonio Rueda-Toicen , Gerard de Melo

Digital Twins (DTs) represent digital counterparts of physical systems, assets, or processes, referred to as the actual twin (AT). DTs integrate heterogeneous data, models, and semantic technologies to support monitoring, simulation,…

Software Engineering · Computer Science 2026-05-21 Faima Abbasi , Jean-Sébastien Sottet , Cedric Pruski

Unified multimodal models (UMMs) strive to consolidate visual understanding and visual generation within a single architecture. However, prevailing training paradigms independently optimize understanding via sparse text signals and…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Songsong Yu , Yuxin Chen , Ying Shan , Yanwei Li

Recent unified models integrate multimodal understanding and generation within a single framework. However, an "understanding-generation gap" persists, where models can capture user intent but often fail to translate this semantic knowledge…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Qingyang Liu , Bingjie Gao , Canmiao Fu , Zhipeng Huang , Chen Li , Feng Wang , Shuochen Chang , Shaobo Wang , Yali Wang , Keming Ye , Jiangtong Li , Li Niu

The visual dialog task attempts to train an agent to answer multi-turn questions given an image, which requires the deep understanding of interactions between the image and dialog history. Existing researches tend to employ the…

Computation and Language · Computer Science 2022-02-23 Tong Ye , Shijing Si , Jianzong Wang , Rui Wang , Ning Cheng , Jing Xiao

Unified models can handle both multimodal understanding and generation within a single architecture, yet they typically operate in a single pass without iteratively refining their outputs. Many multimodal tasks, especially those involving…

Computer Vision and Pattern Recognition · Computer Science 2026-02-13 Leon Liangyu Chen , Haoyu Ma , Zhipeng Fan , Ziqi Huang , Animesh Sinha , Xiaoliang Dai , Jialiang Wang , Zecheng He , Jianwei Yang , Chunyuan Li , Junzhe Sun , Chu Wang , Serena Yeung-Levy , Felix Juefei-Xu

The efficient operation of modern cellular networks hinges on the accurate analysis of spatio-temporal traffic data. Mastering these patterns is essential for core network functions, chiefly forecasting future load to pre-empt congestion…

Machine Learning · Computer Science 2026-05-13 Yichen Zhang , Jun Li

State-of-the-art models in semantic segmentation primarily operate on single, static images, generating corresponding segmentation masks. This one-shot approach leaves little room for error correction, as the models lack the capability to…

Computer Vision and Pattern Recognition · Computer Science 2023-10-16 Foivos I. Diakogiannis , Suzanne Furby , Peter Caccetta , Xiaoliang Wu , Rodrigo Ibata , Ondrej Hlinka , John Taylor

Accurate interpretation and visualization of human instructions are crucial for text-to-image (T2I) synthesis. However, current models struggle to capture semantic variations from word order changes, and existing evaluations, relying on…

Computation and Language · Computer Science 2025-04-18 Xiangru Zhu , Penglei Sun , Yaoxian Song , Yanghua Xiao , Zhixu Li , Chengyu Wang , Jun Huang , Bei Yang , Xiaoxiao Xu

Thanks to its graphical notation and simplicity, Unified Modeling Language (UML) is a de facto standard and a widespread language used in both industry and academia, despite the fact that its semantics is still informal. The Interaction…

Software Engineering · Computer Science 2014-01-23 Aymen Louati , Chadlia Jerad , Kamel Barkaoui

Unified multimodal models (UMMs) were designed to combine the reasoning ability of large language models (LLMs) with the generation capability of vision models. In practice, however, this synergy remains elusive: UMMs fail to transfer…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Songlin Yang , Xianghao Kong , Anyi Rao

In recent years, integrating multimodal understanding and generation into a single unified model has emerged as a promising paradigm. While this approach achieves strong results in text-to-image (T2I) generation, it still struggles with…

Computer Vision and Pattern Recognition · Computer Science 2026-04-09 Ziyun Zeng , David Junhao Zhang , Wei Li , Mike Zheng Shou

Most semantic drift studies report multiple signals e.g., embedding displacement, neighbor changes, distributional divergence, and recursive trajectory instability, without a shared explanatory theory that relates them. This paper proposes…

Computation and Language · Computer Science 2026-02-24 Stephen Russell

Unified multimodal models have recently demonstrated strong generative capabilities, yet whether and when generation improves understanding remains unclear. Existing benchmarks lack a systematic exploration of the specific tasks where…

Computer Vision and Pattern Recognition · Computer Science 2026-03-04 Zimo Wen , Boxiu Li , Wanbo Zhang , Junxiang Lei , Xiaoyu Chen , Yijia Fan , Qi Zhang , Yujiang Wang , Lili Qiu , Bo Li , Ziwei Liu , Caihua Shan , Yifan Yang , Yifei Shen

Unifying text-image contrastive learning and text-to-image (T2I) generation in a single end-to-end model is challenging because the two objectives demand opposing masking regimes: contrastive alignment needs near-complete visible tokens,…

Text-to-Image (T2I) diffusion models have made remarkable advancements in generative modeling; however, they face a trade-off between inference speed and image quality, posing challenges for efficient deployment. Existing distilled T2I…

Computer Vision and Pattern Recognition · Computer Science 2025-05-29 Senmao Li , Lei Wang , Kai Wang , Tao Liu , Jiehang Xie , Joost van de Weijer , Fahad Shahbaz Khan , Shiqi Yang , Yaxing Wang , Jian Yang

Semantic segmentation of large-scale 3D point clouds is crucial for applications such as autonomous driving and urban digital twins. However, the sparse sampling pattern of LiDAR and the view-dependent geometric distortion in image…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Shuai Zhang , Zhecheng Shi , Zhuxiao Li , Jing Ou , Tengxi Wang , Yuan Liu , Wufan Zhao

In computer vision, Image Difference Captioning (IDC) is crucial for accurately describing variations between closely related images. Traditional IDC methods often rely on specialist models, which restrict their applicability across varied…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Erdong Hu , Longteng Guo , Tongtian Yue , Zijia Zhao , Shuning Xue , Jing Liu

Image-text matching (ITM) aims to address the fundamental challenge of aligning visual and textual modalities, which inherently differ in their representations, continuous, high-dimensional image features vs. discrete, structured text. We…

Multimedia · Computer Science 2025-07-14 Junyu Chen , Yihua Gao , Mingyong Li

Currently, the success of large language models (LLMs) illustrates that a unified multitasking approach can significantly enhance model usability, streamline deployment, and foster synergistic benefits across different tasks. However, in…

Computer Vision and Pattern Recognition · Computer Science 2025-10-03 Bin Xia , Yuechen Zhang , Jingyao Li , Chengyao Wang , Yitong Wang , Xinglong Wu , Bei Yu , Jiaya Jia