中文
相关论文

相关论文: JourneyDB: A Benchmark for Generative Image Unders…

200 篇论文

Image representations are often evaluated through disjointed, task-specific protocols, leading to a fragmented understanding of model capabilities. For instance, it is unclear whether an image embedding model adept at clustering images is…

Multimodal retrieval is becoming a crucial component of modern AI applications, yet its evaluation lags behind the demands of more realistic and challenging scenarios. Existing benchmarks primarily probe surface-level semantic…

Recent years have seen remarkable progress in deep learning powered visual content creation. This includes deep generative 3D-aware image synthesis, which produces high-idelity images in a 3D-consistent manner while simultaneously capturing…

计算机视觉与模式识别 · 计算机科学 2023-10-04 Weihao Xia , Jing-Hao Xue

Developing interpretable models for neurodevelopmental disorders (NDDs) diagnosis presents significant challenges in effectively encoding, decoding, and integrating multimodal neuroimaging data. While many existing machine learning…

计算机视觉与模式识别 · 计算机科学 2026-04-24 Yueyang Li , Lei Chen , Wenhao Dong , Shengyu Gong , Zijian Kang , Boyang Wei , Weiming Zeng , Hongjie Yan , Lingbin Bian , Zhiguo Zhang , Wai Ting Siok , Nizhuan Wang

Multimodal large language models (MLLMs) have demonstrated significant progress in semantic scene understanding and text-image alignment, with reasoning variants enhancing performance on more complex tasks involving mathematics and logic.…

计算机视觉与模式识别 · 计算机科学 2026-03-13 Sicheng Feng , Song Wang , Shuyi Ouyang , Lingdong Kong , Zikai Song , Jianke Zhu , Huan Wang , Xinchao Wang

End-to-end In-Image Machine Translation (IIMT) aims to convert text embedded within an image into a target language while preserving the original visual context, layout, and rendering style. However, existing IIMT benchmarks are largely…

计算机视觉与模式识别 · 计算机科学 2026-04-02 Jiahao Lyu , Pei Fu , Zhenhang Li , Weichao Zeng , Shaojie Zhang , Jiahui Yang , Can Ma , Yu Zhou , Zhenbo Luo , Jian Luan

Unifying multimodal understanding and generation is a compelling frontier that is beginning to emerge in the medical field. However, the limited existing unified medical models typically treat understanding and generation as disjoint…

计算机视觉与模式识别 · 计算机科学 2026-05-12 Weiren Zhao , Yi Dong , Cheng Chen

Recent advancements in Large Vision-Language Models (VLMs), have greatly enhanced their capability to jointly process text and images. However, despite extensive benchmarks evaluating visual comprehension (e.g., diagrams, color schemes, OCR…

计算与语言 · 计算机科学 2025-05-27 Benjamin Clavié , Florian Brand

Recent generative models produce images with a level of authenticity that makes them nearly indistinguishable from real photos and artwork. Potential harmful use cases of these models, necessitate the creation of robust synthetic image…

计算机视觉与模式识别 · 计算机科学 2025-01-15 Delyan Boychev , Radostin Cholakov

Recent breakthroughs in the field of language-guided image generation have yielded impressive achievements, enabling the creation of high-quality and diverse images based on user instructions.Although the synthesis performance is…

计算机视觉与模式识别 · 计算机科学 2023-05-24 Jian Ma , Mingjun Zhao , Chen Chen , Ruichen Wang , Di Niu , Haonan Lu , Xiaodong Lin

Multimodal large language models (MLLMs) have achieved impressive performance across various tasks such as image captioning and visual question answer(VQA); however, they often struggle to accurately interpret depth information inherent in…

计算机视觉与模式识别 · 计算机科学 2026-03-09 Hao Yang , Hongbo Zhang , Yanyan Zhao , Bing Qin

Understanding mid-level road semantics, which capture the structural and contextual cues that link low-level perception to high-level planning, is essential for reliable autonomous driving and digital map construction. However, existing…

计算机视觉与模式识别 · 计算机科学 2025-12-01 Xiyan Liu , Han Wang , Yuhu Wang , Junjie Cai , Zhe Cao , Jianzhong Yang , Zhen Lu

Recent advancements in image generation models have enabled the prediction of future Graphical User Interface (GUI) states based on user instructions. However, existing benchmarks primarily focus on general domain visual fidelity, leaving…

Text-to-image generation models have achieved strong performance in culturally homogeneous settings, yet their ability to generate multicultural scenes, where people and landmarks originate from different cultures, remains largely…

计算机视觉与模式识别 · 计算机科学 2026-04-20 Parth Bhalerao , Mounika Yalamarty , Brian Trinh , Oana Ignat

Deep generative models have led to significant advances in cross-modal generation such as text-to-image synthesis. Training these models typically requires paired data with direct correspondence between modalities. We introduce the novel…

计算机视觉与模式识别 · 计算机科学 2019-08-21 Shuang Ma , Daniel McDuff , Yale Song

Recent breakthroughs in large multimodal models (LMMs), such as the impressive GPT-4o-Native, have demonstrated remarkable proficiency in following general-purpose instructions for image generation. However, current benchmarks often lack…

计算机视觉与模式识别 · 计算机科学 2025-05-27 Jiayu Wang , Yang Jiao , Yue Yu , Tianwen Qian , Shaoxiang Chen , Jingjing Chen , Yu-Gang Jiang

Driving World Models (DWMs) have been developing rapidly with the advances of generative models. However, existing DWMs lack 3D scene understanding capabilities and can only generate content conditioned on input data, without the ability to…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Tianchen Deng , Xuefeng Chen , Yi Chen , Qu Chen , Yuyao Xu , Lijin Yang , Le Xu , Yu Zhang , Bo Zhang , Wuxiong Huang , Hesheng Wang

Unified vision large language models (VLLMs) have recently achieved impressive advancements in both multimodal understanding and generation, powering applications such as visual question answering and text-guided image synthesis. However,…

计算与语言 · 计算机科学 2025-09-19 Pengyu Wang , Shaojun Zhou , Chenkun Tan , Xinghao Wang , Wei Huang , Zhen Ye , Zhaowei Li , Botian Jiang , Dong Zhang , Xipeng Qiu

Using multimodal foundation models to analyze table images is a high-value yet challenging application in consumer and enterprise scenarios. Despite its importance, current evaluations rely largely on structured-text tables or clean…

计算机视觉与模式识别 · 计算机科学 2026-05-25 Junzhe Huang , Xiaoxiao Sun , Yan Yang , Yuxuan Hou , Ruotian Zhang , Sirui Li , Hehe Fan , Serena Yeung-Levy , Xin Yu

In the evolving landscape of multimodal language models, understanding the nuanced meanings conveyed through visual cues - such as satire, insult, or critique - remains a significant challenge. Existing evaluation benchmarks primarily focus…

机器学习 · 计算机科学 2025-02-25 Xiaofei Yin , Yijie Hong , Ya Guo , Yi Tu , Weiqiang Wang , Gongshen Liu , Huijia zhu