English
Related papers

Related papers: JourneyDB: A Benchmark for Generative Image Unders…

200 papers

The visual world is fundamentally compositional. Visual scenes are defined by the composition of objects and their relations. Hence, it is essential for computer vision systems to reflect and exploit this compositionality to achieve robust…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Shuhao Fu , Andrew Jun Lee , Anna Wang , Ida Momennejad , Trevor Bihl , Hongjing Lu , Taylor W. Webb

Systematic evaluation of Multimodal Large Language Models (MLLMs) is crucial for advancing Artificial General Intelligence (AGI). However, existing benchmarks remain insufficient for rigorously assessing their reasoning capabilities under…

The rapid advancement of talking-head deepfake generation fueled by advanced generative models has elevated the realism of synthetic videos to a level that poses substantial risks in domains such as media, politics, and finance. However,…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Xinqi Xiong , Prakrut Patel , Qingyuan Fan , Amisha Wadhwa , Sarathy Selvam , Xiao Guo , Luchao Qi , Xiaoming Liu , Roni Sengupta

Generating images conditioned on multiple visual references is critical for real-world applications such as multi-subject composition, narrative illustration, and novel view synthesis, yet current models suffer from severe performance…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Zhekai Chen , Yuqing Wang , Manyuan Zhang , Xihui Liu

Human vision combines low-resolution "gist" information from the visual periphery with sparse but high-resolution information from fixated locations to construct a coherent understanding of a visual scene. In this paper, we introduce…

Computer Vision and Pattern Recognition · Computer Science 2026-02-25 Ritik Raina , Abe Leite , Alexandros Graikos , Seoyoung Ahn , Dimitris Samaras , Gregory J. Zelinsky

Personal photo albums are not merely collections of static images but living, ecological archives defined by temporal continuity, social entanglement, and rich metadata, which makes the personalized photo retrieval non-trivial. However,…

Commonsense reasoning, the ability to make logical assumptions about daily scenes, is one core intelligence of human beings. In this work, we present a novel task and dataset for evaluating the ability of text-to-image generative models to…

Multimedia · Computer Science 2024-01-24 Mianzhi Pan , Jianfei Li , Mingyue Yu , Zheng Ma , Kanzhi Cheng , Jianbing Zhang , Jiajun Chen

Conversational generative vision models (CGVMs) like Visual ChatGPT (Wu et al., 2023) have recently emerged from the synthesis of computer vision and natural language processing techniques. These models enable more natural and interactive…

Computer Vision and Pattern Recognition · Computer Science 2023-05-30 Narjes Nikzad Khasmakhi , Meysam Asgari-Chenaghlu , Nabiha Asghar , Philipp Schaer , Dietlind Zühlke

The field of advanced text-to-image generation is witnessing the emergence of unified frameworks that integrate powerful text encoders, such as CLIP and T5, with Diffusion Transformer backbones. Although there have been efforts to control…

Computer Vision and Pattern Recognition · Computer Science 2025-02-28 Liang Chen , Shuai Bai , Wenhao Chai , Weichu Xie , Haozhe Zhao , Leon Vinci , Junyang Lin , Baobao Chang

It has been shown that accurate representation in media improves the well-being of the people who consume it. By contrast, inaccurate representations can negatively affect viewers and lead to harmful perceptions of other cultures. To…

Computer Vision and Pattern Recognition · Computer Science 2023-04-27 Zhixuan Liu , Youeun Shin , Beverley-Claire Okogwu , Youngsik Yun , Lia Coleman , Peter Schaldenbrand , Jihie Kim , Jean Oh

We present JoyAI-Image, a unified multimodal foundation model for visual understanding, text-to-image generation, and instruction-guided image editing. JoyAI-Image couples a spatially enhanced Multimodal Large Language Model (MLLM) with a…

Recent generative data augmentation methods conditioned on both image and text prompts struggle to balance between fidelity and diversity, as it is challenging to preserve essential image details while aligning with varied text prompts.…

Computer Vision and Pattern Recognition · Computer Science 2025-10-20 Tianchen Zhao , Xuanbai Chen , Zhihua Li , Jun Fang , Dongsheng An , Xiang Xu , Zhuowen Tu , Yifan Xing

Understanding perspective is fundamental to human visual perception, yet the extent to which multimodal large language models (MLLMs) internalize perspective geometry remains unclear. We introduce MMPerspective, the first benchmark…

Computer Vision and Pattern Recognition · Computer Science 2025-11-26 Yolo Y. Tang , Pinxin Liu , Zhangyun Tan , Mingqian Feng , Rui Mao , Chao Huang , Jing Bi , Yunzhong Xiao , Susan Liang , Hang Hua , Ali Vosoughi , Luchuan Song , Zeliang Zhang , Chenliang Xu

Interleaved text-and-image generation has been an intriguing research direction, where the models are required to generate both images and text pieces in an arbitrary order. Despite the emerging advancements in interleaved generation, the…

Computer Vision and Pattern Recognition · Computer Science 2024-10-10 Minqian Liu , Zhiyang Xu , Zihao Lin , Trevor Ashby , Joy Rimchala , Jiaxin Zhang , Lifu Huang

Generating realistic and diverse trajectories is a critical challenge in autonomous driving simulation. While Large Language Models (LLMs) show promise, existing methods often rely on structured data like vectorized maps, which fail to…

Artificial Intelligence · Computer Science 2026-03-06 Mingxuan Mu , Guo Yang , Lei Chen , Ping Wu , Jianxun Cui

GPS trajectory data reveals valuable patterns of human mobility and urban dynamics, supporting a variety of spatial applications. However, traditional methods often struggle to extract deep semantic representations and incorporate…

Computers and Society · Computer Science 2025-06-23 Chunhou Ji , Qiumeng Li

In this paper, we introduce knowledge image generation as a new task, alongside the Massive Multi-Discipline Multi-Tier Knowledge-Image Generation Benchmark (MMMG) to probe the reasoning capability of image generation models. Knowledge…

Computer Vision and Pattern Recognition · Computer Science 2026-01-07 Yuxuan Luo , Yuhui Yuan , Junwen Chen , Haonan Cai , Ziyi Yue , Yuwei Yang , Fatima Zohra Daha , Ji Li , Zhouhui Lian

Diffusion models have gained tremendous success in text-to-image generation, yet still lag behind with visual understanding tasks, an area dominated by autoregressive vision-language models. We propose a large-scale and fully end-to-end…

Computer Vision and Pattern Recognition · Computer Science 2025-04-03 Zijie Li , Henry Li , Yichun Shi , Amir Barati Farimani , Yuval Kluger , Linjie Yang , Peng Wang

Understanding road scenes is essential for autonomous driving, as it enables systems to interpret visual surroundings to aid in effective decision-making. We present Roadscapes, a multitask multimodal dataset consisting of upto 9,000 images…

Computer Vision and Pattern Recognition · Computer Science 2026-02-16 Vijayasri Iyer , Maahin Rathinagiriswaran , Jyothikamalesh S

Product images strongly influence consumer decision-making in online marketplaces. Empowered by multimodal contrastive learning, generative AI can output images that closely align with text prompts. Yet existing generative AI models do not…

Artificial Intelligence · Computer Science 2026-05-28 Xiaohang Feng , Yiling Xie
‹ Prev 1 4 5 6 7 8 10 Next ›