English
Related papers

Related papers: GeoDiv: Framework For Measuring Geographical Diver…

200 papers

Comprehensive and constructive evaluation protocols play an important role in the development of sophisticated text-to-video (T2V) generation models. Existing evaluation protocols primarily focus on temporal consistency and content…

Computer Vision and Pattern Recognition · Computer Science 2024-07-02 Mingxiang Liao , Hannan Lu , Xinyu Zhang , Fang Wan , Tianyu Wang , Yuzhong Zhao , Wangmeng Zuo , Qixiang Ye , Jingdong Wang

The robustness of a model for real-world deployment is decided by how well it performs on unseen data and distinguishes between in-domain and out-of-domain samples. Visual document classifiers have shown impressive performance on…

Computer Vision and Pattern Recognition · Computer Science 2023-05-30 Fnu Mohbat , Mohammed J. Zaki , Catherine Finegan-Dollak , Ashish Verma

Massive web-crawled image-text datasets lay the foundation for recent progress in multimodal learning. These datasets are designed with the goal of training a model to do well on standard computer vision benchmarks, many of which, however,…

Computer Vision and Pattern Recognition · Computer Science 2025-09-16 Thao Nguyen , Matthew Wallingford , Sebastin Santy , Wei-Chiu Ma , Sewoong Oh , Ludwig Schmidt , Pang Wei Koh , Ranjay Krishna

Recent advancements in text-to-image (T2I) generation models have transformed the field. However, challenges persist in generating images that reflect demanding textual descriptions, especially for fine-grained details and unusual…

Multimedia · Computer Science 2025-02-21 Ran Li , Xiaomeng Jin , Heng ji

We propose T2I-ReasonBench, a benchmark evaluating reasoning capabilities of text-to-image (T2I) models. It consists of four dimensions: Idiom Interpretation, Textual Image Design, Entity-Reasoning and Scientific-Reasoning. We propose a…

Computer Vision and Pattern Recognition · Computer Science 2025-08-26 Kaiyue Sun , Rongyao Fang , Chengqi Duan , Xian Liu , Xihui Liu

Recent advances in discriminative and generative pretraining have yielded geometry estimation models with strong generalization capabilities. While discriminative monocular geometry estimation methods rely on large-scale fine-tuning data to…

Computer Vision and Pattern Recognition · Computer Science 2024-06-24 Yongtao Ge , Guangkai Xu , Zhiyue Zhao , Libo Sun , Zheng Huang , Yanlong Sun , Hao Chen , Chunhua Shen

Employing a single, unified model (UM) for both visual understanding (image-to-text: I2T) and visual generation (text-to-image: T2I) has opened a new direction in Visual Language Model (VLM) research. While UMs can also support broader…

Computer Vision and Pattern Recognition · Computer Science 2025-10-07 Sabbir Mollah , Rohit Gupta , Sirnam Swetha , Qingyang Liu , Ahnaf Munir , Mubarak Shah

Socio-economic indicators like regional GDP, population, and education levels, are crucial to shaping policy decisions and fostering sustainable development. This research introduces GeoReg a regression model that integrates diverse data…

Machine Learning · Computer Science 2026-03-20 Kyeongjin Ahn , Sungwon Han , Seungeon Lee , Donghyun Ahn , Hyoshin Kim , Jungwon Kim , Jihee Kim , Sangyoon Park , Meeyoung Cha

Large-scale instance-level training data is scarce, so models are typically trained on domain-specific datasets. Yet in real-world retrieval, they must handle diverse domains, making generalization to unseen data critical. We introduce…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Pavel Suma , Giorgos Kordopatis-Zilos , Yannis Kalantidis , Giorgos Tolias

The increasing frequency and intensity of natural disasters call for rapid and accurate damage assessment. In response, disaster benchmark datasets from high-resolution satellite imagery have been constructed to develop methods for…

Computer Vision and Pattern Recognition · Computer Science 2025-01-22 Kyeongjin Ahn , Sungwon Han , Sungwon Park , Jihee Kim , Sangyoon Park , Meeyoung Cha

While text-to-image (T2I) models can synthesize high-quality images, their performance degrades significantly when prompted with novel or out-of-distribution (OOD) entities due to inherent knowledge cutoffs. We introduce World-To-Image, a…

Computer Vision and Pattern Recognition · Computer Science 2025-10-07 Moo Hyun Son , Jintaek Oh , Sun Bin Mun , Jaechul Roh , Sehyun Choi

Current text-to-image (T2I) generation models achieve promising results, but they fail on the scenarios where the knowledge implied in the text prompt is uncertain. For example, a T2I model released in February would struggle to generate a…

Computer Vision and Pattern Recognition · Computer Science 2025-05-22 Chuanhao Li , Jianwen Sun , Yukang Feng , Mingliang Zhai , Yifan Chang , Kaipeng Zhang

Multilingual vision-language models (VLMs) promise universal image-text retrieval, yet their social biases remain underexplored. We perform the first systematic audit of four public multilingual CLIP variants: M-CLIP, NLLB-CLIP,…

Computation and Language · Computer Science 2025-11-20 Zahraa Al Sahili , Ioannis Patras , Matthew Purver

As Text-to-Image (TTI) diffusion models become increasingly influential in content creation, growing attention is being directed toward their societal and cultural implications. While prior research has primarily examined demographic and…

Computer Vision and Pattern Recognition · Computer Science 2026-02-23 Maria-Teresa De Rosa Palmini , Eva Cetinic

Diffusion-based \textit{image-to-video} (I2V) generation has become a central direction in generative models by turning a reference image, with optional conditions, into a temporally coherent video. Compared with broader video generation…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Xianlong Wang , Wenbo Pan , Shijia Zhou , Ke Li , Yuqi Wang , Zeyu Ye , Hangtao Zhang , Leo Yu Zhang , Xiaohua Jia

Recent advances in large-scale text-to-image generation models have led to a surge in subject-driven text-to-image generation, which aims to produce customized images that align with textual descriptions while preserving the identity of…

Computer Vision and Pattern Recognition · Computer Science 2025-05-29 Kewen Chen , Xiaobin Hu , Wenqi Ren

Large vision-language models (LVLMs) have recently achieved significant progress, demonstrating strong capabilities in open-world visual understanding. However, it is not yet clear how LVLMs address demographic biases in real life,…

Computation and Language · Computer Science 2025-09-23 Xuyang Wu , Yuan Wang , Hsin-Tai Wu , Zhiqiang Tao , Yi Fang

Benchmarks of the multilingual capabilities of text-to-image (T2I) models compare generated images prompted in a test language to an expected image distribution over a concept set. One such benchmark, "Conceptual Coverage Across Languages"…

Computation and Language · Computer Science 2024-03-19 Michael Saxon , Yiran Luo , Sharon Levy , Chitta Baral , Yezhou Yang , William Yang Wang

Image editing aims to edit the given synthetic or real image to meet the specific requirements from users. It is widely studied in recent years as a promising and challenging field of Artificial Intelligence Generative Content (AIGC).…

Computer Vision and Pattern Recognition · Computer Science 2024-06-21 Xincheng Shuai , Henghui Ding , Xingjun Ma , Rongcheng Tu , Yu-Gang Jiang , Dacheng Tao

This paper develops an approach to language identification in which the set of languages considered by the model depends on the geographic origin of the text in question. Given that many digital corpora can be geo-referenced at the country…

Computation and Language · Computer Science 2024-03-18 Jonathan Dunn , Lane Edwards-Brown
‹ Prev 1 8 9 10 Next ›