English
Related papers

Related papers: Crowdsource, Crawl, or Generate? Creating SEA-VL, …

200 papers

Recently, there has been an increasing number of efforts to introduce models capable of generating natural language explanations (NLEs) for their predictions on vision-language (VL) tasks. Such models are appealing, because they can provide…

Computer Vision and Pattern Recognition · Computer Science 2021-08-19 Maxime Kayser , Oana-Maria Camburu , Leonard Salewski , Cornelius Emde , Virginie Do , Zeynep Akata , Thomas Lukasiewicz

In this paper, we introduce the Semantic Environment Atlas (SEA), a novel mapping approach designed to enhance visual navigation capabilities of embodied agents. The SEA utilizes semantic graph maps that intricately delineate the…

Artificial Intelligence · Computer Science 2024-10-15 Nuri Kim , Jeongho Park , Mineui Hong , Songhwai Oh

Coral reefs are vital yet vulnerable ecosystems that require continuous monitoring to support conservation. While coral reef images provide essential information in coral monitoring, interpreting such images remains challenging due to the…

Computer Vision and Pattern Recognition · Computer Science 2025-11-04 Hongyong Han , Wei Wang , Gaowei Zhang , Mingjie Li , Yi Wang

With the rapid development of VR technology, the demand for high-quality 3D models is increasing. Traditional methods struggle with efficiency and quality in large-scale customization. This paper introduces a deep-learning framework that…

Computer Vision and Pattern Recognition · Computer Science 2024-09-05 Jie Fu , Shun Fu , Mick Grierson

Multi-view clustering, a long-standing and important research problem, focuses on mining complementary information from diverse views. However, existing works often fuse multiple views' representations or handle clustering in a common…

Computer Vision and Pattern Recognition · Computer Science 2021-12-06 Jie Xu , Yazhou Ren , Huayi Tang , Xiaorong Pu , Xiaofeng Zhu , Ming Zeng , Lifang He

The rapid progress of large language models (LLMs) raises concerns about cultural bias, fairness, and performance in diverse languages and underrepresented regions. Addressing these gaps requires large-scale resources grounded in…

Computation and Language · Computer Science 2026-04-08 Firoj Alam , Md Arid Hasan , Sahinur Rahman Laskar , Mucahid Kutlu , Kareem Darwish , Shammur Absar Chowdhury

Image-generating AI, which allows users to create images from text, is increasingly used to produce visual content. Despite its advancements, cultural biases in AI-generated images have raised significant concerns. While much research has…

Human-Computer Interaction · Computer Science 2025-04-08 Xingyu Lan , Jiaxi An , Yisu Guo , Chiyou Tong , Xintong Cai , Jun Zhang

Vision-Language Models (VLMs) are pretrained on large, diverse, and noisy web-crawled datasets. This underscores the critical need for dataset pruning, as the quality of these datasets is strongly correlated with the performance of VLMs on…

Computer Vision and Pattern Recognition · Computer Science 2024-03-12 Anas Mahmoud , Mostafa Elhoushi , Amro Abbas , Yu Yang , Newsha Ardalani , Hugh Leather , Ari Morcos

A crucial challenge for generative large language models (LLMs) is diversity: when a user's prompt is under-specified, models may follow implicit assumptions while generating a response, which may result in homogenization of the responses,…

The accelerating growth of photographic collections has outpaced manual cataloguing, motivating the use of vision language models (VLMs) to automate metadata generation. This study examines whether Al-generated catalogue descriptions can…

Human-Computer Interaction · Computer Science 2025-07-11 Line Abele , Gerrit Anders , Tolgahan Aydın , Jürgen Buder , Helen Fischer , Dominik Kimmel , Markus Huff

Despite recent advancements in vision-language models, their performance remains suboptimal on images from non-western cultures due to underrepresentation in training datasets. Various benchmarks have been proposed to test models' cultural…

Computation and Language · Computer Science 2024-07-02 Mehar Bhatia , Sahithya Ravi , Aditya Chinchure , Eunjeong Hwang , Vered Shwartz

Large Language Models (LLMs) encode meanings of words in the form of distributed semantics. Distributed semantics capture common statistical patterns among language tokens (words, phrases, and sentences) from large amounts of data. LLMs…

Computation and Language · Computer Science 2023-06-27 Yuxin Zi , Kaushik Roy , Vignesh Narayanan , Manas Gaur , Amit Sheth

Massive web-crawled image-text datasets lay the foundation for recent progress in multimodal learning. These datasets are designed with the goal of training a model to do well on standard computer vision benchmarks, many of which, however,…

Computer Vision and Pattern Recognition · Computer Science 2025-09-16 Thao Nguyen , Matthew Wallingford , Sebastin Santy , Wei-Chiu Ma , Sewoong Oh , Ludwig Schmidt , Pang Wei Koh , Ranjay Krishna

Assessing and enhancing human learning through question-answering is vital, yet automating this process remains challenging. While large language models (LLMs) excel at summarization and query responses, their ability to generate meaningful…

Computation and Language · Computer Science 2025-02-25 Kimia Noorbakhsh , Joseph Chandler , Pantea Karimi , Mohammad Alizadeh , Hari Balakrishnan

We study cultural and socioeconomic diversity in contrastive vision-language models (VLMs). Using a broad range of benchmark datasets and evaluation metrics, we bring to attention several important findings. First, the common filtering of…

Computer Vision and Pattern Recognition · Computer Science 2024-10-25 Angéline Pouget , Lucas Beyer , Emanuele Bugliarello , Xiao Wang , Andreas Peter Steiner , Xiaohua Zhai , Ibrahim Alabdulmohsin

We introduce MAIA (Multimodal AI Assessment), a native-Italian benchmark designed for fine-grained investigation of the reasoning abilities of visual language models on videos. MAIA differs from other available video benchmarks for its…

While generative multilingual models are rapidly being deployed, their safety and fairness evaluations are largely limited to resources collected in English. This is especially problematic for evaluations targeting inherently socio-cultural…

Computation and Language · Computer Science 2024-03-12 Mukul Bhutani , Kevin Robinson , Vinodkumar Prabhakaran , Shachi Dave , Sunipa Dev

Causal thinking enables humans to understand not just what is seen, but why it happens. To replicate this capability in modern AI systems, we introduce the task of visual causal discovery. It requires models to infer cause-and-effect…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Yize Zhang , Meiqi Chen , Sirui Chen , Bo Peng , Yanxi Zhang , Tianyu Li , Chaochao Lu

Ensuring cultural values alignment in Large Language Models (LLMs) remains a critical challenge, as these models often embed Western-centric biases from their training data, leading to misrepresentations and fairness concerns in…

Computation and Language · Computer Science 2025-05-09 Wonduk Seo , Zonghao Yuan , Yi Bu

Visual question answering is an important task in both natural language and vision understanding. However, in most of the public visual question answering datasets such as VQA, CLEVR, the questions are human generated that specific to the…

Computation and Language · Computer Science 2022-08-08 Bingning Wang , Feiyang Lv , Ting Yao , Yiming Yuan , Jin Ma , Yu Luo , Haijin Liang