English
Related papers

Related papers: SciDraw-6K: A Multilingual Scientific Illustration…

200 papers

This paper introduces a large-scale multimodal and multilingual dataset that aims to facilitate research on grounding words to images in their contextual usage in language. The dataset consists of images selected to unambiguously illustrate…

Computation and Language · Computer Science 2022-06-20 Josiah Wang , Pranava Madhyastha , Josiel Figueiredo , Chiraag Lala , Lucia Specia

In this work, we announce a comprehensive well curated and opensource dataset with millions of samples for pre-college and college level problems in mathematicsand science. A preliminary set of results using transformer architecture with…

Numerical Analysis · Mathematics 2021-10-01 Neeraj Kollepara , Snehith Kumar Chatakonda , Pawan Kumar

Lecture slide presentations, a sequence of pages that contain text and figures accompanied by speech, are constructed and presented carefully in order to optimally transfer knowledge to students. Previous studies in multimedia and…

Artificial Intelligence · Computer Science 2022-08-18 Dong Won Lee , Chaitanya Ahuja , Paul Pu Liang , Sanika Natu , Louis-Philippe Morency

This paper introduces a public dataset of 1.4 million procedurally-generated bicycle designs represented parametrically, as JSON files, and as rasterized images. The dataset is created through the use of a rendering engine which harnesses…

Computer Vision and Pattern Recognition · Computer Science 2024-02-13 Lyle Regenwetter , Yazan Abu Obaideh , Amin Heyrani Nobari , Faez Ahmed

Visual Genome is a dataset connecting structured image information with English language. We present ``Hindi Visual Genome'', a multimodal dataset consisting of text and images suitable for English-Hindi multimodal machine translation task…

Computation and Language · Computer Science 2019-07-23 Shantipriya Parida , Ondřej Bojar , Satya Ranjan Dash

The rapid advancement of multimodal large language models (MLLMs) offers new opportunities for complex scientific challenges, yet their application in earth science-especially at the graduate level-remains underexplored due to a lack of…

Artificial Intelligence · Computer Science 2026-05-05 Xiangyu Zhao , Wanghan Xu , Bo Liu , Yuhao Zhou , Fenghua Ling , Ben Fei , Xiaoyu Yue , Lei Bai , Wenlong Zhang , Xiao-Ming Wu

The ever-increasing volume of paper submissions makes it difficult to stay informed about the latest state-of-the-art research. To address this challenge, we introduce LEGOBench, a benchmark for evaluating systems that generate scientific…

Computation and Language · Computer Science 2024-02-22 Shruti Singh , Shoaib Alam , Husain Malwat , Mayank Singh

Foundation models have shown remarkable success across scientific domains, yet their impact in chemistry remains limited due to the absence of diverse, large-scale, high-quality datasets that reflect the field's multifaceted nature. We…

The automatic clinical caption generation problem is referred to as proposed model combining the analysis of frontal chest X-Ray scans with structured patient information from the radiology records. We combine two language models, the…

Computer Vision and Pattern Recognition · Computer Science 2022-09-29 Alexander Selivanov , Oleg Y. Rogov , Daniil Chesakov , Artem Shelmanov , Irina Fedulova , Dmitry V. Dylov

We study personalized figure caption generation using author profile data from scientific papers. Our experiments demonstrate that rich author profile data, combined with relevant metadata, can significantly improve the personalization…

Computation and Language · Computer Science 2025-10-01 Jaeyoung Kim , Jongho Lee , Hongjun Choi , Sion Jang

In the rapidly evolving field of Artificial Intelligence Generated Content (AIGC), a central challenge is distinguishing AI-synthesized images from natural ones. Despite the impressive capabilities of advanced generative models in producing…

Artificial Intelligence · Computer Science 2025-08-12 Renyang Liu , Ziyu Lyu , Wei Zhou , See-Kiong Ng

Artificial-intelligence (AI) agent frameworks have been developed for autonomous scientific simulations, but most current agent frameworks are tailored to a single or a small set of software packages. Herein, FermiLink, a unified and…

Despite significant advancements in general AI, its effectiveness in the medical domain is limited by the lack of specialized medical knowledge. To address this, we formulate GMAI-VL-5.5M, a multimodal medical dataset created by converting…

Computer Vision and Pattern Recognition · Computer Science 2025-03-28 Tianbin Li , Yanzhou Su , Wei Li , Bin Fu , Zhe Chen , Ziyan Huang , Guoan Wang , Chenglong Ma , Ying Chen , Ming Hu , Yanjun Li , Pengcheng Chen , Xiaowei Hu , Zhongying Deng , Yuanfeng Ji , Jin Ye , Yu Qiao , Junjun He

The proliferation of open-source scientific software for science and research presents opportunities and challenges. In this paper, we introduce the SciCat dataset -- a comprehensive collection of Free-Libre Open Source Software (FLOSS)…

Software Engineering · Computer Science 2023-12-12 Addi Malviya-Thakur , Reed Milewicz , Lavinia Paganini , Ahmed Samir Imam Mahmoud , Audris Mockus

Multimodal generative AI systems like Stable Diffusion, DALL-E, and MidJourney have fundamentally changed how synthetic images are created. These tools drive innovation but also enable the spread of misleading content, false information,…

Although image captioning has a vast array of applications, it has not reached its full potential in languages other than English. Arabic, for instance, although the native language of more than 400 million people, remains largely…

Computer Vision and Pattern Recognition · Computer Science 2023-11-16 Abdelrahman Mohamed , Fakhraddin Alwajih , El Moatez Billah Nagoudi , Alcides Alcoba Inciarte , Muhammad Abdul-Mageed

We introduce GLAMI-1M: the largest multilingual image-text classification dataset and benchmark. The dataset contains images of fashion products with item descriptions, each in 1 of 13 languages. Categorization into 191 classes has…

Computer Vision and Pattern Recognition · Computer Science 2022-11-29 Vaclav Kosar , Antonín Hoskovec , Milan Šulc , Radek Bartyzal

Multi-modal models are data hungry. While datasets with natural images are abundant, medical image datasets can not afford the same luxury. To enable representation learning for medical images at scale, we turn to YouTube, a platform with a…

Computer Vision and Pattern Recognition · Computer Science 2026-01-15 Wisdom O. Ikezogwo , Kevin Zhang , Mehmet Saygin Seyfioglu , Fatemeh Ghezloo , Linda Shapiro , Ranjay Krishna

The integration of artificial intelligence into scientific research has reached a new pinnacle with GPT-4V, a large language model featuring enhanced vision capabilities, accessible through ChatGPT or an API. This study demonstrates the…

Artificial Intelligence · Computer Science 2025-08-19 Zhiling Zheng , Zhiguo He , Omar Khattab , Nakul Rampal , Matei A. Zaharia , Christian Borgs , Jennifer T. Chayes , Omar M. Yaghi