English
Related papers

Related papers: BIOSCAN-5M: A Multimodal Dataset for Insect Biodiv…

200 papers

Typographic attacks exploit the interplay between text and visual content in multimodal foundation models, causing misclassifications when misleading text is embedded within images. Existing datasets are limited in size and diversity,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Justus Westerhoff , Erblina Purelku , Jakob Hackstein , Jonas Loos , Leo Pinetzki , Erik Rodner , Lorenz Hufe

Depth estimation and 3D reconstruction have been extensively studied as core topics in computer vision. Starting from rigid objects with relatively simple geometric shapes, such as vehicles, the research has expanded to address general…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Muhammad Aamir , Naoya Muramatsu , Sangyun Shin , Matthew Wijers , Jia-Xing Zhong , Xinyu Hou , Amir Patel , Andrew Loveridge , Andrew Markham

Cluster closure, defined as the progressive filling of gaps between the berries in a grape bunch, is a key trait in vineyard management, impacting disease risk. However, traditional visual scoring methods are labor-intensive, subjective,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Xiangzhi Tong , Chengrui Zhang , Mac Flaherty , Andre Matteo Garcia , Dominic Gorman , Jonathan Jaramillo , Justine E. Vanden Heuvel , Yu Jiang

The ability to perform complex tasks from detailed instructions is a key to many remarkable achievements of our species. As humans, we are not only capable of performing a wide variety of tasks but also very complex ones that may entail…

Artificial Intelligence · Computer Science 2024-07-23 Xiaoxuan Lei , Lucas Gomez , Hao Yuan Bai , Pouya Bashivan

Understanding the spatio-temporal distribution of species is a cornerstone of ecology and conservation. By pairing species observations with geographic and environmental predictors, researchers can model the relationship between an…

Computer Vision and Pattern Recognition · Computer Science 2025-10-01 Christophe Botella , Benjamin Deneu , Diego Marcos , Maximilien Servajean , Theo Larcher , Cesar Leblanc , Joaquim Estopinan , Pierre Bonnet , Alexis Joly

As Artificial Intelligence (AI) has developed rapidly over the past few decades, the new generation of AI, Large Language Models (LLMs) trained on massive datasets, has achieved ground-breaking performance in many applications. Further…

Computer Vision and Pattern Recognition · Computer Science 2024-10-29 Yijiashun Qi , Shuzhang Cai , Zunduo Zhao , Jiaming Li , Yanbin Lin , Zhiqiang Wang

Combining multiple sensors enables a robot to maximize its perceptual awareness of environments and enhance its robustness to external disturbance, crucial to robotic navigation. This paper proposes the FusionPortable benchmark, a complete…

The difficulty to measure or predict species community composition at fine spatio-temporal resolution and over large spatial scales severely hampers our ability to understand species assemblages and take appropriate conservation measures.…

The convergence of IoT sensing, edge computing, and machine learning is transforming precision livestock farming. Yet bioacoustic data streams remain underused because of computational complexity and ecological validity challenges. We…

Sound · Computer Science 2025-10-17 Mayuri Kate , Suresh Neethirajan

Recent advancements in biological research leverage the integration of molecules, proteins, and natural language to enhance drug discovery. However, current models exhibit several limitations, such as the generation of invalid molecular…

Computation and Language · Computer Science 2024-01-30 Qizhi Pei , Wei Zhang , Jinhua Zhu , Kehan Wu , Kaiyuan Gao , Lijun Wu , Yingce Xia , Rui Yan

Trust and interpretability are crucial for the use of Artificial Intelligence (AI) in scientific research, but current models often operate as black boxes offering limited transparency and justifications for their outputs. We introduce…

Synthesizing information from multiple data sources plays a crucial role in the practice of modern medicine. Current applications of artificial intelligence in medicine often focus on single-modality data due to a lack of publicly…

Medicine is inherently multimodal, with rich data modalities spanning text, imaging, genomics, and more. Generalist biomedical artificial intelligence (AI) systems that flexibly encode, integrate, and interpret this data at scale can…

Spatial domain identification requires jointly modeling molecular signatures and physical coordinates, yet current tools frequently over-smooth biological boundaries, require user-specified cluster numbers, and lack principled multimodal…

Applications · Statistics 2026-05-18 Xin Li , Xiaofei Dong , Zhenke Duan , Lulu Shang , Xiao Wang , Xinyuan Song , Hanwen Ning , Guanyu Hu

Multimodal Large Language Models demonstrate strong performance on natural image understanding, yet exhibit limited capability in interpreting scientific images, including but not limited to schematic diagrams, experimental…

Computer Vision and Pattern Recognition · Computer Science 2026-02-17 Haoyi Tao , Chaozheng Huang , Nan Wang , Han Lyu , Linfeng Zhang , Guolin Ke , Xi Fang

Although deep learning models have achieved state-of-the-art performance on a number of vision tasks, generalization over high dimensional multi-modal data, and reliable predictive uncertainty estimation are still active areas of research.…

Machine Learning · Computer Science 2020-12-10 Samarth Sinha , Homanga Bharadhwaj , Anirudh Goyal , Hugo Larochelle , Animesh Garg , Florian Shkurti

Despite recent advances in AI, multimodal culturally diverse datasets are still limited, particularly for regions in the Middle East and Africa. In this paper, we introduce EgMM-Corpus, a multimodal dataset dedicated to Egyptian culture. By…

Computation and Language · Computer Science 2025-10-21 Mohamed Gamil , Abdelrahman Elsayed , Abdelrahman Lila , Ahmed Gad , Hesham Abdelgawad , Mohamed Aref , Ahmed Fares

This paper presents the multi-modal BigEarthNet (BigEarthNet-MM) benchmark archive made up of 590,326 pairs of Sentinel-1 and Sentinel-2 image patches to support the deep learning (DL) studies in multi-modal multi-label remote sensing (RS)…

Computer Vision and Pattern Recognition · Computer Science 2021-06-18 Gencer Sumbul , Arne de Wall , Tristan Kreuziger , Filipe Marcelino , Hugo Costa , Pedro Benevides , Mário Caetano , Begüm Demir , Volker Markl

Bytes form the basis of the digital world and thus are a promising building block for multimodal foundation models. Recently, Byte Language Models (BLMs) have emerged to overcome tokenization, yet the excessive length of bytestreams…

Computation and Language · Computer Science 2025-02-21 Eric Egli , Matteo Manica , Jannis Born

This paper presents a dataset of agricultural pest images captured over five years by thousands of small holder farmers and farming extension workers across India. The dataset has been used to support a mobile application that relies on…

Computer Vision and Pattern Recognition · Computer Science 2023-04-04 Jerome White , Chandan Agrawal , Anmol Ojha , Apoorv Agnihotri , Makkunda Sharma , Jigar Doshi