English
Related papers

Related papers: FungiTastic: A multi-modal dataset and benchmark f…

200 papers

Food image segmentation is a critical and indispensible task for developing health-related applications such as estimating food calories and nutrients. Existing food image segmentation models are underperforming due to two reasons: (1)…

Computer Vision and Pattern Recognition · Computer Science 2021-05-13 Xiongwei Wu , Xin Fu , Ying Liu , Ee-Peng Lim , Steven C. H. Hoi , Qianru Sun

The development of foundation vision models has pushed the general visual recognition to a high level, but cannot well address the fine-grained recognition in specialized domain such as invasive species classification. Identifying and…

Computer Vision and Pattern Recognition · Computer Science 2023-10-31 Wei He , Kai Han , Ying Nie , Chengcheng Wang , Yunhe Wang

We introduce GLAMI-1M: the largest multilingual image-text classification dataset and benchmark. The dataset contains images of fashion products with item descriptions, each in 1 of 13 languages. Categorization into 191 classes has…

Computer Vision and Pattern Recognition · Computer Science 2022-11-29 Vaclav Kosar , Antonín Hoskovec , Milan Šulc , Radek Bartyzal

The scarcity of high-quality, labelled retinal imaging data, which presents a significant challenge in the development of machine learning models for ophthalmology, hinders progress in the field. Existing methods for synthesising Colour…

Image and Video Processing · Electrical Eng. & Systems 2025-07-18 Junzhi Ning , Cheng Tang , Kaijing Zhou , Diping Song , Lihao Liu , Ming Hu , Wei Li , Huihui Xu , Yanzhou Su , Tianbin Li , Jiyao Liu , Jin Ye , Sheng Zhang , Yuanfeng Ji , Junjun He

Smartphone clip-on microscopes turn everyday devices into low-cost, portable imaging systems that can even reveal fungal structures at the microscopic level, enabling mold inspection beyond unaided visual checks. In this paper, we introduce…

Computer Vision and Pattern Recognition · Computer Science 2026-04-13 Dinh Nam Pham , Leonard Prokisch , Bennet Meyer , Jonas Thumbs

Recent advances in multimodal models have demonstrated remarkable text-guided image editing capabilities, with systems like GPT-4o and Nano-Banana setting new benchmarks. However, the research community's progress remains constrained by the…

Computer Vision and Pattern Recognition · Computer Science 2025-10-23 Yusu Qian , Eli Bocek-Rivele , Liangchen Song , Jialing Tong , Yinfei Yang , Jiasen Lu , Wenze Hu , Zhe Gan

We apply deep metric learning for the first time to the problem of classifying planktic foraminifer shells on microscopic images. This species recognition task is an important information source and scientific pillar for reconstructing past…

Computer Vision and Pattern Recognition · Computer Science 2022-03-25 Tayfun Karaderi , Tilo Burghardt , Allison Y. Hsiang , Jacob Ramaer , Daniela N. Schmidt

Robust machine learning depends on clean data, yet current image data cleaning benchmarks rely on synthetic noise or narrow human studies, limiting comparison and real-world relevance. We introduce CleanPatrick, the first large-scale…

Artificial intelligence (AI)-enabled diagnostics in maxillofacial pathology require structured, high-quality multimodal datasets. However, existing resources provide limited ameloblastoma coverage and lack the format consistency needed for…

Artificial Intelligence · Computer Science 2026-02-06 Ajo Babu George , Anna Mariam John , Athul Anoop , Balu Bhasuran

Robust whole-slide image (WSI) analysis under strict data-governance remains challenging due to substantial cross-institutional stain heterogeneity. Domain generalization (DG) mitigates these shifts but typically requires centralized data,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Fengyi Zhang , Junya Zhang , Wenzhuo Sun

Accurate classification of second-trimester fetal ultrasound images remains challenging due to low image quality, high intra-class variability, and significant class imbalance. In this work, we introduce a simple yet powerful, biologically…

Image and Video Processing · Electrical Eng. & Systems 2025-06-11 Rinat Prochii , Elizaveta Dakhova , Pavel Birulin , Maxim Sharaev

Stomata play a crucial role in regulating plant physiological processes and reflecting environmental responses. However, accurate and high-throughput stomatal phenotyping remains challenging, as conventional approaches rely on destructive…

Computer Vision and Pattern Recognition · Computer Science 2026-04-22 Quanling Zhao , Meng'en Qin , Yanfeng Sun , Yuan Miao , Xiaohui Yang

Accurate segmentation and classification of brain tumors from Magnetic Resonance Imaging (MRI) remain key challenges in medical image analysis, primarily due to the lack of high-quality, balanced, and diverse datasets with expert…

Image and Video Processing · Electrical Eng. & Systems 2026-01-29 Amirreza Fateh , Yasin Rezvani , Sara Moayedi , Sadjad Rezvani , Fatemeh Fateh , Mansoor Fateh , Vahid Abolghasemi

Recent years have witnessed increasing attention in cartoon media, powered by the strong demands of industrial applications. As the first step to understand this media, cartoon face recognition is a crucial but less-explored task with few…

Computer Vision and Pattern Recognition · Computer Science 2020-06-30 Yi Zheng , Yifan Zhao , Mengyuan Ren , He Yan , Xiangju Lu , Junhui Liu , Jia Li

Recent advances in large vision-language models (LVLMs) have demonstrated strong performance on general-purpose medical tasks. However, their effectiveness in specialized domains such as dentistry remains underexplored. In particular,…

Computer Vision and Pattern Recognition · Computer Science 2025-09-12 Jing Hao , Yuxuan Fan , Yanpeng Sun , Kaixin Guo , Lizhuo Lin , Jinrong Yang , Qi Yong H. Ai , Lun M. Wong , Hao Tang , Kuo Feng Hung

Semi-iNat is a challenging dataset for semi-supervised classification with a long-tailed distribution of classes, fine-grained categories, and domain shifts between labeled and unlabeled data. This dataset is behind the second iteration of…

Computer Vision and Pattern Recognition · Computer Science 2021-06-24 Jong-Chyi Su , Subhransu Maji

Foundation models are becoming increasingly effective in the medical domain, offering pre-trained models on large datasets that can be readily adapted for downstream tasks. Despite progress, fetal ultrasound images remain a challenging…

Food classification serves as the basic step of image-based dietary assessment to predict the types of foods in each input image. However, food image predictions in a real world scenario are usually long-tail distributed among different…

Computer Vision and Pattern Recognition · Computer Science 2022-10-27 Jiangpeng He , Luotao Lin , Heather Eicher-Miller , Fengqing Zhu

Assessing scientific claims requires identifying, extracting, and reasoning with multimodal data expressed in information-rich figures in scientific literature. Despite the large body of work in scientific QA, figure captioning, and other…

Computation and Language · Computer Science 2025-07-31 Yash Kumar Lal , Manikanta Bandham , Mohammad Saqib Hasan , Apoorva Kashi , Mahnaz Koupaee , Niranjan Balasubramanian

We introduce a new dataset, MELINDA, for Multimodal biomEdicaL experImeNt methoD clAssification. The dataset is collected in a fully automated distant supervision manner, where the labels are obtained from an existing curated database, and…

Computation and Language · Computer Science 2020-12-18 Te-Lin Wu , Shikhar Singh , Sayan Paul , Gully Burns , Nanyun Peng
‹ Prev 1 4 5 6 7 8 10 Next ›