中文
相关论文

相关论文: FungiTastic: A multi-modal dataset and benchmark f…

200 篇论文

Food image segmentation is a critical and indispensible task for developing health-related applications such as estimating food calories and nutrients. Existing food image segmentation models are underperforming due to two reasons: (1)…

计算机视觉与模式识别 · 计算机科学 2021-05-13 Xiongwei Wu , Xin Fu , Ying Liu , Ee-Peng Lim , Steven C. H. Hoi , Qianru Sun

The development of foundation vision models has pushed the general visual recognition to a high level, but cannot well address the fine-grained recognition in specialized domain such as invasive species classification. Identifying and…

计算机视觉与模式识别 · 计算机科学 2023-10-31 Wei He , Kai Han , Ying Nie , Chengcheng Wang , Yunhe Wang

We introduce GLAMI-1M: the largest multilingual image-text classification dataset and benchmark. The dataset contains images of fashion products with item descriptions, each in 1 of 13 languages. Categorization into 191 classes has…

计算机视觉与模式识别 · 计算机科学 2022-11-29 Vaclav Kosar , Antonín Hoskovec , Milan Šulc , Radek Bartyzal

The scarcity of high-quality, labelled retinal imaging data, which presents a significant challenge in the development of machine learning models for ophthalmology, hinders progress in the field. Existing methods for synthesising Colour…

图像与视频处理 · 电气工程与系统科学 2025-07-18 Junzhi Ning , Cheng Tang , Kaijing Zhou , Diping Song , Lihao Liu , Ming Hu , Wei Li , Huihui Xu , Yanzhou Su , Tianbin Li , Jiyao Liu , Jin Ye , Sheng Zhang , Yuanfeng Ji , Junjun He

Smartphone clip-on microscopes turn everyday devices into low-cost, portable imaging systems that can even reveal fungal structures at the microscopic level, enabling mold inspection beyond unaided visual checks. In this paper, we introduce…

计算机视觉与模式识别 · 计算机科学 2026-04-13 Dinh Nam Pham , Leonard Prokisch , Bennet Meyer , Jonas Thumbs

Recent advances in multimodal models have demonstrated remarkable text-guided image editing capabilities, with systems like GPT-4o and Nano-Banana setting new benchmarks. However, the research community's progress remains constrained by the…

计算机视觉与模式识别 · 计算机科学 2025-10-23 Yusu Qian , Eli Bocek-Rivele , Liangchen Song , Jialing Tong , Yinfei Yang , Jiasen Lu , Wenze Hu , Zhe Gan

We apply deep metric learning for the first time to the problem of classifying planktic foraminifer shells on microscopic images. This species recognition task is an important information source and scientific pillar for reconstructing past…

计算机视觉与模式识别 · 计算机科学 2022-03-25 Tayfun Karaderi , Tilo Burghardt , Allison Y. Hsiang , Jacob Ramaer , Daniela N. Schmidt

Robust machine learning depends on clean data, yet current image data cleaning benchmarks rely on synthetic noise or narrow human studies, limiting comparison and real-world relevance. We introduce CleanPatrick, the first large-scale…

Artificial intelligence (AI)-enabled diagnostics in maxillofacial pathology require structured, high-quality multimodal datasets. However, existing resources provide limited ameloblastoma coverage and lack the format consistency needed for…

人工智能 · 计算机科学 2026-02-06 Ajo Babu George , Anna Mariam John , Athul Anoop , Balu Bhasuran

Robust whole-slide image (WSI) analysis under strict data-governance remains challenging due to substantial cross-institutional stain heterogeneity. Domain generalization (DG) mitigates these shifts but typically requires centralized data,…

计算机视觉与模式识别 · 计算机科学 2026-05-15 Fengyi Zhang , Junya Zhang , Wenzhuo Sun

Accurate classification of second-trimester fetal ultrasound images remains challenging due to low image quality, high intra-class variability, and significant class imbalance. In this work, we introduce a simple yet powerful, biologically…

图像与视频处理 · 电气工程与系统科学 2025-06-11 Rinat Prochii , Elizaveta Dakhova , Pavel Birulin , Maxim Sharaev

Stomata play a crucial role in regulating plant physiological processes and reflecting environmental responses. However, accurate and high-throughput stomatal phenotyping remains challenging, as conventional approaches rely on destructive…

计算机视觉与模式识别 · 计算机科学 2026-04-22 Quanling Zhao , Meng'en Qin , Yanfeng Sun , Yuan Miao , Xiaohui Yang

Accurate segmentation and classification of brain tumors from Magnetic Resonance Imaging (MRI) remain key challenges in medical image analysis, primarily due to the lack of high-quality, balanced, and diverse datasets with expert…

图像与视频处理 · 电气工程与系统科学 2026-01-29 Amirreza Fateh , Yasin Rezvani , Sara Moayedi , Sadjad Rezvani , Fatemeh Fateh , Mansoor Fateh , Vahid Abolghasemi

Recent years have witnessed increasing attention in cartoon media, powered by the strong demands of industrial applications. As the first step to understand this media, cartoon face recognition is a crucial but less-explored task with few…

计算机视觉与模式识别 · 计算机科学 2020-06-30 Yi Zheng , Yifan Zhao , Mengyuan Ren , He Yan , Xiangju Lu , Junhui Liu , Jia Li

Recent advances in large vision-language models (LVLMs) have demonstrated strong performance on general-purpose medical tasks. However, their effectiveness in specialized domains such as dentistry remains underexplored. In particular,…

计算机视觉与模式识别 · 计算机科学 2025-09-12 Jing Hao , Yuxuan Fan , Yanpeng Sun , Kaixin Guo , Lizhuo Lin , Jinrong Yang , Qi Yong H. Ai , Lun M. Wong , Hao Tang , Kuo Feng Hung

Semi-iNat is a challenging dataset for semi-supervised classification with a long-tailed distribution of classes, fine-grained categories, and domain shifts between labeled and unlabeled data. This dataset is behind the second iteration of…

计算机视觉与模式识别 · 计算机科学 2021-06-24 Jong-Chyi Su , Subhransu Maji

Foundation models are becoming increasingly effective in the medical domain, offering pre-trained models on large datasets that can be readily adapted for downstream tasks. Despite progress, fetal ultrasound images remain a challenging…

Food classification serves as the basic step of image-based dietary assessment to predict the types of foods in each input image. However, food image predictions in a real world scenario are usually long-tail distributed among different…

计算机视觉与模式识别 · 计算机科学 2022-10-27 Jiangpeng He , Luotao Lin , Heather Eicher-Miller , Fengqing Zhu

Assessing scientific claims requires identifying, extracting, and reasoning with multimodal data expressed in information-rich figures in scientific literature. Despite the large body of work in scientific QA, figure captioning, and other…

We introduce a new dataset, MELINDA, for Multimodal biomEdicaL experImeNt methoD clAssification. The dataset is collected in a fully automated distant supervision manner, where the labels are obtained from an existing curated database, and…

计算与语言 · 计算机科学 2020-12-18 Te-Lin Wu , Shikhar Singh , Sayan Paul , Gully Burns , Nanyun Peng