中文
相关论文

相关论文: FungiTastic: A multi-modal dataset and benchmark f…

200 篇论文

Acquiring aligned visuo-tactile datasets is slow and costly, requiring specialised hardware and large-scale data collection. Synthetic generation is promising, but prior methods are typically single-modality, limiting cross-modal learning.…

计算机视觉与模式识别 · 计算机科学 2026-02-24 Sirine Bhouri , Lan Wei , Jian-Qing Zheng , Dandan Zhang

Foundation language models obtain the instruction-following ability through supervised fine-tuning (SFT). Diversity and complexity are considered critical factors of a successful SFT dataset, while their definitions remain obscure and lack…

计算与语言 · 计算机科学 2023-08-16 Keming Lu , Hongyi Yuan , Zheng Yuan , Runji Lin , Junyang Lin , Chuanqi Tan , Chang Zhou , Jingren Zhou

Multimodal image matching seeks pixel-level correspondences between images of different modalities, crucial for cross-modal perception, fusion and analysis. However, the significant appearance differences between modalities make this task…

计算机视觉与模式识别 · 计算机科学 2025-09-22 Meng Yang , Fan Fan , Zizhuo Li , Songchu Deng , Yong Ma , Jiayi Ma

Better understanding and modelling of building interiors and the emergence of more impressive AR/VR technology has brought up the need for automatic parsing of floorplan images. However, there is a clear lack of representative datasets to…

计算机视觉与模式识别 · 计算机科学 2019-04-04 Ahti Kalervo , Juha Ylioinas , Markus Häikiö , Antti Karhu , Juho Kannala

Large multimodal models (LMMs) have achieved impressive progress in vision-language understanding, yet they face limitations in real-world applications requiring complex reasoning over a large number of images. Existing benchmarks for…

计算机视觉与模式识别 · 计算机科学 2024-12-09 Jun Chen , Dannong Xu , Junjie Fei , Chun-Mei Feng , Mohamed Elhoseiny

Progress in remote PhotoPlethysmoGraphy (rPPG) is limited by the critical issues of existing publicly available datasets: small size, privacy concerns with facial videos, and lack of diversity in conditions. The paper introduces a novel…

计算机视觉与模式识别 · 计算机科学 2025-10-28 Konstantin Egorov , Stepan Botman , Pavel Blinov , Galina Zubkova , Anton Ivaschenko , Alexander Kolsanov , Andrey Savchenko

Remote photoplethysmography (rPPG) emerges as a promising method for non-invasive, convenient measurement of vital signs, utilizing the widespread presence of cameras. Despite advancements, existing datasets fall short in terms of size and…

计算机视觉与模式识别 · 计算机科学 2024-04-09 Jiankai Tang , Xinyi Li , Jiacheng Liu , Xiyuxing Zhang , Zeyu Wang , Yuntao Wang

Geo-spatial analysis of our world benefits from a multimodal approach, as every single geographic location can be described in numerous ways (images from various viewpoints, textual descriptions, geographic coordinates, etc.). Current…

计算机视觉与模式识别 · 计算机科学 2026-04-29 Oskar Kristoffersen , Alba Reinders Sánchez , Morten Rieger Hannemose , Anders Bjorholm Dahl , Dim P. Papadopoulos

We propose a unified cross-domain transfer learning framework that leverages knowledge from multiple heterogeneous medical imaging datasets to improve performance across segmentation, classification, and object detection tasks. Our approach…

计算机视觉与模式识别 · 计算机科学 2026-05-05 Ceausescu Ciprian-Mihai , Anghelina Ion-Marian , Alexe Dumitru-Bogdan

We present Urban-ImageNet, a large-scale multi-modal dataset and evaluation benchmark for urban space perception from user-generated social media imagery. The corpus contains over 2 Million public social media images and paired textual…

计算机视觉与模式识别 · 计算机科学 2026-05-12 Yiwei Ou , Chung Ching Cheung , Jun Yang Ang , Xiaobin Ren , Ronggui Sun , Guansong Gao , Kaiqi Zhao , Manfredo Manfredini

Identification of fossil species is crucial to evolutionary studies. Recent advances from deep learning have shown promising prospects in fossil image identification. However, the quantity and quality of labeled fossil images are often…

计算机视觉与模式识别 · 计算机科学 2024-02-05 Chengbin Hou , Xinyu Lin , Hanhui Huang , Sheng Xu , Junxuan Fan , Yukun Shi , Hairong Lv

Robust mammography registration is essential for clinical applications like tracking disease progression and monitoring longitudinal changes in breast tissue. However, progress has been limited by the absence of public datasets and…

计算机视觉与模式识别 · 计算机科学 2025-12-22 Svetlana Krasnova , Emiliya Starikova , Ilia Naletov , Andrey Krylov , Dmitry Sorokin

Fine-grained visual classification (FGVC) requires distinguishing between visually similar categories through subtle, localized features - a task that remains challenging due to high intra-class variability and limited inter-class…

计算机视觉与模式识别 · 计算机科学 2025-09-22 Johann Schmidt , Sebastian Stober , Joachim Denzler , Paul Bodesheim

We present Flat'n'Fold, a novel large-scale dataset for garment manipulation that addresses critical gaps in existing datasets. Comprising 1,212 human and 887 robot demonstrations of flattening and folding 44 unique garments across 8…

机器人学 · 计算机科学 2024-09-30 Lipeng Zhuang , Shiyu Fan , Yingdong Ru , Florent Audonnet , Paul Henderson , Gerardo Aragon-Camarasa

Fine-grained image classification (FGIC) is a challenging task in computer vision for due to small visual differences among inter-subcategories, but, large intra-class variations. Deep learning methods have achieved remarkable success in…

计算机视觉与模式识别 · 计算机科学 2024-04-02 Asish Bera , Debotosh Bhattacharjee , Mita Nasipuri

Understanding clothes from a single image has strong commercial and cultural impacts on modern societies. However, this task remains a challenging computer vision problem due to wide variations in the appearance, style, brand and layering…

计算机视觉与模式识别 · 计算机科学 2019-04-11 Shuai Zheng , Fan Yang , M. Hadi Kiapour , Robinson Piramuthu

Integrating various data modalities brings valuable insights into underlying phenomena. Multimodal factor analysis (FA) uncovers shared axes of variation underlying different simple data modalities, where each sample is represented by a…

机器学习 · 计算机科学 2025-04-29 Małgorzata Łazęcka , Ewa Szczurek

Virtual try-on (VTON) has advanced single-garment visualization, yet real-world fashion centers on full outfits with multiple garments, accessories, fine-grained categories, layering, and diverse styling, remaining beyond current VTON…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Junyao Hu , Zhongwei Cheng , Waikeung Wong , Xingxing Zou

Recent advances in training vision-language models have demonstrated unprecedented robustness and transfer learning effectiveness; however, standard computer vision datasets are image-only, and therefore not well adapted to such training…

计算机视觉与模式识别 · 计算机科学 2023-02-22 Andre Nakkab , Benjamin Feuer , Chinmay Hegde
‹ 上一页 1 8 9 10 下一页 ›