中文
相关论文

相关论文: Comprehensive Machine Learning Benchmarking for Fr…

200 篇论文

The advent of foundation models, particularly Vision-Language Models (VLMs) and Multi-modal Large Language Models (MLLMs), has redefined the frontiers of artificial intelligence, enabling remarkable generalization across diverse tasks with…

计算机视觉与模式识别 · 计算机科学 2025-06-02 Redwan Sony , Parisa Farmanifard , Hamzeh Alzwairy , Nitish Shukla , Arun Ross

Synthetic images are an option for augmenting limited medical imaging datasets to improve the performance of various machine learning models. A common metric for evaluating synthetic image quality is the Fr\'echet Inception Distance (FID)…

图像与视频处理 · 电气工程与系统科学 2025-07-30 Thomas Wallace , Ik Siong Heng , Senad Subasic , Chris Messenger

Tapered optical fibers (TFs), with diameters gradually reduced from hundreds of microns to the micron scale, offer key advantages over conventional flat optical fibers (FFs), including uniform illumination, efficient long-range signal…

光学 · 物理学 2026-02-03 Mingliang Xu , Fangyuan Li , Yuxin Leng , Ruxin Li , Fei He

We propose a high-performance fully convolutional neural network (FCN) for historical document segmentation that is designed to process a single page in one step. The advantage of this model beside its speed is its ability to directly learn…

计算机视觉与模式识别 · 计算机科学 2018-07-25 Christoph Wick , Frank Puppe

Blind Face Restoration (BFR) aims to construct a high-quality (HQ) face image from its corresponding low-quality (LQ) input. Recently, many BFR methods have been proposed and they have achieved remarkable success. However, these methods are…

计算机视觉与模式识别 · 计算机科学 2022-06-09 Puyang Zhang , Kaihao Zhang , Wenhan Luo , Changsheng Li , Guoren Wang

One of the most challenging problems in fingerprint recognition continues to be establishing the identity of a suspect associated with partial and smudgy fingerprints left at a crime scene (i.e., latent prints or fingermarks). Despite the…

计算机视觉与模式识别 · 计算机科学 2023-09-11 Steven A. Grosz , Anil K. Jain

Cutting out an object and estimating its opacity mask, known as image matting, is a key task in many image editing applications. Deep learning approaches have made significant progress by adapting the encoder-decoder architecture of…

计算机视觉与模式识别 · 计算机科学 2020-03-18 Marco Forte , François Pitié

Foundation models have revolutionized artificial intelligence by providing robust, versatile architectures pre-trained on large-scale datasets. However, adapting these massive models to specific downstream tasks requires fine-tuning, which…

机器学习 · 计算机科学 2025-05-01 Jieming Bian , Yuanzhe Peng , Lei Wang , Yin Huang , Jie Xu

Pose-invariant face recognition has become a challenging problem for modern AI-based face recognition systems. It aims at matching a profile face captured in the wild with a frontal face registered in a database. Existing methods perform…

计算机视觉与模式识别 · 计算机科学 2025-05-23 Nikolay Stanishev , Yuhang Lu , Touradj Ebrahimi

Recent deep monocular depth estimation approaches based on supervised regression have achieved remarkable performance. However, they require costly ground truth annotations during training. To cope with this issue, in this paper we present…

计算机视觉与模式识别 · 计算机科学 2019-09-18 Andrea Pilzer , Stéphane Lathuilière , Dan Xu , Mihai Marian Puscas , Elisa Ricci , Nicu Sebe

3D human pose and shape estimation (a.k.a. "human mesh recovery") has achieved substantial progress. Researchers mainly focus on the development of novel algorithms, while less attention has been paid to other critical factors involved.…

计算机视觉与模式识别 · 计算机科学 2022-09-22 Hui En Pang , Zhongang Cai , Lei Yang , Tianwei Zhang , Ziwei Liu

Image matting aims to obtain an alpha matte that separates foreground objects from the background accurately. Recently, trimap-free matting has been well studied because it requires only the original image without any extra input. Such…

计算机视觉与模式识别 · 计算机科学 2024-05-29 Leo Shan Wenzhang Zhou Grace Zhao

Wireless localization of permanent magnets enables occlusion-free guidance for medical interventions, yet its practical accuracy is fundamentally limited by two coupled challenges: the poor observability of conventional planar sensor arrays…

机器人学 · 计算机科学 2026-04-27 Wenxuan Xie , Yuelin Zhang , Qingpeng Ding , Jianghua Chen , Jiewen Tan , Jiwei Shan , Shing Shin Cheng

Neural networks that map between low dimensional spaces are ubiquitous in computer graphics and scientific computing; however, in their naive implementation, they are unable to learn high frequency information. We present a comprehensive…

计算机视觉与模式识别 · 计算机科学 2025-04-21 Samuel Audia , Soheil Feizi , Matthias Zwicker , Dinesh Manocha

We propose a new framework for processing Fringe Patterns (FP). Our novel approach builds upon the hypothesis that the denoising and normalisation of FPs can be learned by a deep neural network if enough pairs of corrupted and ideal FPs are…

图像与视频处理 · 电气工程与系统科学 2020-10-29 Alan Reyes-Figueroa , Mariano Rivera

Diffusion MRI microstructure fitting is nonconvex and often performed voxelwise, which limits fiber peak recovery in narrow crossings. This work introduces PRISM, a differentiable analysis-by-synthesis framework that fits an explicit…

计算机视觉与模式识别 · 计算机科学 2026-04-02 Mohamed Abouagour , Atharva Shah , Eleftherios Garyfallidis

Multimodal LLMs (MLLMs) are capable of performing complex data analysis, visual question answering, generation, and reasoning tasks. However, their ability to analyze biometric data is relatively underexplored. In this work, we investigate…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Ekta Gavas , Sudipta Banerjee , Chinmay Hegde , Nasir Memon

For low-level computer vision and image processing ML tasks, training on large datasets is critical for generalization. However, the standard practice of relying on real-world images primarily from the Internet comes with image quality,…

计算机视觉与模式识别 · 计算机科学 2022-12-09 Gyeongmin Choe , Beibei Du , Seonghyeon Nam , Xiaoyu Xiang , Bo Zhu , Rakesh Ranjan

Automated face recognition and identification softwares are becoming part of our daily life; it finds its abode not only with Facebook's auto photo tagging, Apple's iPhoto, Google's Picasa, Microsoft's Kinect, but also in Homeland Security…

计算机视觉与模式识别 · 计算机科学 2015-06-22 Poonam Yadav

Federated learning (FL) offers a privacy-preserving paradigm for collaborative medical image analysis without sharing raw data. However, the absence of standardized benchmarks for medical image segmentation hinders fair and comprehensive…

计算机视觉与模式识别 · 计算机科学 2026-03-13 Meilu Zhu , Zhiwei Wang , Axiu Mao , Yuxing Li , Xiaohan Xing , Yixuan Yuan , Edmund Y. Lam