中文
相关论文

相关论文: MICo-150K: A Comprehensive Dataset Advancing Multi…

200 篇论文

Multiple instance learning (MIL) has shown significant promise in histopathology whole slide image (WSI) analysis for cancer diagnosis and prognosis. However, the inherent spatial heterogeneity of WSIs presents critical challenges, as…

计算机视觉与模式识别 · 计算机科学 2026-01-13 Junjian Li , Jin Liu , Hulin Kuang , Hailin Yue , Mengshen He , Jianxin Wang

Image generation has witnessed significant advancements in the past few years. However, evaluating the performance of image generation models remains a formidable challenge. In this paper, we propose ICE-Bench, a unified and comprehensive…

计算机视觉与模式识别 · 计算机科学 2025-08-19 Yulin Pan , Xiangteng He , Chaojie Mao , Zhen Han , Zeyinzi Jiang , Jingfeng Zhang , Yu Liu

We introduce 3D-COCO, an extension of the original MS-COCO dataset providing 3D models and 2D-3D alignment annotations. 3D-COCO was designed to achieve computer vision tasks such as 3D reconstruction or image detection configurable with…

计算机视觉与模式识别 · 计算机科学 2024-07-17 Maxence Bideaux , Alice Phe , Mohamed Chaouch , Bertrand Luvison , Quoc-Cuong Pham

How well can Multimodal Large Language Models (MLLMs) understand composite images? Composite images (CIs) are synthetic visuals created by merging multiple visual elements, such as charts, posters, or screenshots, rather than being captured…

计算机视觉与模式识别 · 计算机科学 2024-12-09 Xiaohui Chen , Satya Narayan Shukla , Mahmoud Azab , Aashu Singh , Qifan Wang , David Yang , ShengYun Peng , Hanchao Yu , Shen Yan , Xuewen Zhang , Baosheng He

We present MIX'EM, a novel solution for unsupervised image classification. MIX'EM generates representations that by themselves are sufficient to drive a general-purpose clustering algorithm to deliver high-quality classification. This is…

计算机视觉与模式识别 · 计算机科学 2020-10-06 Ali Varamesh , Tinne Tuytelaars

Recent advances in visual generative models have enabled high-fidelity image editing guided by human instructions. However, these models often struggle with complex instructions involving combinatorial editing operations or inter-step…

计算机视觉与模式识别 · 计算机科学 2026-05-12 Zilai Zeng , Mingdeng Cao , Zijie Li , Xiaochen Lian , Yichun Shi , Peihao Zhu , Chen Sun , Peng Wang

The performance of unified multimodal models for image generation and editing is fundamentally constrained by the quality and comprehensiveness of their training data. While existing datasets have covered basic tasks like style transfer and…

计算机视觉与模式识别 · 计算机科学 2025-09-30 Zhihong Chen , Xuehai Bai , Yang Shi , Chaoyou Fu , Huanyu Zhang , Haotian Wang , Xiaoyan Sun , Zhang Zhang , Liang Wang , Yuanxing Zhang , Pengfei Wan , Yi-Fan Zhang

The integration of neural-network-based systems into clinical practice is limited by challenges related to domain generalization and robustness. The computer vision community established benchmarks such as ImageNet-C as a fundamental…

图像与视频处理 · 电气工程与系统科学 2024-07-24 Francesco Di Salvo , Sebastian Doerrich , Christian Ledig

We introduce MuirBench, a comprehensive benchmark that focuses on robust multi-image understanding capabilities of multimodal LLMs. MuirBench consists of 12 diverse multi-image tasks (e.g., scene understanding, ordering) that involve 10…

Recent advancements in image quality assessment (IQA), driven by sophisticated deep neural network designs, have significantly improved the ability to approach human perceptions. However, most existing methods are obsessed with fitting the…

计算机视觉与模式识别 · 计算机科学 2025-08-26 Shunyu Yao , Ming Liu , Zhilu Zhang , Zhaolin Wan , Zhilong Ji , Jinfeng Bai , Wangmeng Zuo

Existing Multimodal Large Language Models (MLLMs) suffer from increased inference costs due to the additional vision tokens introduced by image inputs. In this work, we propose Visual Consistency Learning (ViCO), a novel training algorithm…

计算机视觉与模式识别 · 计算机科学 2025-12-16 Long Cui , Weiyun Wang , Jie Shao , Zichen Wen , Gen Luo , Linfeng Zhang , Yanting Zhang , Yu Qiao , Wenhai Wang

Training visual embeddings with labeled data supervision has been the de facto setup for representation learning in computer vision. Inspired by recent success of adopting masked image modeling (MIM) in self-supervised representation…

计算机视觉与模式识别 · 计算机科学 2023-12-05 Kaifeng Chen , Daniel Salz , Huiwen Chang , Kihyuk Sohn , Dilip Krishnan , Mojtaba Seyedhosseini

Recent advancements in image customization exhibit a wide range of application prospects due to stronger customization capabilities. However, since we humans are more sensitive to faces, a significant challenge remains in preserving…

计算机视觉与模式识别 · 计算机科学 2025-09-09 Yufeng Cheng , Wenxu Wu , Shaojin Wu , Mengqi Huang , Fei Ding , Qian He

We introduce the Multi-Instance Generation (MIG) task, which focuses on generating multiple instances within a single image, each accurately placed at predefined positions with attributes such as category, color, and shape, strictly…

计算机视觉与模式识别 · 计算机科学 2025-05-06 Dewei Zhou , You Li , Fan Ma , Zongxin Yang , Yi Yang

The automated evaluation of cognitive status utilizing multimedia technologies presents a promising frontier in early dementia diagnosis. However, the development of robust machine learning models for cognitive impairment detection is…

数据库 · 计算机科学 2026-04-03 Liuyu Wu , Rui Feng , Jie Li , Wentao Xiang , Yi Zhang , Yin Cao , Siyang Song , Xiao Gu , Jianqing Li , Wei Wang

Multi-view clustering can explore consistent information from different views to guide clustering. Most existing works focus on pursuing shallow consistency in the feature space and integrating the information of multiple views into a…

机器学习 · 计算机科学 2023-05-18 Yiyang Zhou , Qinghai Zheng , Wenbiao Yan , Yifei Wang , Pengcheng Shi , Jihua Zhu

Macro photography (MP) is a specialized field of photography that captures objects at an extremely close range, revealing tiny details. Although an accurate macro photography image quality assessment (MPIQA) metric can benefit macro…

计算机视觉与模式识别 · 计算机科学 2025-12-16 Jiashuo Chang , Zhengyi Li , Jianxun Lou , Zhen Qiu , Hanhe Lin

Recently, extensive research on image customization (e.g., identity, subject, style, background, etc.) demonstrates strong customization capabilities in large-scale generative models. However, most approaches are designed for specific…

We present Pippo, a generative model capable of producing 1K resolution dense turnaround videos of a person from a single casually clicked photo. Pippo is a multi-view diffusion transformer and does not require any additional inputs - e.g.,…

计算机视觉与模式识别 · 计算机科学 2025-02-12 Yash Kant , Ethan Weber , Jin Kyu Kim , Rawal Khirodkar , Su Zhaoen , Julieta Martinez , Igor Gilitschenski , Shunsuke Saito , Timur Bagautdinov

Generating high-fidelity images of humans with fine-grained control over attributes such as hairstyle and clothing remains a core challenge in personalized text-to-image synthesis. While prior methods emphasize identity preservation from a…

计算机视觉与模式识别 · 计算机科学 2025-10-17 Guocheng Gordon Qian , Daniil Ostashev , Egor Nemchinov , Avihay Assouline , Sergey Tulyakov , Kuan-Chieh Jackson Wang , Kfir Aberman