English
Related papers

Related papers: MICo-150K: A Comprehensive Dataset Advancing Multi…

200 papers

Pretraining robust vision or multimodal foundation models (e.g., CLIP) relies on large-scale datasets that may be noisy, potentially misaligned, and have long-tail distributions. Previous works have shown promising results in augmenting…

Computer Vision and Pattern Recognition · Computer Science 2024-10-17 Qingqing Cao , Mahyar Najibi , Sachin Mehta

We propose a novel method for unsupervised semantic image segmentation based on mutual information maximization between local and global high-level image features. The core idea of our work is to leverage recent progress in self-supervised…

Computer Vision and Pattern Recognition · Computer Science 2021-10-08 Robert Harb , Patrick Knöbelreiter

Recent text-to-image diffusion models are able to learn and synthesize images containing novel, personalized concepts (e.g., their own pets or specific items) with just a few examples for training. This paper tackles two interconnected…

Computer Vision and Pattern Recognition · Computer Science 2024-02-26 Chun-Hsiao Yeh , Ta-Ying Cheng , He-Yen Hsieh , Chuan-En Lin , Yi Ma , Andrew Markham , Niki Trigoni , H. T. Kung , Yubei Chen

We present MILO (Metric for Image- and Latent-space Optimization), a lightweight, multiscale, perceptual metric for full-reference image quality assessment (FR-IQA). MILO is trained using pseudo-MOS (Mean Opinion Score) supervision, in…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Uğur Çoğalan , Mojtaba Bemana , Karol Myszkowski , Hans-Peter Seidel , Colin Groth

The recent surge in popularity of Nano-Banana and Seedream 4.0 underscores the community's strong interest in multi-image composition tasks. Compared to single-image editing, multi-image composition presents significantly greater challenges…

Computer Vision and Pattern Recognition · Computer Science 2026-01-23 Hongyang Wei , Hongbo Liu , Zidong Wang , Yi Peng , Baixin Xu , Size Wu , Xuying Zhang , Xianglong He , Zexiang Liu , Peiyu Wang , Xuchen Song , Yangguang Li , Yang Liu , Yahui Zhou

Foundation models in computer vision have demonstrated exceptional performance in zero-shot and few-shot tasks by extracting multi-purpose features from large-scale datasets through self-supervised pre-training methods. However, these…

Computer Vision and Pattern Recognition · Computer Science 2024-10-29 Yingjun Shen , Haizhao Dai , Qihe Chen , Yan Zeng , Jiakai Zhang , Yuan Pei , Jingyi Yu

Modern video diffusion models excel at appearance synthesis but still struggle with physical consistency: objects drift, collisions lack realistic rebound, and material responses seldom match their underlying properties. We present PhyCo, a…

Computer Vision and Pattern Recognition · Computer Science 2026-05-01 Sriram Narayanan , Ziyu Jiang , Srinivasa Narasimhan , Manmohan Chandraker

Multi-view clustering aims at exploiting information from multiple heterogeneous views to promote clustering. Most previous works search for only one optimal clustering based on the predefined clustering criterion, but devising such a…

Machine Learning · Computer Science 2020-10-06 Shaowei Wei , Jun Wang , Guoxian Yu , Carlotta Domeniconi , Xiangliang Zhang

Learning generic skills for humanoid robots interacting with 3D scenes by mimicking human data is a key research challenge with significant implications for robotics and real-world applications. However, existing methodologies and…

Robotics · Computer Science 2024-12-24 Yun Liu , Bowen Yang , Licheng Zhong , He Wang , Li Yi

With the continuous advancement of large language models (LLMs), it is essential to create new benchmarks to effectively evaluate their expanding capabilities and identify areas for improvement. This work focuses on multi-image reasoning,…

Computer Vision and Pattern Recognition · Computer Science 2024-06-14 Mehran Kazemi , Nishanth Dikkala , Ankit Anand , Petar Devic , Ishita Dasgupta , Fangyu Liu , Bahare Fatemi , Pranjal Awasthi , Dee Guo , Sreenivas Gollapudi , Ahmed Qureshi

Deep image matting methods have achieved increasingly better results on benchmarks (e.g., Composition-1k/alphamatting.com). However, the robustness, including robustness to trimaps and generalization to images from different domains, is…

Computer Vision and Pattern Recognition · Computer Science 2022-01-19 Yutong Dai , Brian Price , He Zhang , Chunhua Shen

Composed Image Retrieval (CIR) retrieves target images using a multi-modal query that combines a reference image with text describing desired modifications. The primary challenge is effectively fusing this visual and textual information.…

Computer Vision and Pattern Recognition · Computer Science 2025-04-16 Chaoyang Wang , Zeyu Zhang , Long Teng , Zijun Li , Shichao Kan

Image denoising is a fundamental problem in computer vision and medical imaging. However, real-world images are often degraded by structured noise with strong anisotropic correlations that existing methods struggle to remove. Most…

Image and Video Processing · Electrical Eng. & Systems 2025-10-03 Jianxu Wang , Ge Wang

Large Language Models (LLMs) and Large Multimodal Models (LMMs) demonstrate impressive problem-solving skills in many tasks and domains. However, their ability to reason with complex images in academic domains has not been systematically…

Multimedia · Computer Science 2025-10-01 Chenghao Ma , Haihong E. , Junpeng Ding , Jun Zhang , Ziyan Ma , Huang Qing , Bofei Gao , Liang Chen , Yifan Zhu , Meina Song

Contrastive Language-Image Pretraining (CLIP) has demonstrated great zero-shot performance for matching images and text. However, it is still challenging to adapt vision-lanaguage pretrained models like CLIP to compositional image and text…

Computer Vision and Pattern Recognition · Computer Science 2024-04-16 Kenan Jiang , Xuehai He , Ruize Xu , Xin Eric Wang

Many people are interested in taking astonishing photos and sharing with others. Emerging hightech hardware and software facilitate ubiquitousness and functionality of digital photography. Because composition matters in photography,…

Computer Vision and Pattern Recognition · Computer Science 2018-11-13 Farshid Farhat , Mohammad Mahdi Kamani , James Z. Wang

Modeling complex systems is a time-consuming, difficult and fragmented task, often requiring the analyst to work with disparate data, a variety of models, and expert knowledge across a diverse set of domains. Applying a user-centered design…

Human-Computer Interaction · Computer Science 2021-09-09 Fahd Husain , Pascale Proulx , Meng-Wei Chang , Rosa Romero-Gomez , Holland Vasquez

While the field of medical image analysis has undergone a transformative shift with the integration of machine learning techniques, the main challenge of these techniques is often the scarcity of large, diverse, and well-annotated datasets.…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Stefano Woerner , Arthur Jaques , Christian F. Baumgartner

While text-to-image generation has achieved unprecedented fidelity, the vast majority of existing models function fundamentally as static text-to-pixel decoders. Consequently, they often fail to grasp implicit user intentions. Although…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Jun He , Junyan Ye , Zilong Huang , Dongzhi Jiang , Chenjue Zhang , Leqi Zhu , Renrui Zhang , Xiang Zhang , Weijia Li

Recent proprietary models such as Sora2 demonstrate promising progress in generating multi-shot videos conditioned on multiple reference characters. However, academic research on this problem remains limited. We study this task and identify…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Binyuan Huang , Yuning Lu , Weinan Jia , Hualiang Wang , Mu Liu , Daiqing Yang
‹ Prev 1 8 9 10 Next ›