English
Related papers

Related papers: 15M Multimodal Facial Image-Text Dataset

200 papers

Multimodal Large Language Models (mLLMs) are trained on a large amount of text-image data. While most mLLMs are trained on caption-like data only, Alayrac et al. (2022) showed that additionally training them on interleaved sequences of text…

Computation and Language · Computer Science 2025-05-30 Matthieu Futeral , Armel Zebaze , Pedro Ortiz Suarez , Julien Abadji , Rémi Lacroix , Cordelia Schmid , Rachel Bawden , Benoît Sagot

The rapid advancement of GenAI technology over the past few years has significantly contributed towards highly realistic deepfake content generation. Despite ongoing efforts, the research community still lacks a large-scale and reasoning…

Multimedia · Computer Science 2025-06-17 Parul Gupta , Shreya Ghosh , Tom Gedeon , Thanh-Toan Do , Abhinav Dhall

The advent of large Vision-Language Models (VLMs) has significantly advanced multimodal tasks, enabling more sophisticated and accurate reasoning across various applications, including image and video captioning, visual question answering,…

Computer Vision and Pattern Recognition · Computer Science 2024-11-26 Hang Hua , Qing Liu , Lingzhi Zhang , Jing Shi , Zhifei Zhang , Yilin Wang , Jianming Zhang , Jiebo Luo

Medical image segmentation remains challenging due to limited annotations for training, ambiguous anatomical features, and domain shifts. While vision-language models such as CLIP offer strong cross-modal representations, their potential…

Computer Vision and Pattern Recognition · Computer Science 2026-02-25 Taha Koleilat , Hojat Asgariandehkordi , Omid Nejati Manzari , Berardino Barile , Yiming Xiao , Hassan Rivaz

In this paper, we address a fundamental gap between pre-training and fine-tuning of deep neural networks: while pre-training has shifted from unimodal to multimodal learning with enhanced visual understanding, fine-tuning predominantly…

Computer Vision and Pattern Recognition · Computer Science 2026-01-30 Shohei Enomoto , Shin'ya Yamaguchi

Multimodal Large Language Models (MLLMs) show promise for medical applications, yet progress in dermatology lags due to limited training data, narrow task coverage, and lack of clinically-grounded supervision that mirrors expert diagnostic…

Computation and Language · Computer Science 2026-01-06 Jinghan Ru , Siyuan Yan , Yuguo Yin , Yuexian Zou , Zongyuan Ge

Recent progress in face detection (including keypoint detection), and recognition is mainly being driven by (i) deeper convolutional neural network architectures, and (ii) larger datasets. However, most of the large datasets are maintained…

Computer Vision and Pattern Recognition · Computer Science 2017-05-23 Ankan Bansal , Anirudh Nanduri , Carlos Castillo , Rajeev Ranjan , Rama Chellappa

Methods based on Contrastive Language-Image Pre-training (CLIP) are nowadays extensively used in support of vision-and-language tasks involving remote sensing data, such as cross-modal retrieval. The adaptation of CLIP to this specific…

Computer Vision and Pattern Recognition · Computer Science 2024-11-01 João Daniel Silva , Joao Magalhaes , Devis Tuia , Bruno Martins

Multimodal Large Language Models demonstrate strong performance on natural image understanding, yet exhibit limited capability in interpreting scientific images, including but not limited to schematic diagrams, experimental…

Computer Vision and Pattern Recognition · Computer Science 2026-02-17 Haoyi Tao , Chaozheng Huang , Nan Wang , Han Lyu , Linfeng Zhang , Guolin Ke , Xi Fang

Text-to-image models have rapidly evolved from casual creative tools to professional-grade systems, achieving unprecedented levels of image quality and realism. Yet, most models are trained to map short prompts into detailed images,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Eyal Gutflaish , Eliran Kachlon , Hezi Zisman , Tal Hacham , Nimrod Sarid , Alexander Visheratin , Saar Huberman , Gal Davidi , Guy Bukchin , Kfir Goldberg , Ron Mokady

Photorealistic avatars of human faces have come a long way in recent years, yet research along this area is limited by a lack of publicly available, high-quality datasets covering both, dense multi-view camera captures, and rich facial…

The past few years have witnessed renewed interest in NLP tasks at the interface between vision and language. One intensively-studied problem is that of automatically generating text from images. In this paper, we extend this problem to the…

Image representations are often evaluated through disjointed, task-specific protocols, leading to a fragmented understanding of model capabilities. For instance, it is unclear whether an image embedding model adept at clustering images is…

Computer Vision and Pattern Recognition · Computer Science 2025-04-15 Chenghao Xiao , Isaac Chung , Imene Kerboua , Jamie Stirling , Xin Zhang , Márton Kardos , Roman Solomatin , Noura Al Moubayed , Kenneth Enevoldsen , Niklas Muennighoff

Recent advancements in deep learning have significantly enhanced content-based retrieval methods, notably through models like CLIP that map images and texts into a shared embedding space. However, these methods often struggle with…

Computer Vision and Pattern Recognition · Computer Science 2025-02-11 Nicola Messina , Lucia Vadicamo , Leo Maltese , Claudio Gennaro

Existing text-to-image synthesis methods generally are only applicable to words in the training dataset. However, human faces are so variable to be described with limited words. So this paper proposes the first free-style text-to-face…

Computer Vision and Pattern Recognition · Computer Science 2022-03-30 Jianxin Sun , Qiyao Deng , Qi Li , Muyi Sun , Min Ren , Zhenan Sun

Instruction-based speech processing is becoming popular. Studies show that training with multiple tasks boosts performance, but collecting diverse, large-scale tasks and datasets is expensive. Thus, it is highly desirable to design a…

Computation and Language · Computer Science 2024-08-27 Chien-yu Huang , Min-Han Shih , Ke-Han Lu , Chi-Yuan Hsiao , Hung-yi Lee

Recently, face recognition in the wild has achieved remarkable success and one key engine is the increasing size of training data. For example, the largest face dataset, WebFace42M contains about 2 million identities and 42 million faces.…

Computer Vision and Pattern Recognition · Computer Science 2021-09-13 Yunze Chen , Junjie Huang , Jiagang Zhu , Zheng Zhu , Tian Yang , Guan Huang , Dalong Du

Text-guided human body animation has advanced rapidly, yet facial animation lags due to the scarcity of well-annotated, text-paired facial corpora. To close this gap, we leverage foundation generative models to synthesize a large, balanced…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Luchuan Song , Pinxin Liu , Haiyang Liu , Zhenchao Jin , Yolo Yunlong Tang , Zichong Xu , Susan Liang , Jing Bi , Jason J Corso , Chenliang Xu

Multimodal foundation models have significantly improved feature representation by integrating information from multiple modalities, making them highly suitable for a broader set of applications. However, the exploration of multimodal…

Computer Vision and Pattern Recognition · Computer Science 2025-04-15 Kaiwen Zheng , Xuri Ge , Junchen Fu , Jun Peng , Joemon M. Jose

Face benchmarks empower the research community to train and evaluate high-performance face recognition systems. In this paper, we contribute a new million-scale recognition benchmark, containing uncurated 4M identities/260M faces…

Computer Vision and Pattern Recognition · Computer Science 2022-04-22 Zheng Zhu , Guan Huang , Jiankang Deng , Yun Ye , Junjie Huang , Xinze Chen , Jiagang Zhu , Tian Yang , Dalong Du , Jiwen Lu , Jie Zhou
‹ Prev 1 3 4 5 6 7 10 Next ›