English
Related papers

Related papers: MoCHA: Denoising Caption Supervision for Motion-Te…

200 papers

We introduce Quantized Language-Image Pretraining (QLIP), a visual tokenization method that combines state-of-the-art reconstruction quality with state-of-the-art zero-shot image understanding. QLIP trains a…

Computer Vision and Pattern Recognition · Computer Science 2025-02-10 Yue Zhao , Fuzhao Xue , Scott Reed , Linxi Fan , Yuke Zhu , Jan Kautz , Zhiding Yu , Philipp Krähenbühl , De-An Huang

Large-scale multi-modal contrastive pre-training has demonstrated great utility to learn transferable features for a range of downstream tasks by mapping multiple modalities into a shared embedding space. Typically, this has employed…

Computer Vision and Pattern Recognition · Computer Science 2022-07-27 Haoxuan You , Luowei Zhou , Bin Xiao , Noel Codella , Yu Cheng , Ruochen Xu , Shih-Fu Chang , Lu Yuan

Recent advances in vision-language models (VLMs) have demonstrated remarkable zero-shot capabilities, yet adapting these models to specialized domains remains a significant challenge. Building on recent theoretical insights suggesting that…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Pranav Mantini , Shishir K. Shah

Multimodal learning from document data has achieved great success lately as it allows to pre-train semantically meaningful features as a prior into a learnable downstream task. In this paper, we approach the document classification problem…

Computer Vision and Pattern Recognition · Computer Science 2023-05-12 Souhail Bakkali , Zuheng Ming , Mickael Coustaty , Marçal Rusiñol , Oriol Ramos Terrades

Computer vision tasks such as object detection and semantic/instance segmentation rely on the painstaking annotation of large training datasets. In this paper, we propose LocTex that takes advantage of the low-cost localized textual…

Computer Vision and Pattern Recognition · Computer Science 2021-08-27 Zhijian Liu , Simon Stent , Jie Li , John Gideon , Song Han

In many applications, we desire neural networks to exhibit invariance or equivariance to certain groups due to symmetries inherent in the data. Recently, frame-averaging methods emerged to be a unified framework for attaining symmetries…

Machine Learning · Computer Science 2024-11-05 George Ma , Yifei Wang , Derek Lim , Stefanie Jegelka , Yisen Wang

CLIP models perform remarkably well on zero-shot classification and retrieval tasks. But recent studies have shown that learnt representations in CLIP are not well suited for dense prediction tasks like object detection, semantic…

Computer Vision and Pattern Recognition · Computer Science 2024-05-16 Pavan Kumar Anasosalu Vasu , Hadi Pouransari , Fartash Faghri , Oncel Tuzel

Generating accurate and coherent image captions in a continual learning setting remains a major challenge due to catastrophic forgetting and the difficulty of aligning evolving visual concepts with language over time. In this work, we…

Computer Vision and Pattern Recognition · Computer Science 2025-11-19 Bertram Taetz , Gal Bordelius

Automatically describing a video with natural language is regarded as a fundamental challenge in computer vision. The problem nevertheless is not trivial especially when a video contains multiple events to be worthy of mention, which often…

Computer Vision and Pattern Recognition · Computer Science 2018-04-24 Yehao Li , Ting Yao , Yingwei Pan , Hongyang Chao , Tao Mei

Face Anti-Spoofing (FAS) is essential for the security of facial recognition systems in diverse scenarios such as payment processing and surveillance. Current multimodal FAS methods often struggle with effective generalization, mainly due…

Computer Vision and Pattern Recognition · Computer Science 2025-05-15 Yingjie Ma , Xun Lin , Zitong Yu , Xin Liu , Xiaochen Yuan , Weicheng Xie , Linlin Shen

Bilingual lexicon induction, translating words from the source language to the target language, is a long-standing natural language processing task. Recent endeavors prove that it is promising to employ images as pivot to learn the lexicon…

Computation and Language · Computer Science 2019-06-04 Shizhe Chen , Qin Jin , Alexander Hauptmann

Multi-modal semantic understanding requires integrating information from different modalities to extract users' real intention behind words. Most previous work applies a dual-encoder structure to separately encode image and text, but fails…

Computation and Language · Computer Science 2024-03-12 Ming Zhang , Ke Chang , Yunfang Wu

Hashing techniques have been applied broadly in retrieval tasks due to their low storage requirements and high speed of processing. Many hashing methods based on a single view have been extensively studied for information retrieval.…

Machine Learning · Computer Science 2020-01-07 Jun Yu , Xiao-Jun Wu , Josef Kittler

We present MoCA-Video, a training-free framework for semantic mixing in videos. Operating in the latent space of a frozen video diffusion model, MoCA-Video utilizes class-agnostic segmentation with diagonal denoising scheduler to localize…

Computer Vision and Pattern Recognition · Computer Science 2025-12-15 Tong Zhang , Juan C Leon Alcazar , Victor Escorcia , Bernard Ghanem

Subject-driven text-to-image diffusion models empower users to tailor the model to new concepts absent in the pre-training dataset using a few sample images. However, prevalent subject-driven models primarily rely on single-concept input…

Computer Vision and Pattern Recognition · Computer Science 2024-02-16 Junjie Shentu , Matthew Watson , Noura Al Moubayed

Motion, as the most distinct phenomenon in a video to involve the changes over time, has been unique and critical to the development of video representation learning. In this paper, we ask the question: how important is the motion…

Computer Vision and Pattern Recognition · Computer Science 2022-01-12 Rui Li , Yiheng Zhang , Zhaofan Qiu , Ting Yao , Dong Liu , Tao Mei

Training data is at the core of any successful text-to-image models. The quality and descriptiveness of image text are crucial to a model's performance. Given the noisiness and inconsistency in web-scraped datasets, recent works shifted…

Computer Vision and Pattern Recognition · Computer Science 2025-06-23 Manuel Brack , Sudeep Katakol , Felix Friedrich , Patrick Schramowski , Hareesh Ravi , Kristian Kersting , Ajinkya Kale

This paper learns multi-modal embeddings from text, audio, and video views/modes of data in order to improve upon down-stream sentiment classification. The experimental framework also allows investigation of the relative contributions of…

Information Retrieval · Computer Science 2019-07-23 Zhongkai Sun , Prathusha K Sarma , William Sethares , Erik P. Bucy

In this paper, we present Change3D, a framework that reconceptualizes the change detection and captioning tasks through video modeling. Recent methods have achieved remarkable success by regarding each pair of bi-temporal images as separate…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Duowang Zhu , Xiaohu Huang , Haiyan Huang , Hao Zhou , Zhenfeng Shao

Diffusion models have recently advanced human motion generation, producing realistic and diverse animations from textual prompts. However, adapting these models to unseen actions or styles typically requires additional motion capture data…

Computer Vision and Pattern Recognition · Computer Science 2025-10-09 Girolamo Macaluso , Lorenzo Mandelli , Mirko Bicchierai , Stefano Berretti , Andrew D. Bagdanov