English
Related papers

Related papers: Video Joint-Embedding Predictive Architectures for…

200 papers

Facial expressions of emotion are a major channel in our daily communications, and it has been subject of intense research in recent years. To automatically infer facial expressions, convolutional neural network based approaches has become…

Computer Vision and Pattern Recognition · Computer Science 2020-08-14 Bita Houshmand , Naimul Khan

Although state-of-the-art classifiers for facial expression recognition (FER) can achieve a high level of accuracy, they lack interpretability, an important feature for end-users. Experts typically associate spatial action units (AUs) from…

Computer Vision and Pattern Recognition · Computer Science 2024-10-04 Soufiane Belharbi , Marco Pedersoli , Alessandro Lameiras Koerich , Simon Bacon , Eric Granger

Utilizing large pre-trained models for specific tasks has yielded impressive results. However, fully fine-tuning these increasingly large models is becoming prohibitively resource-intensive. This has led to a focus on more…

Computer Vision and Pattern Recognition · Computer Science 2024-08-06 Shreyank N Gowda , Boyan Gao , David A. Clifton

Two competing paradigms exist for self-supervised learning of data representations. Joint Embedding Predictive Architecture (JEPA) is a class of architectures in which semantically similar inputs are encoded into representations that are…

Machine Learning · Computer Science 2024-07-08 Etai Littwin , Omid Saremi , Madhu Advani , Vimal Thilak , Preetum Nakkiran , Chen Huang , Joshua Susskind

This paper proposes a novel 4D Facial Expression Recognition (FER) method using Collaborative Cross-domain Dynamic Image Network (CCDN). Given a 4D data of face scans, we first compute its geometrical images, and then combine their…

Computer Vision and Pattern Recognition · Computer Science 2020-02-10 Muzammil Behzad , Nhat Vo , Xiaobai Li , Guoying Zhao

Critical obstacles in training classifiers to detect facial actions are the limited sizes of annotated video databases and the relatively low frequencies of occurrence of many actions. To address these problems, we propose an approach that…

Computer Vision and Pattern Recognition · Computer Science 2020-10-22 Koichiro Niinuma , Itir Onal Ertugrul , Jeffrey F Cohn , László A Jeni

Throughout the various ages, facial expressions have become one of the universal ways of non-verbal communication. The ability to recognize facial expressions would pave the path for many novel applications. Despite the success of…

Computer Vision and Pattern Recognition · Computer Science 2020-06-09 Raghu Vamshi. N , Bharathi Raja S

Dynamic facial expression recognition (DFER) is essential to the development of intelligent and empathetic machines. Prior efforts in this field mainly fall into supervised learning paradigm, which is severely restricted by the limited…

Computer Vision and Pattern Recognition · Computer Science 2023-08-09 Licai Sun , Zheng Lian , Bin Liu , Jianhua Tao

One of the most universal ways that people communicate is through facial expressions. In this paper, we take a deep dive, implementing multiple deep learning models for facial expression recognition (FER). Our goals are twofold: we aim not…

Computer Vision and Pattern Recognition · Computer Science 2020-04-27 Amil Khanzada , Charles Bai , Ferhat Turker Celepcikay

World models for partially observed environments must imagine multiple compatible hidden futures and steer between them under counterfactual actions. Joint Embedding Predictive Architectures (JEPAs) do this in latent space, but a…

Machine Learning · Computer Science 2026-05-26 Santosh Kumar Radha , Oktay Goktas

Language representation learning has emerged as a promising approach for sequential recommendation, thanks to its ability to learn generalizable representations. However, despite its advantages, this approach still struggles with data…

Information Retrieval · Computer Science 2025-08-08 Minh-Anh Nguyen , Dung D. Le

This study investigates the key characteristics and suitability of widely used Facial Expression Recognition (FER) datasets for training deep learning models. In the field of affective computing, FER is essential for interpreting human…

Computer Vision and Pattern Recognition · Computer Science 2025-03-27 F. Xavier Gaya-Morey , Cristina Manresa-Yee , Célia Martinie , Jose M. Buades-Rubio

This paper presents a novel visual-language model called DFER-CLIP, which is based on the CLIP model and designed for in-the-wild Dynamic Facial Expression Recognition (DFER). Specifically, the proposed DFER-CLIP consists of a visual part…

Computer Vision and Pattern Recognition · Computer Science 2024-11-27 Zengqun Zhao , Ioannis Patras

Multi-view facial expression recognition (FER) is a challenging task because the appearance of an expression varies in poses. To alleviate the influences of poses, recent methods either perform pose normalization or learn separate FER…

Computer Vision and Pattern Recognition · Computer Science 2019-05-27 Yuanyuan Liu , Jiyao Peng , Jiabei Zeng , Shiguang Shan

Foundation models have recently attracted significant attention for their impressive generalizability across diverse downstream tasks. However, these models are demonstrated to exhibit great limitations in representing high-frequency…

Image and Video Processing · Electrical Eng. & Systems 2025-04-18 Yuetan Chu , Yilan Zhang , Zhongyi Han , Changchun Yang , Longxi Zhou , Gongning Luo , Chao Huang , Xin Gao

Genomic Foundation Models (GFMs) typically rely on Masked Language Modeling (MLM) or Next-Token Prediction (NTP) to learn the "Laws of Nature". While effective at capturing local syntax, these generative paradigms prioritize token-level…

In this work we propose a novel joint training method for Visual Place Recognition (VPR), which simultaneously learns a global descriptor and a pair classifier for re-ranking. The pair classifier can predict whether a given pair of images…

Robotics · Computer Science 2025-03-04 Stephen Hausler , Peyman Moghadam

Modern Text-to-Image (T2I) generation increasingly relies on token-centric architectures that are trained with self-supervision, yet effectively fusing text with visual tokens remains a challenge. We propose \textbf{JEPA-T}, a unified…

Computer Vision and Pattern Recognition · Computer Science 2025-10-02 Siheng Wan , Zhengtao Yao , Zhengdao Li , Junhao Dong , Yanshu Li , Yikai Li , Linshan Li , Haoyan Xu , Yijiang Li , Zhikang Dong , Huacan Wang , Jifeng Shen

We propose WirelessJEPA, a novel wireless foundation model (WFM) that uses the Joint Embedding Predictive Architecture (JEPA). WirelessJEPA learns general-purpose representations directly from real-world multi-antenna IQ data by predicting…

Signal Processing · Electrical Eng. & Systems 2026-01-29 Viet Chu , Omar Mashaal , Hatem Abou-Zeid

Invariance-based and generative methods have shown a conspicuous performance for 3D self-supervised representation learning (SSRL). However, the former relies on hand-crafted data augmentations that introduce bias not universally applicable…

Computer Vision and Pattern Recognition · Computer Science 2024-09-25 Naiwen Hu , Haozhe Cheng , Yifan Xie , Shiqi Li , Jihua Zhu
‹ Prev 1 4 5 6 7 8 10 Next ›