English
Related papers

Related papers: Learning Spatial-Temporal Coherent Correlations fo…

200 papers

Multimodal Emotion Recognition in Conversations (MERC) is a crucial task for understanding human interactions, where multimodal approaches integrating language, facial expressions, and vocal tone have achieved significant progress. However,…

Machine Learning · Computer Science 2026-05-22 Phuong-Anh Nguyen , The-Son Le , Duc-Trong Le , Cam-Van Thi Nguyen

Lip-reading aims to recognize speech content from videos via visual analysis of speakers' lip movements. This is a challenging task due to the existence of homophemes-words which involve identical or highly similar lip movements, as well as…

Computer Vision and Pattern Recognition · Computer Science 2019-09-04 Chenhao Wang

In contrastive self-supervised learning, the common way to learn discriminative representation is to pull different augmented "views" of the same image closer while pushing all other images further apart, which has been proven to be…

Computer Vision and Pattern Recognition · Computer Science 2022-12-14 Kaiyou Song , Shan Zhang , Zihao An , Zimeng Luo , Tong Wang , Jin Xie

Accurate molecular property prediction requires integrating complementary information from molecular structure and chemical semantics. In this work, we propose LGM-CL, a local-global multimodal contrastive learning framework that jointly…

Machine Learning · Computer Science 2026-02-02 Xiayu Liu , Zhengyi Lu , Yunhong Liao , Chan Fan , Hou-biao Li

Facial emotional recognition is one of the essential tools used by recognition psychology to diagnose patients. Face and facial emotional recognition are areas where machine learning is excelling. Facial Emotion Recognition in an…

Computer Vision and Pattern Recognition · Computer Science 2022-08-08 Nitesh Banskota , Abeer Alsadoon , P. W. C. Prasad , Ahmed Dawoud , Tarik A. Rashid , Omar Hisham Alsadoon

Understanding human affective behaviour, especially in the dynamics of real-world settings, requires Facial Expression Recognition (FER) models to continuously adapt to individual differences in user expression, contextual attributions, and…

Computer Vision and Pattern Recognition · Computer Science 2023-05-12 Nikhil Churamani , Tolga Dimlioglu , German I. Parisi , Hatice Gunes

Time-Scale Modification (TSM) of speech aims to alter the playback rate of audio without changing its pitch. While classical methods like Waveform Similarity-based Overlap-Add (WSOLA) provide strong baselines, they often introduce artifacts…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-06 Dyah A. M. G. Wisnu , Ryandhimas E. Zezario , Stefano Rini , Fo-Rui Li , Yan-Tsung Peng , Hsin-Min Wang , Yu Tsao

Computer-supported simulation enables a practical alternative for medical training purposes. This study investigates the co-occurrence of facial-recognition-derived emotions and socially shared regulation of learning (SSRL) interactions in…

Human-Computer Interaction · Computer Science 2025-10-21 Xiaoshan Huang , Tianlong Zhong , Haolun Wu , Yeyu Wang , Ethan Churchill , Xue Liu , David Williamson Shaffer

Video editing models have advanced significantly, but evaluating their performance remains challenging. Traditional metrics, such as CLIP text and image scores, often fall short: text scores are limited by inadequate training data and…

Computer Vision and Pattern Recognition · Computer Science 2025-01-14 Varun Biyyala , Bharat Chanderprakash Kathuria , Jialu Li , Youshan Zhang

Data-driven soft sensors have been widely applied in complex industrial processes. However, the interpretable spatio-temporal features extraction by soft sensors remains a challenge. In this light, this work introduces a novel method termed…

Systems and Control · Electrical Eng. & Systems 2025-11-04 Qianchao Wang , Peng Sha , Leena Heistrene , Yuxuan Ding , Yaping Du

Text-Based Person Search (TBPS) aims to retrieve target person images from a large-scale gallery using natural language descriptions, posing fundamental challenges in cross-modal representation learning. Existing methods often struggle to…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Jing Liu , Donglai Wei , Yang Liu , Sipeng Zhang , Tong Yang , Wei Zhou , Weiping Ding , Victor C. M. Leung

In text recognition, self-supervised pre-training emerges as a good solution to reduce dependence on expansive annotated real data. Previous studies primarily focus on local visual representation by leveraging mask image modeling or…

Computer Vision and Pattern Recognition · Computer Science 2024-05-14 Zuan Gao , Yuxin Wang , Yadong Qu , Boqiang Zhang , Zixiao Wang , Jianjun Xu , Hongtao Xie

The current state-of-the-art image-sentence retrieval methods implicitly align the visual-textual fragments, like regions in images and words in sentences, and adopt attention modules to highlight the relevance of cross-modal semantic…

Computer Vision and Pattern Recognition · Computer Science 2021-08-06 Xuri Ge , Fuhai Chen , Joemon M. Jose , Zhilong Ji , Zhongqin Wu , Xiao Liu

Speech-driven facial animation aims to synthesize lip-synchronized 3D talking faces following the given speech signal. Prior methods to this task mostly focus on pursuing realism with deterministic systems, yet characterizing the…

Computer Vision and Pattern Recognition · Computer Science 2024-10-01 Chunzhi Gu , Shigeru Kuriyama , Katsuya Hotta

Existing video-language pre-training methods primarily focus on instance-level alignment between video clips and captions via global contrastive learning but neglect rich fine-grained local information in both videos and text, which is of…

Computer Vision and Pattern Recognition · Computer Science 2024-09-10 Yuanhao Xiong , Long Zhao , Boqing Gong , Ming-Hsuan Yang , Florian Schroff , Ting Liu , Cho-Jui Hsieh , Liangzhe Yuan

In low-resource computing contexts, such as smartphones and other tiny devices, Both deep learning and machine learning are being used in a lot of identification systems. as authentication techniques. The transparent, contactless, and…

Computer Vision and Pattern Recognition · Computer Science 2023-05-02 Pangoth Santhosh Kumar , Garika Akshay

Steady-State Visual Evoked Potential (SSVEP) spellers are a promising communication tool for individuals with disabilities. This Brain-Computer Interface utilizes scalp potential data from (electroencephalography) EEG electrodes on a…

Human-Computer Interaction · Computer Science 2024-12-31 Joseph Zhang , Ruiming Zhang , Kipngeno Koech , David Hill , Kateryna Shapovalenko

Language-queried video actor segmentation aims to predict the pixel-level mask of the actor which performs the actions described by a natural language query in the target frames. Existing methods adopt 3D CNNs over the video clip as a…

Computer Vision and Pattern Recognition · Computer Science 2021-05-17 Tianrui Hui , Shaofei Huang , Si Liu , Zihan Ding , Guanbin Li , Wenguan Wang , Jizhong Han , Fei Wang

Dynamic facial expression recognition (DFER) infers emotions from the temporal evolution of expressions, unlike static facial expression recognition (SFER), which relies solely on a single snapshot. This temporal analysis provides richer…

Computer Vision and Pattern Recognition · Computer Science 2025-10-31 Yin Chen , Jia Li , Yu Zhang , Zhenzhen Hu , Shiguang Shan , Meng Wang , Richang Hong

Self-supervised learning has been widely used to obtain transferrable representations from unlabeled images. Especially, recent contrastive learning methods have shown impressive performances on downstream image classification tasks. While…

Computer Vision and Pattern Recognition · Computer Science 2021-04-29 Byungseok Roh , Wuhyun Shin , Ildoo Kim , Sungwoong Kim