中文
相关论文

相关论文: Rep3Net: An Approach Exploiting Multimodal Represe…

200 篇论文

Multimodal emotion recognition (MMER) systems typically outperform unimodal systems by leveraging the inter- and intra-modal relationships between, e.g., visual, textual, physiological, and auditory modalities. This paper proposes an MMER…

计算机视觉与模式识别 · 计算机科学 2024-04-23 Paul Waligora , Haseeb Aslam , Osama Zeeshan , Soufiane Belharbi , Alessandro Lameiras Koerich , Marco Pedersoli , Simon Bacon , Eric Granger

Learning effective numerical representations, or embeddings, of programs is a fundamental prerequisite for applying machine learning to automate and enhance compiler optimization. Prevailing paradigms, however, present a dilemma. Static…

机器学习 · 计算机科学 2025-10-16 Haolin Pan , Jinyuan Dong , Hongbin Zhang , Hongyu Lin , Mingjie Xing , Yanjun Wu

Predicting the performance of deep learning (DL) models, such as execution time and resource utilization, is crucial for Neural Architecture Search (NAS), DL cluster schedulers, and other technologies that advance deep learning. The…

性能 · 计算机科学 2025-02-04 Xinlong Zhao , Jiande Sun , Jia Zhang , Sujuan Hou , Shuai Li , Tong Liu , Ke Liu

This work focuses on complete 3D facial geometry prediction, including 3D facial alignment via 3D face modeling and face orientation estimation using the proposed multi-task, multi-modal, and multi-representation landmark refinement network…

计算机视觉与模式识别 · 计算机科学 2021-04-20 Cho-Ying Wu , Qiangeng Xu , Ulrich Neumann

We propose Int3DNet, a scene-aware network that predicts 3D intention areas directly from scene geometry and head-hand motion cues, enabling robust human intention prediction without explicit object-level perception. In Mixed Reality (MR),…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Taewook Ha , Woojin Cho , Dooyoung Kim , Woontack Woo

Prompt learning has emerged as a promising paradigm for adapting pre-trained vision-language models (VLMs) to few-shot whole slide image (WSI) classification by aligning visual features with textual representations, thereby reducing…

计算机视觉与模式识别 · 计算机科学 2025-10-01 Junjie Zhou , Wei Shao , Yagao Yue , Wei Mu , Peng Wan , Qi Zhu , Daoqiang Zhang

The extraction of chemical-gene relations plays a pivotal role in understanding the intricate interactions between chemical compounds and genes, with significant implications for drug discovery, disease understanding, and biomedical…

计算与语言 · 计算机科学 2026-02-05 Mai H. Nguyen , Shibani Likhite , Jiawei Tang , Darshini Mahendran , Bridget T. McInnes

The pretraining-finetuning paradigm has powered major advances in domains such as natural language processing and computer vision, with representative examples including masked language modeling and next-token prediction. In molecular…

机器学习 · 计算机科学 2025-10-21 Shaoheng Yan , Zian Li , Muhan Zhang

Loss function is crucial for model training and feature representation learning, conventional models usually regard facial attractiveness recognition task as a regression problem, and adopt MSE loss or Huber variant loss as supervision to…

多媒体 · 计算机科学 2020-10-22 Lu Xu , Jinhai Xiang

The research and applications of multimodal emotion recognition have become increasingly popular recently. However, multimodal emotion recognition faces the challenge of lack of data. To solve this problem, we propose to use transfer…

计算与语言 · 计算机科学 2022-07-13 Zihan Zhao , Yanfeng Wang , Yu Wang

Multi-scale features are essential for dense prediction tasks, such as object detection, instance segmentation, and semantic segmentation. The prevailing methods usually utilize a classification backbone to extract multi-scale features and…

计算机视觉与模式识别 · 计算机科学 2023-11-01 Gang Zhang , Ziyi Li , Chufeng Tang , Jianmin Li , Xiaolin Hu

In this study, we present a novel computational method for generating molecular fingerprints using multiparameter persistent homology (MPPH). This technique holds considerable significance for drug discovery and materials science, where…

机器学习 · 计算机科学 2023-12-14 Andac Demir , Francis Prael , Bulent Kiziltan

Grounded Multimodal Named Entity Recognition (GMNER) is an emerging information extraction (IE) task, aiming to simultaneously extract entity spans, types, and corresponding visual regions of entities from given sentence-image pairs data.…

信息检索 · 计算机科学 2025-01-28 Jielong Tang , Zhenxing Wang , Ziyang Gong , Jianxing Yu , Xiangwei Zhu , Jian Yin

The human face is a silent communicator, expressing emotions and thoughts through its facial expressions. With the advancements in computer vision in recent years, facial emotion recognition technology has made significant strides, enabling…

计算机视觉与模式识别 · 计算机科学 2024-11-26 Arnab Kumar Roy , Hemant Kumar Kathania , Adhitiya Sharma , Abhishek Dey , Md. Sarfaraj Alam Ansari

Visual complexity prediction is a fundamental problem in computer vision with applications in image compression, retrieval, and classification. Understanding what makes humans perceive an image as complex is also a long-standing question in…

计算机视觉与模式识别 · 计算机科学 2025-11-24 Jonathan Skaza , Parsa Madinei , Ziqi Wen , Miguel Eckstein

Micro-expression has emerged as a promising modality in affective computing due to its high objectivity in emotion detection. Despite the higher recognition accuracy provided by the deep learning models, there are still significant scope…

计算机视觉与模式识别 · 计算机科学 2022-01-25 Viswanatha Reddy Gajjala , Sai Prasanna Teja Reddy , Snehasis Mukherjee , Shiv Ram Dubey

Effective molecular representation learning is of great importance to facilitate molecular property prediction, which is a fundamental task for the drug and material industry. Recent advances in graph neural networks (GNNs) have shown great…

机器学习 · 计算机科学 2022-05-17 Xiaomin Fang , Lihang Liu , Jieqiong Lei , Donglong He , Shanzhuo Zhang , Jingbo Zhou , Fan Wang , Hua Wu , Haifeng Wang

Multimodal emotion recognition in conversation (MERC) requires representations that effectively integrate signals from multiple modalities. These signals include modality-specific cues, information shared across modalities, and interactions…

机器学习 · 计算机科学 2026-01-22 Anh-Tuan Mai , Cam-Van Thi Nguyen , Duc-Trong Le

Discovering gene-disease associations is crucial for understanding disease mechanisms, yet identifying these associations remains challenging due to the time and cost of biological experiments. Computational methods are increasingly vital…

人工智能 · 计算机科学 2025-01-15 Wentao Cui , Shoubo Li , Chen Fang , Qingqing Long , Chengrui Wang , Xuezhi Wang , Yuanchun Zhou

Face images appeared in multimedia applications, e.g., social networks and digital entertainment, usually exhibit dramatic pose, illumination, and expression variations, resulting in considerable performance degradation for traditional face…

计算机视觉与模式识别 · 计算机科学 2016-11-17 Changxing Ding , Dacheng Tao