中文
相关论文

相关论文: Most Important Person-guided Dual-branch Cross-Pat…

200 篇论文

The ability to monitor audience reactions is critical when delivering presentations. However, current videoconferencing platforms offer limited solutions to support this. This work leverages recent advances in affect sensing to capture and…

人机交互 · 计算机科学 2021-02-01 Prasanth Murali , Javier Hernandez , Daniel McDuff , Kael Rowan , Jina Suh , Mary Czerwinski

Multimodal Machine Translation (MMT) typically enhances text-only translation by incorporating aligned visual features. Despite the remarkable progress, state-of-the-art MMT approaches often rely on paired image-text inputs at inference and…

计算机视觉与模式识别 · 计算机科学 2025-12-05 Jie Wang , Zhendong Yang , Liansong Zong , Xiaobo Zhang , Dexian Wang , Ji Zhang

Depression, a prevalent and serious mental health issue, affects approximately 3.8\% of the global population. Despite the existence of effective treatments, over 75\% of individuals in low- and middle-income countries remain untreated,…

计算与语言 · 计算机科学 2024-07-19 Shengjie Li , Yinhao Xiao

Visual attention plays a critical role when our visual system executes active visual tasks by interacting with the physical scene. However, how to encode the visual object relationship in the psychological world of our brain deserves to be…

计算机视觉与模式识别 · 计算机科学 2025-11-26 Kai-Fu Yang , Yong-Jie Li

Human activity recognition in videos has been widely studied and has recently gained significant advances with deep learning approaches; however, it remains a challenging task. In this paper, we propose a novel framework that simultaneously…

计算机视觉与模式识别 · 计算机科学 2021-01-25 Dong-Gyu Lee , Seong-Whan Lee

In vision and linguistics; the main input modalities are facial expressions, speech patterns, and the words uttered. The issue with analysis of any one mode of expression (Visual, Verbal or Vocal) is that lot of contextual information can…

计算机视觉与模式识别 · 计算机科学 2022-11-29 Kunjal Panchal

We present a novel facial expression recognition network, called Distract your Attention Network (DAN). Our method is based on two key observations. Firstly, multiple classes share inherently similar underlying facial appearance, and their…

计算机视觉与模式识别 · 计算机科学 2023-05-23 Zhengyao Wen , Wenzhong Lin , Tao Wang , Ge Xu

Human behavior understanding requires looking at minute details in the large context of a scene containing multiple input modalities. It is necessary as it allows the design of more human-like machines. While transformer approaches have…

计算机视觉与模式识别 · 计算机科学 2022-12-09 Tanay Agrawal , Michal Balazia , Philipp Müller , François Brémond

Multi-modal crowd counting is a crucial task that uses multi-modal cues to estimate the number of people in crowded scenes. To overcome the gap between different modalities, we propose a modal emulation-based two-pass multi-modal…

计算机视觉与模式识别 · 计算机科学 2024-07-30 Chenhao Wang , Xiaopeng Hong , Zhiheng Ma , Yupeng Wei , Yabin Wang , Xiaopeng Fan

Group portrait editing is highly desirable since users constantly want to add a person, delete a person, or manipulate existing persons. It is also challenging due to the intricate dynamics of human interactions and the diverse gestures. In…

计算机视觉与模式识别 · 计算机科学 2024-09-24 Yuming Jiang , Nanxuan Zhao , Qing Liu , Krishna Kumar Singh , Shuai Yang , Chen Change Loy , Ziwei Liu

Crowd estimation is a very challenging problem. The most recent study tries to exploit auditory information to aid the visual models, however, the performance is limited due to the lack of an effective approach for feature extraction and…

计算机视觉与模式识别 · 计算机科学 2021-09-07 Usman Sajid , Xiangyu Chen , Hasan Sajid , Taejoon Kim , Guanghui Wang

Speech emotion recognition is crucial to human-computer interaction. The temporal regions that represent different emotions scatter in different parts of the speech locally. Moreover, the temporal scales of important information may vary…

声音 · 计算机科学 2023-03-06 Shuaiqi Chen , Xiaofen Xing , Weibin Zhang , Weidong Chen , Xiangmin Xu

Multimodal emotion recognition (MER) aims to infer human affect by jointly modeling audio and visual cues; however, existing approaches often struggle with temporal misalignment, weakly discriminative feature representations, and suboptimal…

多媒体 · 计算机科学 2026-01-21 Joe Dhanith P R , Shravan Venkatraman , Vigya Sharma , Santhosh Malarvannan

Consumers often react expressively to products such as food samples, perfume, jewelry, sunglasses, and clothing accessories. This research discusses a multimodal affect recognition system developed to classify whether a consumer likes or…

人机交互 · 计算机科学 2017-05-09 Amol S Patwardhan , Gerald M Knapp

Personality computing and affective computing have gained recent interest in many research areas. The datasets for the task generally have multiple modalities like video, audio, language and bio-signals. In this paper, we propose a flexible…

计算机视觉与模式识别 · 计算机科学 2023-01-13 Tanay Agrawal , Dhruv Agarwal , Michal Balazia , Neelabh Sinha , Francois Bremond

Multimodal affective computing, learning to recognize and interpret human affects and subjective information from multiple data sources, is still challenging because: (i) it is hard to extract informative features to represent human affects…

计算与语言 · 计算机科学 2018-05-23 Yue Gu , Kangning Yang , Shiyu Fu , Shuhong Chen , Xinyu Li , Ivan Marsic

Expression recognition in in-the-wild video data remains challenging due to substantial variations in facial appearance, background conditions, audio noise, and the inherently dynamic nature of human affect. Relying on a single modality,…

计算机视觉与模式识别 · 计算机科学 2026-03-19 Junhyeong Byeon , Jeongyeol Kim , Sejoon Lim

Deep implicit functions (DIFs) have emerged as a potent and articulate means of representing 3D shapes. However, methods modeling object categories or non-rigid entities have mainly focused on single-object scenarios. In this work, we…

计算机视觉与模式识别 · 计算机科学 2023-12-19 Yuchun Liu , Benjamin Planche , Meng Zheng , Zhongpai Gao , Pierre Sibut-Bourde , Fan Yang , Terrence Chen , Ziyan Wu

In the personalization process of large-scale text-to-image models, overfitting often occurs when learning specific subject from a limited number of images. Existing methods, such as DreamBooth, mitigate this issue through a class-specific…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Seulgi Jeong , Jaeil Kim

The human visual perception system has very strong robustness and contextual awareness in a variety of image processing tasks. This robustness and the perception ability of contextual awareness is closely related to the characteristics of…

计算机视觉与模式识别 · 计算机科学 2019-12-24 Aiqing Fang , Xinbo Zhao , Yanning Zhang