中文
相关论文

相关论文: AU-LLM: Micro-Expression Action Unit Detection via…

200 篇论文

Multimodal Large Language Models (MLLMs) have demonstrated remarkable multimodal emotion recognition capabilities, integrating multimodal cues from visual, acoustic, and linguistic contexts in the video to recognize human emotional states.…

计算机视觉与模式识别 · 计算机科学 2024-07-25 Liyun Zhang

Large Language Models (LLMs) have so far impressed the world, with unprecedented capabilities that emerge in models at large scales. On the vision side, transformer models (i.e., ViT) are following the same trend, achieving the best…

计算机视觉与模式识别 · 计算机科学 2023-10-30 Mustafa Shukor , Corentin Dancette , Matthieu Cord

Large audio-language models (LALMs) can produce expressive speech, yet reliable emotion control remains elusive: conversions often miss the target affect and may degrade linguistic fidelity through refusals, hallucinations, or paraphrase.…

计算与语言 · 计算机科学 2026-03-19 Xiutian Zhao , Ismail Rasim Ulgen , Philipp Koehn , Björn Schuller , Berrak Sisman

Understanding emotions is a fundamental ability for intelligent systems to be able to interact with humans. Vision-language models (VLMs) have made tremendous progress in the last few years for many visual tasks, potentially offering a…

计算机视觉与模式识别 · 计算机科学 2026-04-17 Madhav Agarwal , Sotirios A. Tsaftaris , Laura Sevilla-Lara , Steven McDonagh

Generating realistic human motion sequences from text descriptions is a challenging task that requires capturing the rich expressiveness of both natural language and human motion.Recent advances in diffusion models have enabled significant…

计算机视觉与模式识别 · 计算机科学 2023-12-22 Beibei Jing , Youjia Zhang , Zikai Song , Junqing Yu , Wei Yang

Large Language Models (LLMs) are known to hallucinate and generate non-factual outputs which can undermine user trust. Traditional methods to directly mitigate hallucinations, such as representation editing and contrastive decoding, often…

机器学习 · 计算机科学 2025-03-11 Prasenjit Dey , Srujana Merugu , Sivaramakrishnan Kaveri

Micro-expressions are spontaneous, rapid and subtle facial movements that can neither be forged nor suppressed. They are very important nonverbal communication clues, but are transient and of low intensity thus difficult to recognize.…

计算机视觉与模式识别 · 计算机科学 2023-04-11 Zhijun Zhai , Jianhui Zhao , Chengjiang Long , Wenju Xu , Shuangjiang He , Huijuan Zhao

After the inception of emotion recognition or affective computing, it has increasingly become an active research topic due to its broad applications. Over the past couple of decades, emotion recognition models have gradually migrated from…

计算与语言 · 计算机科学 2023-08-23 Zixing Zhang , Liyizhe Peng , Tao Pang , Jing Han , Huan Zhao , Bjorn W. Schuller

Emotion AI is the ability of computers to understand human emotional states. Existing works have achieved promising progress, but two limitations remain to be solved: 1) Previous studies have been more focused on short sequential video…

计算机视觉与模式识别 · 计算机科学 2025-02-05 Deng Li , Xin Liu , Bohao Xing , Baiqiang Xia , Yuan Zong , Bihan Wen , Heikki Kälviäinen

Micro-expression recognition (MER) is crucial in the affective computing field due to its wide application in medical diagnosis, lie detection, and criminal investigation. Despite its significance, obtaining micro-expression (ME)…

计算机视觉与模式识别 · 计算机科学 2025-08-14 Jiateng Liu , Hengcan Shi , Feng Chen , Zhiwen Shao , Yaonan Wang , Jianfei Cai , Wenming Zheng

The sensor-based recognition of Activities of Daily Living (ADLs) in smart home environments enables several applications in the areas of energy management, safety, well-being, and healthcare. ADLs recognition is typically based on deep…

人工智能 · 计算机科学 2025-03-24 Gabriele Civitarese , Michele Fiori , Priyankar Choudhary , Claudio Bettini

The intensity estimation of facial action units (AUs) is challenging due to subtle changes in the person's facial appearance. Previous approaches mainly rely on probabilistic models or predefined rules for modeling co-occurrence…

计算机视觉与模式识别 · 计算机科学 2020-04-22 Yingruo Fan , Jacqueline C. K. Lam , Victor O. K. Li

Large language models (LLMs) have enabled the creation of multi-modal LLMs that exhibit strong comprehension of visual data such as images and videos. However, these models usually rely on extensive visual tokens from visual encoders,…

计算机视觉与模式识别 · 计算机科学 2025-07-30 Yiwu Zhong , Zhuoming Liu , Yin Li , Liwei Wang

We present an Audio-Visual Language Model (AVLM) for expressive speech generation by integrating full-face visual cues into a pre-trained expressive speech model. We explore multiple visual encoders and multimodal fusion strategies during…

计算与语言 · 计算机科学 2025-08-29 Weiting Tan , Jiachen Lian , Hirofumi Inaguma , Paden Tomasello , Philipp Koehn , Xutai Ma

Misinformation is prevalent in various fields such as education, politics, health, etc., causing significant harm to society. However, current methods for cross-domain misinformation detection rely on effort- and resource-intensive…

计算与语言 · 计算机科学 2025-10-28 Zhiwei Liu , Kailai Yang , Qianqian Xie , Christine de Kock , Sophia Ananiadou , Eduard Hovy

Pressure ulcers (PUs) are a serious and prevalent healthcare concern. Accurate classification of PU severity (Stages I-IV) is essential for proper treatment but remains challenging due to subtle visual distinctions and subjective…

计算机视觉与模式识别 · 计算机科学 2025-10-30 Reza Saadati Fard , Emmanuel Agu , Palawat Busaranuvong , Deepak Kumar , Shefalika Gautam , Bengisu Tulu , Diane Strong , Lorraine Loretz

Large Language Models (LLMs) have demonstrated impressive capabilities in language generation and general task performance. However, their application to spoken language understanding (SLU) remains challenging, particularly for token-level…

计算与语言 · 计算机科学 2025-10-09 Shangjian Yin , Peijie Huang , Jiatian Chen , Haojing Huang , Yuhong Xu

The detection of facial action units (AUs) has been studied as it has the competition due to the wide-ranging applications thereof. In this paper, we propose a novel framework for the AU detection from a single input image by grasping the…

计算机视觉与模式识别 · 计算机科学 2021-04-30 Ziqiang Shi , Liu Liu , Zhongling Liu , Rujie Liu , Xiaoyu Mi , and Kentaro Murase

Large Language Models (LLMs) are playing an increasingly important role in physics research by assisting with symbolic manipulation, numerical computation, and scientific reasoning. However, ensuring the reliability, transparency, and…

人工智能 · 计算机科学 2025-08-19 Yinggan Xu , Hana Kimlee , Yijia Xiao , Di Luo

Facial micro-expressions (MEs) are involuntary movements of the face that occur spontaneously when a person experiences an emotion but attempts to suppress or repress the facial expression, typically found in a high-stakes environment. In…

计算机视觉与模式识别 · 计算机科学 2025-10-16 Xinqi Fan , Jingting Li , John See , Moi Hoon Yap , Wen-Huang Cheng , Xiaobai Li , Xiaopeng Hong , Su-Jing Wang , Adrian K. Davision