中文
相关论文

相关论文: TAG: Thinking with Action Unit Grounding for Facia…

200 篇论文

Foundation models have ushered in a new era for multimodal video understanding by enabling the extraction of rich spatiotemporal and semantic representations. In this work, we introduce a novel graph-based framework that integrates a…

计算机视觉与模式识别 · 计算机科学 2025-10-08 Fatemeh Ziaeetabar , Florentin Wörgötter

Multi-image reasoning and grounding require understanding complex cross-image relationships at both object levels and image levels. Current Large Visual Language Models (LVLMs) face two critical challenges: the lack of cross-image reasoning…

计算机视觉与模式识别 · 计算机科学 2025-09-29 Lihao Zheng , Jiawei Chen , Xintian Shen , Hao Ma , Tao Wei

Recent advances in reasoning language models and reinforcement learning with verifiable rewards have significantly enhanced multi-step reasoning capabilities. This progress motivates the extension of reasoning paradigms to remote sensing…

计算机视觉与模式识别 · 计算机科学 2026-03-23 Shuchang Lyu , Haiquan Wen , Guangliang Cheng , Meng Li , Zheng Zhou , You Zhou , Dingding Yao , Zhenwei Shi

Object-aware reasoning in vision-language tasks poses significant challenges for current models, particularly in handling unseen objects, reducing hallucinations, and capturing fine-grained relationships in complex visual scenes. To address…

计算机视觉与模式识别 · 计算机科学 2024-12-17 Antonio Carlos Rivera , Anthony Moore , Steven Robinson

Recent advancements in Multimodal Large Language Models (MLLMs) have generated significant interest in their ability to autonomously interact with and interpret Graphical User Interfaces (GUIs). A major challenge in these systems is…

计算机视觉与模式识别 · 计算机科学 2024-12-17 Hai-Ming Xu , Qi Chen , Lei Wang , Lingqiao Liu

Understanding and predicting emotion from videos has gathered significant attention in recent studies, driven by advancements in video large language models (VideoLLMs). While advanced methods have made progress in video emotion analysis,…

计算机视觉与模式识别 · 计算机科学 2025-11-05 Zhicheng Zhang , Weicheng Wang , Yongjie Zhu , Wenyu Qin , Pengfei Wan , Di Zhang , Jufeng Yang

Video reasoning requires a fine-grained understanding of the temporal dependencies and event-level relations between objects and events in videos. Current Multimodal Large Language Models (MLLMs) are prone to severe temporal hallucinations…

计算机视觉与模式识别 · 计算机科学 2026-05-14 Zixu Cheng , Da Li , Jian Hu , Yuhang Zang , Ziquan Liu , Shaogang Gong , Wei Li

Facial expression synthesis or editing has recently received increasing attention in the field of affective computing and facial expression modeling. However, most existing facial expression synthesis works are limited in paired training…

计算机视觉与模式识别 · 计算机科学 2020-01-03 Zhilei Liu , Diyi Liu , Yunpeng Wu

The field of Automatic Facial Expression Analysis has grown rapidly in recent years. However, despite progress in new approaches as well as benchmarking efforts, most evaluations still focus on either posed expressions, near-frontal…

计算机视觉与模式识别 · 计算机科学 2017-02-15 Michel F. Valstar , Enrique Sánchez-Lozano , Jeffrey F. Cohn , László A. Jeni , Jeffrey M. Girard , Zheng Zhang , Lijun Yin , Maja Pantic

We propose a margin-based loss for tuning joint vision-language models so that their gradient-based explanations are consistent with region-level annotations provided by humans for relatively smaller grounding datasets. We refer to this…

计算机视觉与模式识别 · 计算机科学 2024-01-09 Ziyan Yang , Kushal Kafle , Franck Dernoncourt , Vicente Ordonez

Key to tasks that require reasoning about natural language in visual contexts is grounding words and phrases to image regions. However, observing this grounding in contemporary models is complex, even if it is generally expected to take…

计算与语言 · 计算机科学 2024-06-03 Noriyuki Kojima , Hadar Averbuch-Elor , Yoav Artzi

Traditional augmented reality (AR) systems predominantly rely on fixed class detectors or fiducial markers, limiting their ability to interpret complex, open-vocabulary natural language queries. We present a modular AR agent system that…

计算机视觉与模式识别 · 计算机科学 2025-12-02 Lixing Guo , Tobias Höllerer

Graphical User Interface (GUI) grounding is commonly framed as a coordinate prediction task -- given a natural language instruction, generate on-screen coordinates for actions such as clicks and keystrokes. However, recent Vision Language…

Joint video-language learning has received increasing attention in recent years. However, existing works mainly focus on single or multiple trimmed video clips (events), which makes human-annotated event boundaries necessary during…

计算机视觉与模式识别 · 计算机科学 2023-05-18 Teng Wang , Jinrui Zhang , Feng Zheng , Wenhao Jiang , Ran Cheng , Ping Luo

Many existing facial expression recognition (FER) systems encounter substantial performance degradation when faced with variations in head pose. Numerous frontalization methods have been proposed to enhance these systems' performance under…

计算机视觉与模式识别 · 计算机科学 2024-04-16 Omar Ikne , Benjamin Allaert , Ioan Marius Bilasco , Hazem Wannous

Spoken Question Answering (Spoken QA) presents a challenging cross-modal problem: effectively aligning acoustic queries with textual knowledge while avoiding the latency and error propagation inherent in cascaded ASR-based systems. In this…

计算与语言 · 计算机科学 2026-03-19 Ke Yang , Bolin Chen , Yuejie Li , Yueying Hua , Jianhao Nie , Yueping He , Bowen Li , Chengjun Mao

A novel Identity-Free conditional Generative Adversarial Network (IF-GAN) was proposed for Facial Expression Recognition (FER) to explicitly reduce high inter-subject variations caused by identity-related facial attributes, e.g., age, race,…

计算机视觉与模式识别 · 计算机科学 2021-05-24 Jie Cai , Zibo Meng , Ahmed Shehab Khan , Zhiyuan Li , James O'Reilly , Shizhong Han , Yan Tong

Foundation Models (FMs) are rapidly transforming Affective Computing (AC), with Vision Language Models (VLMs) now capable of recognising emotions in zero shot settings. This paper probes a critical but underexplored question: what visual…

计算机视觉与模式识别 · 计算机科学 2025-12-01 Iosif Tsangko , Andreas Triantafyllopoulos , Adem Abdelmoula , Adria Mallol-Ragolta , Bjoern W. Schuller

Analyzing human affect is vital for human-computer interaction systems. Most methods are developed in restricted scenarios which are not practical for in-the-wild settings. The Affective Behavior Analysis in-the-wild (ABAW) 2021 Contest…

计算机视觉与模式识别 · 计算机科学 2021-07-16 Yue Jin , Tianqing Zheng , Chao Gao , Guoqiang Xu

Capsule neural network is a new and popular technique in deep learning. However, the traditional capsule neural network does not extract features sufficiently before the dynamic routing between the capsules. In this paper, the one Double…

计算机视觉与模式识别 · 计算机科学 2019-12-06 Shan Cao , Yuqian Yao , Gaoyun An