中文
相关论文

相关论文: AUD-TGN: Advancing Action Unit Detection with Temp…

200 篇论文

This paper presents our Facial Action Units (AUs) detection submission to the fifth Affective Behavior Analysis in-the-wild Competition (ABAW). Our approach consists of three main modules: (i) a pre-trained facial representation encoder…

计算机视觉与模式识别 · 计算机科学 2023-06-06 Zihan Wang , Siyang Song , Cheng Luo , Yuzhi Zhou , Shiling Wu , Weicheng Xie , Linlin Shen

Audio-Visual Video Parsing (AVVP) task aims to parse the event categories and occurrence times from audio and visual modalities in a given video. Existing methods usually focus on implicitly modeling audio and visual features through weak…

多媒体 · 计算机科学 2025-05-06 Yaru Chen , Peiliang Zhang , Fei Li , Faegheh Sardari , Ruohao Guo , Zhenbo Li , Wenwu Wang

This paper proposes a hybrid fusion-based deep learning approach based on two different modalities, audio and video, to improve human activity recognition and violence detection in public places. To take advantage of audiovisual fusion,…

计算机视觉与模式识别 · 计算机科学 2024-08-06 Pooya Janani , Amirabolfazl Suratgar , Afshin Taghvaeipour

Semantic change detection is an important task in geoscience and earth observation. By producing a semantic change map for each temporal phase, both the land use land cover categories and change information can be interpreted. Recently some…

计算机视觉与模式识别 · 计算机科学 2024-06-18 Chenyao Zhou , Haotian Zhang , Han Guo , Zhengxia Zou , Zhenwei Shi

Active speaker detection and speech enhancement have become two increasingly attractive topics in audio-visual scenario understanding. According to their respective characteristics, the scheme of independently designed architecture has been…

声音 · 计算机科学 2022-07-08 Junwen Xiong , Yu Zhou , Peng Zhang , Lei Xie , Wei Huang , Yufei Zha

Voice Activity Detection (VAD) plays a key role in speech processing, often utilizing hand-crafted or neural features. This study examines the effectiveness of Mel-Frequency Cepstral Coefficients (MFCCs) and pre-trained model (PTM)…

声音 · 计算机科学 2025-06-03 Kumud Tripathi , Chowdam Venkata Kumar , Pankaj Wasnik

Emotion recognition in real-world environments is hindered by partial occlusions, missing modalities, and severe class imbalance. To address these issues, particularly for the Affective Behavior Analysis in-the-wild (ABAW) Expression…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Jun Yu , Naixiang Zheng , Guoyuan Wang , Yunxiang Zhang , Lingsi Zhu , Jiaen Liang , Wei Huang , Shengping Liu

We address the problem of text-guided video temporal grounding, which aims to identify the time interval of a certain event based on a natural language description. Different from most existing methods that only consider RGB images as…

计算机视觉与模式识别 · 计算机科学 2021-11-01 Yi-Wen Chen , Yi-Hsuan Tsai , Ming-Hsuan Yang

Weakly supervised violence detection refers to the technique of training models to identify violent segments in videos using only video-level labels. Among these approaches, multimodal violence detection, which integrates modalities such as…

计算机视觉与模式识别 · 计算机科学 2025-03-17 Wenping Jin , Li Zhu , Jing Sun

This paper describes an approach to the facial action units detections. The involved action units (AU) include AU1 (Inner Brow Raiser), AU2 (Outer Brow Raiser), AU4 (Brow Lowerer), AU6 (Cheek Raise), AU12 (Lip Corner Puller), AU15 (Lip…

计算机视觉与模式识别 · 计算机科学 2020-02-11 Xianpeng Ji , Yu Ding , Lincheng Li , Yu Chen , Changjie Fan

We address the Ambivalence/Hesitancy (A/H) Video Recognition Challenge at the 10th ABAW Competition (CVPR 2026). We propose a divergence-based multimodal fusion that explicitly measures cross-modal conflict between visual, audio, and…

With the rapid proliferation of information across digital platforms, stance detection has emerged as a pivotal challenge in social media analysis. While most of the existing approaches focus solely on textual data, real-world social media…

计算机视觉与模式识别 · 计算机科学 2025-09-11 Lata Pangtey , Omkar Kabde , Shahid Shafi Dar , Nagendra Kumar

News text classification is a crucial task in natural language processing, essential for organizing and filtering the massive volume of digital content. Traditional methods typically rely on statistical features like term frequencies or…

计算与语言 · 计算机科学 2025-11-24 Mohammad Zare

State of the art architectures for untrimmed video Temporal Action Localization (TAL) have only considered RGB and Flow modalities, leaving the information-rich audio modality totally unexploited. Audio fusion has been explored for the…

计算机视觉与模式识别 · 计算机科学 2021-10-19 Anurag Bagchi , Jazib Mahmood , Dolton Fernandes , Ravi Kiran Sarvadevabhatla

Multimodal medical analysis combining image and tabular data has gained increasing attention. However, effective fusion remains challenging due to cross-modal discrepancies in feature dimensions and modality contributions, as well as the…

计算机视觉与模式识别 · 计算机科学 2025-09-17 Congjing Yu , Jing Ye , Yang Liu , Xiaodong Zhang , Zhiyong Zhang

Facial action unit (AU) detection is a fundamental block for objective facial expression analysis. Supervised learning approaches require a large amount of manual labeling which is costly. The limited labeled data are also not diverse in…

计算机视觉与模式识别 · 计算机科学 2024-03-19 Liupei Lu , Yufeng Yin , Yuming Gu , Yizhen Wu , Pratusha Prasad , Yajie Zhao , Mohammad Soleymani

Detecting actions in untrimmed videos should not be limited to a small, closed set of classes. We present a simple, yet effective strategy for open-vocabulary temporal action detection utilizing pretrained image-text co-embeddings. Despite…

计算机视觉与模式识别 · 计算机科学 2023-01-12 Vivek Rathod , Bryan Seybold , Sudheendra Vijayanarasimhan , Austin Myers , Xiuye Gu , Vighnesh Birodkar , David A. Ross

Temporal Video Grounding (TVG) aims to localize the temporal boundary of a specific segment in an untrimmed video based on a given language query. Since datasets in this domain are often gathered from limited video scenes, models tend to…

计算机视觉与模式识别 · 计算机科学 2023-12-22 Haifeng Huang , Yang Zhao , Zehan Wang , Yan Xia , Zhou Zhao

We introduce a challenging decision-making task that we call active acquisition for multimodal temporal data (A2MT). In many real-world scenarios, input features are not readily available at test time and must instead be acquired at…

Human beings have developed fantastic abilities to integrate information from various sensory sources exploring their inherent complementarity. Perceptual capabilities are therefore heightened, enabling, for instance, the well-known…

计算机视觉与模式识别 · 计算机科学 2021-04-14 Gustavo Assunção , Nuno Gonçalves , Paulo Menezes