English
Related papers

Related papers: Hear Me Out: Fusional Approaches for Audio Augment…

200 papers

Recent advancements in audio tokenization have significantly enhanced the integration of audio capabilities into large language models (LLMs). However, audio understanding and generation are often treated as distinct tasks, hindering the…

Temporal modelling is the key for efficient video action recognition. While understanding temporal information can improve recognition accuracy for dynamic actions, removing temporal redundancy and reusing past features can significantly…

Computer Vision and Pattern Recognition · Computer Science 2021-02-12 Yue Meng , Rameswar Panda , Chung-Ching Lin , Prasanna Sattigeri , Leonid Karlinsky , Kate Saenko , Aude Oliva , Rogerio Feris

Temporal action localization (TAL) is a task of identifying a set of actions in a video, which involves localizing the start and end frames and classifying each action instance. Existing methods have addressed this task by using predefined…

Computer Vision and Pattern Recognition · Computer Science 2022-07-22 Tae-Kyung Kang , Gun-Hee Lee , Seong-Whan Lee

In this paper, we consider the problem of temporal action localization under low-shot (zero-shot & few-shot) scenario, with the goal of detecting and classifying the action instances from arbitrary categories within some untrimmed videos,…

Computer Vision and Pattern Recognition · Computer Science 2023-03-22 Chen Ju , Zeqian Li , Peisen Zhao , Ya Zhang , Xiaopeng Zhang , Qi Tian , Yanfeng Wang , Weidi Xie

Temporal Action Detection (TAD) requires precise localization of action boundaries within long, untrimmed video sequences. While current high-performing methods achieve strong accuracy, they are often characterized by excessive parameter…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Zepeng Sun , Naichuan Zheng , Hailun Xia , Junjie Wu , Liwei Bao , Xiaotai Zhang

Audio and video are two most common modalities in the mainstream media platforms, e.g., YouTube. To learn from multimodal videos effectively, in this work, we propose a novel audio-video recognition approach termed audio video Transformer,…

Computer Vision and Pattern Recognition · Computer Science 2024-01-10 Wentao Zhu

Temporal action localization (TAL), which involves recognizing and locating action instances, is a challenging task in video understanding. Most existing approaches directly predict action classes and regress offsets to boundaries, while…

Computer Vision and Pattern Recognition · Computer Science 2023-09-14 Jiayi Shao , Xiaohan Wang , Ruijie Quan , Junjun Zheng , Jiang Yang , Yi Yang

Multimodal tabular-image fusion is an emerging task that has received increasing attention in various domains. However, existing methods may be hindered by gradient conflicts between modalities, misleading the optimization of the unimodal…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Longfei Huang , Yang Yang

Fusion technique is a key research topic in multimodal sentiment analysis. The recent attention-based fusion demonstrates advances over simple operation-based fusion. However, these fusion works adopt single-scale, i.e., token-level or…

Computation and Language · Computer Science 2021-12-03 Huaishao Luo , Lei Ji , Yanyong Huang , Bin Wang , Shenggong Ji , Tianrui Li

Effective fusion of data from multiple modalities, such as video, speech, and text, is challenging due to the heterogeneous nature of multimodal data. In this paper, we propose adaptive fusion techniques that aim to model context from…

Computation and Language · Computer Science 2021-01-27 Gaurav Sahu , Olga Vechtomova

Literature on self-assessment in machine learning mainly focuses on the production of well-calibrated algorithms through consensus frameworks i.e. calibration is seen as a problem. Yet, we observe that learning to be properly confident…

Machine Learning · Computer Science 2020-11-16 Guillaume Vaudaux-Ruth , Adrien Chan-Hon-Tong , Catherine Achard

Visual SLAM is particularly challenging in environments affected by noise, varying lighting conditions, and darkness. Learning-based optical flow algorithms can leverage multiple modalities to address these challenges, but traditional…

Computer Vision and Pattern Recognition · Computer Science 2025-06-25 Youjie Zhou , Guofeng Mei , Yiming Wang , Yi Wan , Fabio Poiesi

Automatic speaker naming is the problem of localizing as well as identifying each speaking character in a TV/movie/live show video. This is a challenging problem mainly attributes to its multimodal nature, namely face cue alone is…

Computer Vision and Pattern Recognition · Computer Science 2015-07-20 Yongtao Hu , Jimmy Ren , Jingwen Dai , Chang Yuan , Li Xu , Wenping Wang

IMU-based Human Activity Recognition (HAR) has enabled a wide range of ubiquitous computing applications, yet its dominant clip classification paradigm cannot capture the rich temporal structure of real-world behaviors. This motivates a…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Pei Li , Jiaxi Yin , Lei Ouyang , Shihan Pan , Ge Wang , Han Ding , Fei Wang

The main idea of multimodal recommendation is the rational utilization of the item's multimodal information to improve the recommendation performance. Previous works directly integrate item multimodal features with item ID embeddings,…

Information Retrieval · Computer Science 2023-04-25 Yan Zhou , Jie Guo , Hao Sun , Bin Song , Fei Richard Yu

Temporal Action Localization (TAL) is a challenging task in video understanding that aims to identify and localize actions within a video sequence. Recent studies have emphasized the importance of applying long-term temporal context…

Computer Vision and Pattern Recognition · Computer Science 2023-03-17 Tuan N. Tang , Kwonyoung Kim , Kwanghoon Sohn

This paper studies the joint learning of action recognition and temporal localization in long, untrimmed videos. We employ a multi-task learning framework that performs the three highly related steps of action proposal, action recognition,…

Computer Vision and Pattern Recognition · Computer Science 2017-04-05 Yi Zhu , Shawn Newsam

Automated incident management is critical for microservice reliability. While recent unified frameworks leverage multimodal data for joint optimization, they unrealistically assume perfect data completeness. In practice, network…

Machine Learning · Computer Science 2026-03-30 Wenzhuo Qian , Hailiang Zhao , Ziqi Wang , Zhipeng Gao , Jiayi Chen , Zhiwei Ling , Shuiguang Deng

Sentiment analysis, mostly based on text, has been rapidly developing in the last decade and has attracted widespread attention in both academia and industry. However, the information in the real world usually comes from multiple…

Computation and Language · Computer Science 2019-12-12 Feiyang Chen , Ziqian Luo , Yanyan Xu , Dengfeng Ke

We present our solution to the BinEgo-360 Challenge at ICCV 2025, which focuses on temporal action localization (TAL) in multi-perspective and multi-modal video settings. The challenge provides a dataset containing panoramic, third-person,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-15 Anh-Kiet Duong , Petra Gomez-Krämer