中文
相关论文

相关论文: Temporal-Spatial Decouple before Act: Disentangled…

200 篇论文

Reasoning about spatial audio with large language models requires a spatial audio encoder as an acoustic front-end to obtain audio embeddings for further processing. Such an encoder needs to capture all information required to detect the…

音频与语音处理 · 电气工程与系统科学 2025-11-04 Kevin Wilkinghoff , Zheng-Hua Tan

The task of language-guided video temporal grounding is to localize the particular video clip corresponding to a query sentence in an untrimmed video. Though progress has been made continuously in this field, some issues still need to be…

计算机视觉与模式识别 · 计算机科学 2020-09-24 Binjie Zhang , Yu Li , Chun Yuan , Dejing Xu , Pin Jiang , Ying Shan

Magnetic resonance (MR) protocols rely on several sequences to assess pathology and organ status properly. Despite advances in image analysis, we tend to treat each sequence, here termed modality, in isolation. Taking advantage of the…

计算机视觉与模式识别 · 计算机科学 2020-11-11 Agisilaos Chartsias , Giorgos Papanastasiou , Chengjia Wang , Scott Semple , David E. Newby , Rohan Dharmakumar , Sotirios A. Tsaftaris

Audio-visual multi-modal modeling has been demonstrated to be effective in many speech related tasks, such as speech recognition and speech enhancement. This paper introduces a new time-domain audio-visual architecture for target speaker…

音频与语音处理 · 电气工程与系统科学 2019-09-24 Jian Wu , Yong Xu , Shi-Xiong Zhang , Lian-Wu Chen , Meng Yu , Lei Xie , Dong Yu

Existing backdoor attacks on multivariate time series (MTS) forecasting enforce strict temporal and dimensional coupling between triggers and target patterns, requiring synchronous activation at fixed positions across variables. However,…

密码学与安全 · 计算机科学 2026-01-09 Zhixin Liu , Xuanlin Liu , Sihan Xu , Yaqiong Qiao , Ying Zhang , Xiangrui Cai

In the realm of multimodal data integration, feature alignment plays a pivotal role. This paper introduces an innovative approach to feature alignment that revolutionizes the fusion of multimodal information. Our method employs a novel…

计算机视觉与模式识别 · 计算机科学 2024-09-27 Jiahao Qin , Yitao Xu , Zong Lu , Xiaojun Zhang

In vision and linguistics; the main input modalities are facial expressions, speech patterns, and the words uttered. The issue with analysis of any one mode of expression (Visual, Verbal or Vocal) is that lot of contextual information can…

计算机视觉与模式识别 · 计算机科学 2022-11-29 Kunjal Panchal

Spatio-temporal traffic forecasting is challenging due to complex temporal patterns, dynamic spatial structures, and diverse input formats. Although Transformer-based models offer strong global modeling, they often struggle with rigid…

人工智能 · 计算机科学 2025-08-20 Jiayu Fang , Zhiqi Shao , S T Boris Choy , Junbin Gao

Temporal action detection (TAD) is an important yet challenging task in video analysis. Most existing works draw inspiration from image object detection and tend to reformulate it as a proposal generation - classification problem. However,…

计算机视觉与模式识别 · 计算机科学 2022-03-04 Chen Zhao , Merey Ramazanova , Mengmeng Xu , Bernard Ghanem

Traditional sentiment analysis has long been a unimodal task, relying solely on text. This approach overlooks non-verbal cues such as vocal tone and prosody that are essential for capturing true emotional intent. We introduce Dynamic…

计算与语言 · 计算机科学 2025-09-30 Sadia Abdulhalim , Muaz Albaghdadi , Moshiur Farazi

Due to its ability to accurately predict emotional state using multimodal features, audiovisual emotion recognition has recently gained more interest from researchers. This paper proposes two methods to predict emotional attributes from…

音频与语音处理 · 电气工程与系统科学 2022-07-22 Bagus Tris Atmaja , Masato Akagi

This paper proposes a method for long-term action anticipation (LTA), the task of predicting action labels and their duration in a video given the observation of an initial untrimmed video interval. We build on an encoder-decoder…

计算机视觉与模式识别 · 计算机科学 2024-12-30 Alberto Maté , Mariella Dimiccoli

The widespread availability of complex time series data in various domains such as environmental science, epidemiology, and economics demands robust causal discovery methods that can identify intricate contemporaneous and lagged…

机器学习 · 计算机科学 2026-05-12 Omar Faruque , Sahara Ali , Xue Zheng , Jianwu Wang

Dynamic mode decomposition (DMD) is a leading tool for equation-free analysis of high-dimensional dynamical systems from observations. In this work, we focus on a combination of delay-coordinates embedding and DMD, i.e., delay-coordinates…

动力系统 · 数学 2022-12-21 Emil Bronstein , Aviad Wiegner , Doron Shilo , Ronen Talmon

Language-queried video actor segmentation aims to predict the pixel-level mask of the actor which performs the actions described by a natural language query in the target frames. Existing methods adopt 3D CNNs over the video clip as a…

计算机视觉与模式识别 · 计算机科学 2021-05-17 Tianrui Hui , Shaofei Huang , Si Liu , Zihan Ding , Guanbin Li , Wenguan Wang , Jizhong Han , Fei Wang

Inspired by the dual-stream theory of the human visual system (HVS) - where the ventral stream is responsible for object recognition and detail analysis, while the dorsal stream focuses on spatial relationships and motion perception - an…

计算机视觉与模式识别 · 计算机科学 2025-04-22 Li Yu , Situo Wang , Wei Zhou , Moncef Gabbouj

Spatio-temporal knowledge graphs (STKGs) enhance traditional KGs by integrating temporal and spatial annotations, enabling precise reasoning over questions with spatio-temporal dependencies. Despite their potential, research on…

计算与语言 · 计算机科学 2025-12-17 Xinbang Dai , Huiying Li , Nan Hu , Yongrui Chen , Rihui Jin , Huikang Hu , Guilin Qi

Temporal reasoning is an important aspect of video analysis. 3D CNN shows good performance by exploring spatial-temporal features jointly in an unconstrained way, but it also increases the computational cost a lot. Previous works try to…

计算机视觉与模式识别 · 计算机科学 2019-10-01 Chenxu Luo , Alan Yuille

Audio-visual emotion recognition (AVER) methods typically fuse utterance-level features, and even frame-level attention models seldom address the frame-rate mismatch across modalities. In this paper, we propose a Transformer-based framework…

多媒体 · 计算机科学 2026-03-13 Inyong Koo , yeeun Seong , Minseok Son , Jaehyuk Jang , Changick Kim

Multimodal sentiment analysis has been studied under the assumption that all modalities are available. However, such a strong assumption does not always hold in practice, and most of multimodal fusion models may fail when partial modalities…

机器学习 · 计算机科学 2022-05-02 Jiandian Zeng , Tianyi Liu , Jiantao Zhou