中文
相关论文

相关论文: Spatial-Temporal Graph Mamba for Music-Guided Danc…

200 篇论文

In the field of multimodal medical data analysis, leveraging diverse types of data and understanding their hidden relationships continues to be a research focus. The main challenges lie in effectively modeling the complex interactions…

计算机视觉与模式识别 · 计算机科学 2025-08-26 Xuhao Shan , Ruiquan Ge , Jikui Liu , Linglong Wu , Chi Zhang , Siqi Liu , Wenjian Qin , Wenwen Min , Ahmed Elazab , Changmiao Wang

Magnetic resonance (MR)-to-computed tomography (CT) translation offers significant advantages, including the elimination of radiation exposure associated with CT scans and the mitigation of imaging artifacts caused by patient motion. The…

图像与视频处理 · 电气工程与系统科学 2025-08-08 Chaohui Gong , Zhiying Wu , Zisheng Huang , Gaofeng Meng , Zhen Lei , Hongbin Liu

Due to recent advances in pose-estimation methods, human motion can be extracted from a common video in the form of 3D skeleton sequences. Despite wonderful application opportunities, effective and efficient content-based access to large…

计算机视觉与模式识别 · 计算机科学 2023-10-05 Nicola Messina , Jan Sedmidubsky , Fabrizio Falchi , Tomáš Rebok

Mamba has recently emerged as a promising alternative to Transformers, offering near-linear complexity in processing sequential data. However, while channels in time series (TS) data have no specific order in general, recent studies have…

机器学习 · 计算机科学 2024-11-01 Seunghan Lee , Juri Hong , Kibok Lee , Taeyoung Park

This paper introduces Bio-Inspired Mamba (BIM), a novel online learning framework for selective state space models that integrates biological learning principles with the Mamba architecture. BIM combines Real-Time Recurrent Learning (RTRL)…

神经与进化计算 · 计算机科学 2024-09-18 Jiahao Qin

Recently, the Mamba architecture based on state space models has demonstrated remarkable performance in a series of natural language processing tasks and has been rapidly applied to remote sensing change detection (CD) tasks. However, most…

计算机视觉与模式识别 · 计算机科学 2025-05-20 Haotian Zhang , Keyan Chen , Chenyang Liu , Hao Chen , Zhengxia Zou , Zhenwei Shi

Temporal human action detection aims to identify and localize action segments within untrimmed videos, serving as a pivotal task in video understanding. Despite the progress achieved by prior architectures like CNN and Transformer models,…

计算机视觉与模式识别 · 计算机科学 2026-04-13 Yicheng Qiu , Keiji Yanai

We present Split-then-Merge (StM), a novel framework designed to enhance control in generative video composition and address its data scarcity problem. Unlike conventional methods relying on annotated datasets or handcrafted rules, StM…

计算机视觉与模式识别 · 计算机科学 2025-11-27 Ozgur Kara , Yujia Chen , Ming-Hsuan Yang , James M. Rehg , Wen-Sheng Chu , Du Tran

Change detection (CD) in multitemporal remote sensing imagery presents significant challenges for fine-grained recognition, owing to heterogeneity and spatiotemporal misalignment. However, existing methodologies based on vision transformers…

计算机视觉与模式识别 · 计算机科学 2025-11-26 Lei Ding , Tong Liu , Xuanguang Liu , Xiangyun Liu , Haitao Guo , Jun Lu

Spatio-Temporal video grounding (STVG) focuses on retrieving the spatio-temporal tube of a specific object depicted by a free-form textual expression. Existing approaches mainly treat this complicated task as a parallel frame-grounding…

计算机视觉与模式识别 · 计算机科学 2022-12-02 Yang Jin , Yongzhi Li , Zehuan Yuan , Yadong Mu

Bio-inspired Spiking Neural Networks (SNNs) provide an energy-efficient way to extract 3D spatio-temporal features. However, existing 3D SNNs have struggled with long-range dependencies until the recent emergence of Mamba, which offers…

计算机视觉与模式识别 · 计算机科学 2025-06-27 Peixi Wu , Bosong Chai , Menghua Zheng , Wei Li , Zhangchi Hu , Jie Chen , Zheyu Zhang , Hebei Li , Xiaoyan Sun

Existing methods for instance segmentation in videos typically involve multi-stage pipelines that follow the tracking-by-detection paradigm and model a video clip as a sequence of images. Multiple networks are used to detect objects in…

计算机视觉与模式识别 · 计算机科学 2023-09-04 Ali Athar , Sabarinath Mahadevan , Aljoša Ošep , Laura Leal-Taixé , Bastian Leibe

Modelling various spatio-temporal dependencies is the key to recognising human actions in skeleton sequences. Most existing methods excessively relied on the design of traversal rules or graph topologies to draw the dependencies of the…

计算机视觉与模式识别 · 计算机科学 2021-11-02 Tailin Chen , Shidong Wang , Desen Zhou , Yu Guan

The rapid proliferation of surveillance cameras has increased the demand for automated violence detection. While CNNs and Transformers have shown success in extracting spatio-temporal features, they struggle with long-term dependencies and…

计算机视觉与模式识别 · 计算机科学 2025-09-29 Damith Chamalke Senadeera , Xiaoyun Yang , Shibo Li , Muhammad Awais , Dimitrios Kollias , Gregory Slabaugh

Despite remarkable advances in video generative models, they still struggle to generate physically realistic videos, frequently exhibiting appearance drift, implausible motion, and temporal inconsistencies. In this work, we address this…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Manjin Kim , Suha Kwak , Minsu Cho

3D Hand reconstruction from a single RGB image is challenging due to the articulated motion, self-occlusion, and interaction with objects. Existing SOTA methods employ attention-based transformers to learn the 3D hand pose and shape, yet…

计算机视觉与模式识别 · 计算机科学 2025-11-27 Haoye Dong , Aviral Chharia , Wenbo Gou , Francisco Vicente Carrasco , Fernando De la Torre

Human trajectory forecasting is crucial for safe navigation in crowded environments, requiring models that balance accuracy with computational efficiency. Efficiently modeling social interactions is key to performance in dense crowds. Yet,…

计算机视觉与模式识别 · 计算机科学 2026-05-18 Po-Chien Luan , Wuyang Li , Yang Gao , Alexandre Alahi

While recent Transformer and Mamba architectures have advanced point cloud representation learning, they are typically developed for single-task or single-domain settings. Directly applying them to multi-task domain generalization (DG)…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Jincen Jiang , Qianyu Zhou , Yuhang Li , Kui Su , Meili Wang , Jian Chang , Jian Jun Zhang , Xuequan Lu

We present Scene-Graph Based Multi-Modal Traffic Agent (SGTA), a modular framework for traffic video understanding that combines structured scene graphs with multi-modal reasoning. It constructs a traffic scene graph from roadside videos…

计算机视觉与模式识别 · 计算机科学 2026-04-07 Xingcheng Zhou , Mingyu Liu , Walter Zimmer , Jiajie Zhang , Alois Knoll

Next Point-of-Interest (POI) recommendation is a critical task in location-based services, yet it faces the fundamental challenge of coupled spatiotemporal asymmetry inherent in urban mobility. Specifically, transition intents between…

机器学习 · 计算机科学 2026-03-03 Zhuoxuan Li , Tangwei Ye , Jieyuan Pei , Haina Liang , Zhongyuan Lai , Zihan Liu , Yiming Wu , Qi Zhang , Liang Hu
‹ 上一页 1 8 9 10 下一页 ›