中文
相关论文

相关论文: TemPose-TF-ASF: Two-Stage Bidirectional Stroke Con…

200 篇论文

This study addresses the critical challenge of error accumulation in spatio-temporal auto-regressive (AR) predictions within scientific machine learning models by exploring temporal integration schemes and adaptive multi-step rollout…

机器学习 · 计算机科学 2025-09-25 Sunwoong Yang , Ricardo Vinuesa , Namwoo Kang

We propose Adaptive Multi-Style Fusion (AMSF), a reference-based training-free framework that enables controllable fusion of multiple reference styles in diffusion models. Most of the existing reference-based methods are limited by (a)…

计算机视觉与模式识别 · 计算机科学 2025-09-24 Xu Liu , Yibo Lu , Xinxian Wang , Xinyu Wu

Trajectory prediction is a challenging task that aims to predict the future trajectory of vehicles or pedestrians over a short time horizon based on their historical positions. The main reason is that the trajectory is a kind of complex…

计算机视觉与模式识别 · 计算机科学 2025-07-18 Pengqian Han , Jiamou Liu , Tianzhe Bao , Yifei Wang

Gait recognition is a biometric technology that has received extensive attention. Most existing gait recognition algorithms are unimodal, and a few multimodal gait recognition algorithms perform multimodal fusion only once. None of these…

计算机视觉与模式识别 · 计算机科学 2023-12-25 Shinan Zou , Jianbo Xiong , Chao Fan , Shiqi Yu , Jin Tang

This paper introduces Gate-Shift-Pose, an enhanced version of Gate-Shift-Fuse networks, designed for athlete fall classification in figure skating by integrating skeleton pose data alongside RGB frames. We evaluate two fusion strategies:…

计算机视觉与模式识别 · 计算机科学 2025-10-06 Edoardo Bianchi , Oswald Lanz

Foundation models have impressive performance and generalization capabilities across a wide range of applications. The increasing size of the models introduces great challenges for the training. Tensor parallelism is a critical technique…

分布式、并行与集群计算 · 计算机科学 2023-01-23 Shenggan Cheng , Ziming Liu , Jiangsu Du , Yang You

The integration of event cameras and spiking neural networks (SNNs) promises energy-efficient visual intelligence, yet scarce event data and the sparsity of DVS outputs hinder effective training. Prior knowledge transfers from RGB to DVS…

计算机视觉与模式识别 · 计算机科学 2025-11-18 Yuqi Xie , Shuhan Ye , Yi Yu , Chong Wang , Qixin Zhang , Jiazhen Xu , Le Shen , Yuanbin Qian , Jiangbo Qian , Guoqi Li

Computer vision based object tracking has been used to annotate and augment sports video. For sports learning and training, video replay is often used in post-match review and training review for tactical analysis and movement analysis. For…

Multimodal time series forecasting is crucial in real-world applications, where decisions depend on both numerical data and contextual signals. The core challenge is to effectively combine temporal numerical patterns with the context…

机器学习 · 计算机科学 2026-02-04 Huu Hiep Nguyen , Minh Hoang Nguyen , Dung Nguyen , Hung Le

Accurate sports prediction is a crucial skill for professional coaches, which can assist in developing effective training strategies and scientific competition tactics. Traditional methods often use complex mathematical statistical…

机器学习 · 计算机科学 2024-09-17 Hui Liu , Jiacheng Gu , Xiyuan Huang , Junjie Shi , Tongtong Feng , Ning He

Identifying significant shots in a rally is important for evaluating players' performance in badminton matches. While there are several studies that have quantified player performance in other sports, analyzing badminton data is remained…

机器学习 · 计算机科学 2021-09-15 Wei-Yao Wang , Teng-Fong Chan , Hui-Kuo Yang , Chih-Chuan Wang , Yao-Chung Fan , Wen-Chih Peng

Text-to-image diffusion inference typically follows synchronized schedules, where the numerical integrator advances the latent state to the same timestep at which the denoiser is conditioned. We propose an asynchronous inference mechanism…

计算机视觉与模式识别 · 计算机科学 2025-12-23 Longhuan Xu , Feng Yin , Cunjian Chen

This paper presents a table tennis stroke detection method from videos. The method relies on a two-stream Convolutional Neural Network processing in parallel the RGB Stream and its computed optical flow. The method has been developed as…

计算机视觉与模式识别 · 计算机科学 2021-12-23 Anam Zahra , Pierre-Etienne Martin

In recent years, the multiple-stage strategy has become a popular trend for visual tracking. This strategy first utilizes a base tracker to coarsely locate the target and then exploits a refinement module to obtain more accurate results.…

计算机视觉与模式识别 · 计算机科学 2021-04-06 Bin Yan , Dong Wang , Huchuan Lu , Xiaoyun Yang

Understanding of the pathophysiology of obstructive lung disease (OLD) is limited by available methods to examine the relationship between multi-omic molecular phenomena and clinical outcomes. Integrative factorization methods for…

统计方法学 · 统计学 2022-12-01 Sarah Samorodnitsky , Chris H. Wendt , Eric F. Lock

Consecutive frames in a video contain redundancy, but they may also contain relevant complementary information for the detection task. The objective of our work is to leverage this complementary information to improve detection. Therefore,…

计算机视觉与模式识别 · 计算机科学 2024-02-19 Noreen Anwar , Guillaume-Alexandre Bilodeau , Wassim Bouachir

The paper addresses the problem of recognition of actions in video with low inter-class variability such as Table Tennis strokes. Two stream, "twin" convolutional neural networks are used with 3D convolutions both on RGB data and optical…

计算机视觉与模式识别 · 计算机科学 2020-12-11 Pierre-Etienne Martin , Jenny Benois-Pineau , Renaud Péteri , Julien Morlier

Diffusion models achieve remarkable success in processing images and text, and have been extended to special domains such as time series forecasting (TSF). Existing diffusion-based approaches for TSF primarily focus on modeling…

计算与语言 · 计算机科学 2025-04-29 Chen Su , Yuanhe Tian , Yan Song

In the realm of multimodal data integration, feature alignment plays a pivotal role. This paper introduces an innovative approach to feature alignment that revolutionizes the fusion of multimodal information. Our method employs a novel…

计算机视觉与模式识别 · 计算机科学 2024-09-27 Jiahao Qin , Yitao Xu , Zong Lu , Xiaojun Zhang

Pre-training of Large Language Models is often prohibitively expensive and inefficient at scale, requiring complex and invasive modifications in order to achieve high data throughput. In this work, we present Token-Superposition Training…

计算与语言 · 计算机科学 2026-05-20 Bowen Peng , Théo Gigant , Jeffrey Quesnelle