中文
相关论文

相关论文: Multivariate Gaussian Representation Learning for …

200 篇论文

Cross view action recognition (CVAR) seeks to recognize a human action when observed from a previously unseen viewpoint. This is a challenging problem since the appearance of an action changes significantly with the viewpoint. Applications…

计算机视觉与模式识别 · 计算机科学 2023-05-04 Yuexi Zhang , Dan Luo , Balaji Sundareshan , Octavia Camps , Mario Sznaier

Deep neural networks have achieved remarkable success for video-based action recognition. However, most of existing approaches cannot be deployed in practice due to the high computational cost. To address this challenge, we propose a new…

计算机视觉与模式识别 · 计算机科学 2020-06-18 Kun Liu , Wu Liu , Huadong Ma , Mingkui Tan , Chuang Gan

Recent advances in multimodal ECG representation learning center on aligning ECG signals with paired free-text reports. However, suboptimal alignment persists due to the complexity of medical language and the reliance on a full 12-lead…

机器学习 · 计算机科学 2025-02-26 Che Liu , Cheng Ouyang , Zhongwei Wan , Haozhe Wang , Wenjia Bai , Rossella Arcucci

Extracting compact, physically interpretable representations from high-dimensional scientific data is a persistent challenge due to the complex, nonlinear structures inherent in physical systems. We propose a Gaussian Mixture Variational…

机器学习 · 计算机科学 2025-12-01 Tiffany Fan , Murray Cutforth , Marta D'Elia , Alexandre Cortiella , Alireza Doostan , Eric Darve

Recent work on action recognition leverages 3D features and textual information to achieve state-of-the-art performance. However, most of the current few-shot action recognition methods still rely on 2D frame-level representations, often…

计算机视觉与模式识别 · 计算机科学 2023-11-13 Yutao Tang , Benjamin Bejar , Rene Vidal

Video-based Clinical Gait Analysis often suffers from poor generalization as models overfit environmental biases instead of capturing pathological motion. To address this, we propose BioGait-VLM, a tri-modal Vision-Language-Biomechanics…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Erdong Chen , Yuyang Ji , Jacob K. Greenberg , Benjamin Steel , Faraz Arkam , Abigail Lewis , Pranay Singh , Feng Liu

Magnetic Resonance Imaging (MRI) is a crucial non-invasive imaging modality. In routine clinical practice, multi-stack thick-slice acquisitions are widely used to reduce scan time and motion sensitivity, particularly in challenging…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Kangyuan Zheng , Xuan Cai , Jiangqi Wang , Guixing Fu , Zhuoshuo Li , Yazhou Chen , Xinting Ge , Liangqiong Qu , Mengting Liu

Vision-language-action (VLA) policies have advanced language-conditioned robotic manipulation by transferring semantic priors from pretrained vision-language models to action generation. However, standard action-imitation learning often…

机器人学 · 计算机科学 2026-05-29 Zijian Zhang , Yuqing Jiang , Qian Cheng , Xiaofan Li , Si Liu , Ding Zhao , Ping Luo , Weitao Zhou , Haibao Yu

Generalist Medical AI (GMAI) systems have demonstrated expert-level performance in biomedical perception tasks, yet their clinical utility remains limited by inadequate multi-modal explainability and suboptimal prognostic capabilities.…

计算机视觉与模式识别 · 计算机科学 2025-05-13 Honglong Yang , Shanshan Song , Yi Qin , Lehan Wang , Haonan Wang , Xinpeng Ding , Qixiang Zhang , Bodong Du , Xiaomeng Li

Multimodal learning leverages complementary information derived from different modalities, thereby enhancing performance in medical image segmentation. However, prevailing multimodal learning methods heavily rely on extensive well-annotated…

计算机视觉与模式识别 · 计算机科学 2024-09-05 Xiaogen Zhou , Yiyou Sun , Min Deng , Winnie Chiu Wing Chu , Qi Dou

3D Gaussian Splatting (3DGS) has made significant strides in scene representation and neural rendering, with intense efforts focused on adapting it for dynamic scenes. Despite delivering remarkable rendering quality and speed, existing…

计算机视觉与模式识别 · 计算机科学 2025-03-25 Sangwoon Kwak , Joonsoo Kim , Jun Young Jeong , Won-Sik Cheong , Jihyong Oh , Munchurl Kim

3D skeleton-based human action recognition has emerged as a powerful alternative to traditional RGB and depth-based approaches, offering robustness to environmental variations, computational efficiency, and enhanced privacy. Despite…

计算机视觉与模式识别 · 计算机科学 2025-09-16 Yang Liu , Jiyao Yang , Madhawa Perera , Pan Ji , Dongwoo Kim , Min Xu , Tianyang Wang , Saeed Anwar , Tom Gedeon , Lei Wang , Zhenyue Qin

3D Gaussian splats have emerged as a revolutionary, effective, learned representation for static 3D scenes. In this work, we explore using 2D Gaussian splats as a new primitive for representing videos. We propose GSVC, an approach to…

计算机视觉与模式识别 · 计算机科学 2025-01-23 Longan Wang , Yuang Shi , Wei Tsang Ooi

Automated surgical gesture recognition is of great importance in robot-assisted minimally invasive surgery. However, existing methods assume that training and testing data are from the same domain, which suffers from severe performance…

计算机视觉与模式识别 · 计算机科学 2021-07-20 Xueying Shi , Yueming Jin , Qi Dou , Jing Qin , Pheng-Ann Heng

The visualization of volumetric medical data is crucial for enhancing diagnostic accuracy and improving surgical planning and education. Cinematic rendering techniques significantly enrich this process by providing high-quality…

计算机视觉与模式识别 · 计算机科学 2025-07-10 Chengkun Li , Yuqi Tong , Kai Chen , Zhenya Yang , Ruiyang Li , Shi Qiu , Jason Ying-Kuen Chan , Pheng-Ann Heng , Qi Dou

Recently, the popularity of depth-sensors such as Kinect has made depth videos easily available while its advantages have not been fully exploited. This paper investigates, for gesture recognition, to explore the spatial and temporal…

计算机视觉与模式识别 · 计算机科学 2016-11-29 Jiali Duan , Shuai Zhou , Jun Wan , Xiaoyuan Guo , Stan Z. Li

Volumetric video represents a transformative advancement in visual media, enabling users to freely navigate immersive virtual experiences and narrowing the gap between digital and real worlds. However, the need for extensive manual…

图形学 · 计算机科学 2024-09-16 Yuheng Jiang , Zhehao Shen , Yu Hong , Chengcheng Guo , Yize Wu , Yingliang Zhang , Jingyi Yu , Lan Xu

We propose multivariate nonstationary Gaussian processes for jointly modeling multiple clinical variables, where the key parameters, length-scales, standard deviations and the correlations between the observed output, are all time…

统计方法学 · 统计学 2019-10-15 Rui Meng , Braden Soper , Herbert Lee , Vincent X. Liu , John D. Greene , Priyadip Ray

A long-standing objective in humanoid robotics is the realization of versatile agents capable of following diverse multimodal instructions with human-level flexibility. Despite advances in humanoid control, bridging high-level multimodal…

计算机视觉与模式识别 · 计算机科学 2026-01-01 Nan Jiang , Zimo He , Wanhe Yu , Lexi Pang , Yunhao Li , Hongjie Li , Jieming Cui , Yuhan Li , Yizhou Wang , Yixin Zhu , Siyuan Huang

2D Gaussian Splatting (2DGS) has recently become a promising paradigm for high-quality video representation. However, existing methods employ content-agnostic or spatio-temporal feature overlapping embeddings to predict canonical Gaussian…

计算机视觉与模式识别 · 计算机科学 2026-04-15 Jierun Lin , Jiacong Chen , Qingyu Mao , Shuai Liu , Xiandong Meng , Fanyang Meng , Yongsheng Liang