中文
相关论文

相关论文: Multivariate Gaussian Representation Learning for …

200 篇论文

We propose a new action and gesture recognition method based on spatio-temporal covariance descriptors and a weighted Riemannian locality preserving projection approach that takes into account the curved space formed by the descriptors. The…

计算机视觉与模式识别 · 计算机科学 2013-03-26 Andres Sanin , Conrad Sanderson , Mehrtash T. Harandi , Brian C. Lovell

Computer-Assisted Intervention (CAI) has the potential to revolutionize modern surgery, with surgical scene understanding serving as a critical component in supporting decision-making, improving procedural efficacy, and ensuring…

Deep learning models have achieved excellent recognition results on large-scale video benchmarks. However, they perform poorly when applied to videos with rare scenes or objects, primarily due to the bias of existing video datasets. We…

计算机视觉与模式识别 · 计算机科学 2022-09-21 Haodong Duan , Yue Zhao , Kai Chen , Yuanjun Xiong , Dahua Lin

Volumetric videos offer immersive 4D experiences, but remain difficult to reconstruct, store, and stream at scale. Existing Gaussian Splatting based methods achieve high-quality reconstruction but break down on long sequences, temporal…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Aashish Rai , Angela Xing , Anushka Agarwal , Xiaoyan Cong , Zekun Li , Tao Lu , Aayush Prakash , Srinath Sridhar

Real-time rendering of human head avatars is a cornerstone of many computer graphics applications, such as augmented reality, video games, and films, to name a few. Recent approaches address this challenge with computationally efficient…

计算机视觉与模式识别 · 计算机科学 2024-09-19 Kartik Teotia , Hyeongwoo Kim , Pablo Garrido , Marc Habermann , Mohamed Elgharib , Christian Theobalt

High-fidelity reconstruction of deformable tissues from endoscopic videos remains challenging due to the limitations of existing methods in capturing subtle color variations and modeling global deformations. While 3D Gaussian Splatting…

计算机视觉与模式识别 · 计算机科学 2025-08-27 Qun Ji , Peng Li , Mingqiang Wei

While automatic monitoring and coaching of exercises are showing encouraging results in non-medical applications, they still have limitations such as errors and limited use contexts. To allow the development and assessment of physical…

机器学习 · 计算机科学 2025-01-14 Sao Mai Nguyen , Maxime Devanne , Olivier Remy-Neris , Mathieu Lempereur , André Thepaut

The growing rate of chronic wound occurrence, especially in patients with diabetes, has become a concerning trend in recent years. Chronic wounds are difficult and costly to treat, and have become a serious burden on health care systems…

图像与视频处理 · 电气工程与系统科学 2025-05-30 Bill Cassidy , Christian McBride , Connah Kendrick , Neil D. Reeves , Joseph M. Pappachan , Shaghayegh Raad , Moi Hoon Yap

We propose GC-VASE, a graph convolutional-based variational autoencoder that leverages contrastive learning for subject representation learning from EEG data. Our method successfully learns robust subject-specific latent representations…

信号处理 · 电气工程与系统科学 2025-01-29 Aditya Mishra , Ahnaf Mozib Samin , Ali Etemad , Javad Hashemi

Robust perception and dynamics modeling are fundamental to real-world robotic policy learning. Recent methods employ video diffusion models (VDMs) to enhance robotic policies, improving their understanding and modeling of the physical…

机器人学 · 计算机科学 2026-03-25 Yueru Jia , Jiaming Liu , Shengbang Liu , Rui Zhou , Wanhe Yu , Yuyang Yan , Xiaowei Chi , Yandong Guo , Boxin Shi , Shanghang Zhang

Accurate motion estimation at high acceleration factors enables rapid motion-compensated reconstruction in Magnetic Resonance Imaging (MRI) without compromising the diagnostic image quality. In this work, we introduce an attention-aware…

图像与视频处理 · 电气工程与系统科学 2024-04-30 Aya Ghoul , Jiazhen Pan , Andreas Lingg , Jens Kübler , Patrick Krumm , Kerstin Hammernik , Daniel Rueckert , Sergios Gatidis , Thomas Küstner

3D Gaussian Splatting (3D-GS) has emerged as an efficient 3D representation and a promising foundation for semantic tasks like segmentation. However, existing 3D-GS-based segmentation methods typically rely on high-dimensional category…

计算机视觉与模式识别 · 计算机科学 2025-12-02 An Yang , Chenyu Liu , Jun Du , Jianqing Gao , Jia Pan , Jinshui Hu , Baocai Yin , Bing Yin , Cong Liu

The recent success in human action recognition with deep learning methods mostly adopt the supervised learning paradigm, which requires significant amount of manually labeled data to achieve good performance. However, label collection is an…

计算机视觉与模式识别 · 计算机科学 2018-09-07 Junnan Li , Yongkang Wong , Qi Zhao , Mohan S. Kankanhalli

The canonical approach to video action recognition dictates a neural model to do a classic and standard 1-of-N majority vote task. They are trained to predict a fixed set of predefined categories, limiting their transferable ability on new…

计算机视觉与模式识别 · 计算机科学 2021-09-20 Mengmeng Wang , Jiazheng Xing , Yong Liu

Multi-label classification (MLC) is a prediction task where each sample can have more than one label. We propose a novel contrastive learning boosted multi-label prediction model based on a Gaussian mixture variational autoencoder…

机器学习 · 计算机科学 2022-06-13 Junwen Bai , Shufeng Kong , Carla P. Gomes

In robot-assisted minimally invasive surgery, high-fidelity dynamic endoscopic scene reconstruction and simulation are crucial to enhancing downstream tasks and advancing surgical outcomes. However, existing methods primarily focus on…

计算机视觉与模式识别 · 计算机科学 2026-05-18 Changjing Liu , Yiming Huang , Long Bai , Beilei Cui , Hongliang Ren

We introduce Gaussian Articulated Template Model GART, an explicit, efficient, and expressive representation for non-rigid articulated subject capturing and rendering from monocular videos. GART utilizes a mixture of moving 3D Gaussians to…

计算机视觉与模式识别 · 计算机科学 2023-11-28 Jiahui Lei , Yufu Wang , Georgios Pavlakos , Lingjie Liu , Kostas Daniilidis

Time series analysis has emerged as an important tool for improving patient diagnosis and management in healthcare applications. However, these applications commonly face two critical challenges: time misalignment and data sparsity.…

机器学习 · 统计学 2025-09-25 Dohyun Ku , Catherine D. Chong , Visar Berisha , Todd J. Schwedt , Jing Li

In this paper, we propose XGC-AVis, a multi-agent framework that enhances the audio-video temporal alignment capabilities of multimodal large models (MLLMs) and improves the efficiency of retrieving key video segments through 4 stages:…

多媒体 · 计算机科学 2025-09-30 Yuqin Cao , Xiongkuo Min , Yixuan Gao , Wei Sun , Zicheng Zhang , Jinliang Han , Guangtao Zhai

Vision-language models, while effective in general domains and showing strong performance in diverse multi-modal applications like visual question-answering (VQA), struggle to maintain the same level of effectiveness in more specialized…

计算与语言 · 计算机科学 2024-04-26 Cuong Nhat Ha , Shima Asaadi , Sanjeev Kumar Karn , Oladimeji Farri , Tobias Heimann , Thomas Runkler