中文
相关论文

相关论文: Multivariate Gaussian Representation Learning for …

200 篇论文

This paper introduces the Gaussian multi-Graphical Model, a model to construct sparse graph representations of matrix- and tensor-variate data. We generalize prior work in this area by simultaneously learning this representation across…

机器学习 · 统计学 2024-02-28 Bailey Andrew , David Westhead , Luisa Cutillo

Most brain disorders are very heterogeneous in terms of their underlying biology and developing analysis methods to model such heterogeneity is a major challenge. A promising approach is to use probabilistic regression methods to estimate…

机器学习 · 统计学 2018-12-03 Seyed Mostafa Kia , Christian F. Beckmann , Andre F. Marquand

Understanding 4D point cloud videos is essential for enabling intelligent agents to perceive dynamic environments. However, temporal scale bias across varying frame rates and distributional uncertainty in irregular point clouds make it…

计算机视觉与模式识别 · 计算机科学 2026-03-18 Jiayi Tian , Jiaze Wang

Accurate human motion prediction with well-calibrated uncertainty is critical for safe human-robot collaboration (HRC), where robots must anticipate and react to human movements in real time. We propose a structured multitask variational…

机器人学 · 计算机科学 2026-03-10 Jinger Chong , Xiaotong Zhang , Kamal Youcef-Toumi

Modeling animatable human avatars from videos is a long-standing and challenging problem. While conventional methods require per-instance optimization, recent feed-forward methods have been proposed to generate 3D Gaussians with a learnable…

计算机视觉与模式识别 · 计算机科学 2025-07-28 Yifan Liu , Shengjun Zhang , Chensheng Dai , Yang Chen , Hao Liu , Chen Li , Yueqi Duan

Action anticipation, the task of predicting future actions from partially observed videos, is crucial for advancing intelligent systems. Unlike action recognition, which operates on fully observed videos, action anticipation must handle…

计算机视觉与模式识别 · 计算机科学 2025-06-04 Seulgi Kim , Ghazal Kaviani , Mohit Prabhushankar , Ghassan AlRegib

We propose Video Gaussian Masked Autoencoders (Video-GMAE), a self-supervised approach for representation learning that encodes a sequence of images into a set of Gaussian splats moving over time. Representing a video as a set of Gaussians…

计算机视觉与模式识别 · 计算机科学 2026-01-01 Tanish Baranwal , Himanshu Gaurav Singh , Jathushan Rajasegaran , Jitendra Malik

We present Gaussian See, Gaussian Do, a novel approach for semantic 3D motion transfer from multiview video. Our method enables rig-free, cross-category motion transfer between objects with semantically meaningful correspondence. Building…

计算机视觉与模式识别 · 计算机科学 2025-11-20 Yarin Bekor , Gal Michael Harari , Or Perel , Or Litany

Accurate face parsing under extreme viewing angles remains a significant challenge due to limited labeled data in such poses. Manual annotation is costly and often impractical at scale. We propose a novel label refinement pipeline that…

计算机视觉与模式识别 · 计算机科学 2026-03-23 Ankit Gahlawat , Anirban Mukherjee , Dinesh Babu Jayagopi

Masked video modeling~(MVM) has emerged as a highly effective pre-training strategy for visual foundation models, whereby the model reconstructs masked spatiotemporal tokens using information from visible tokens. However, a key challenge in…

计算机视觉与模式识别 · 计算机科学 2025-08-15 Ayush K. Rai , Kyle Min , Tarun Krishna , Feiyan Hu , Alan F. Smeaton , Noel E. O'Connor

Neural image representations have emerged as a promising approach for encoding and rendering visual data. Combined with learning-based workflows, they demonstrate impressive trade-offs between visual fidelity and memory footprint. Existing…

计算机视觉与模式识别 · 计算机科学 2025-05-08 Yunxiang Zhang , Bingxuan Li , Alexandr Kuznetsov , Akshay Jindal , Stavros Diolatzis , Kenneth Chen , Anton Sochenov , Anton Kaplanyan , Qi Sun

Recognizing surgical gestures in real-time is a stepping stone towards automated activity recognition, skill assessment, intra-operative assistance, and eventually surgical automation. The current robotic surgical systems provide us with…

计算机视觉与模式识别 · 计算机科学 2025-03-21 Jumanh Atoum , Garrison L. H. Johnston , Nabil Simaan , Jie Ying Wu

Quantitative analysis of cardiac motion is crucial for assessing cardiac function. This analysis typically uses imaging modalities such as MRI and Echocardiograms that capture detailed image sequences throughout the heartbeat cycle.…

计算机视觉与模式识别 · 计算机科学 2024-10-29 Jiewen Yang , Yiqun Lin , Bin Pu , Xiaomeng Li

Realistic videos of human actions exhibit rich spatiotemporal structures at multiple levels of granularity: an action can always be decomposed into multiple finer-grained elements in both space and time. To capture this intuition, we…

计算机视觉与模式识别 · 计算机科学 2015-09-01 Tian Lan , Yuke Zhu , Amir Roshan Zamir , Silvio Savarese

In this paper we propose a novel approach to multi-action recognition that performs joint segmentation and classification. This approach models each action using a Gaussian mixture using robust low-dimensional action features. Segmentation…

计算机视觉与模式识别 · 计算机科学 2015-02-09 Johanna Carvajal , Conrad Sanderson , Chris McCool , Brian C. Lovell

Reconstructing articulated objects is essential for building digital twins of interactive environments. However, prior methods typically decouple geometry and motion by first reconstructing object shape in distinct states and then…

计算机视觉与模式识别 · 计算机科学 2025-11-13 Licheng Shen , Saining Zhang , Honghan Li , Peilin Yang , Zihao Huang , Zongzheng Zhang , Hao Zhao

Characterizing the relationship between neural population activity and behavioral data is a central goal of neuroscience. While latent variable models (LVMs) are successful in describing high-dimensional time-series data, they are typically…

机器学习 · 计算机科学 2026-02-12 Rabia Gondur , Usama Bin Sikandar , Evan Schaffer , Mikio Christian Aoi , Stephen L Keeley

Effective evaluation is critical for driving advancements in MLLM research. The surgical action planning (SAP) task, which aims to generate future action sequences from visual inputs, demands precise and sophisticated analytical…

计算机视觉与模式识别 · 计算机科学 2025-06-16 Mengya Xu , Zhongzhen Huang , Dillan Imans , Yiru Ye , Xiaofan Zhang , Qi Dou

We introduce MultiMedEval, an open-source toolkit for fair and reproducible evaluation of large, medical vision-language models (VLM). MultiMedEval comprehensively assesses the models' performance on a broad array of six multi-modal tasks,…

计算机视觉与模式识别 · 计算机科学 2026-03-19 Corentin Royer , Bjoern Menze , Anjany Sekuboyina

The practical deployment of medical vision-language models (Med-VLMs) necessitates seamless integration of textual data with diverse visual modalities, including 2D/3D images and videos, yet existing models typically employ separate…

计算与语言 · 计算机科学 2025-04-22 Songtao Jiang , Yuan Wang , Sibo Song , Yan Zhang , Zijie Meng , Bohan Lei , Jian Wu , Jimeng Sun , Zuozhu Liu