English
Related papers

Related papers: Multivariate Gaussian Representation Learning for …

200 papers

While Multi-Agent Systems (MAS) show potential for complex clinical decision support, the field remains hindered by architectural fragmentation and the lack of standardized multimodal integration. Current medical MAS research suffers from…

Artificial Intelligence · Computer Science 2026-03-20 Yunhang Qian , Xiaobin Hu , Jiaquan Yu , Siyang Xin , Xiaokun Chen , Jiangning Zhang , Peng-Tao Jiang , Jiawei Liu , Hongwei Bran Li

Action Quality Assessment (AQA) -- the task of quantifying how well an action is performed -- has great potential for detecting errors in gym weight training, where accurate feedback is critical to prevent injuries and maximize gains.…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Hao Yin , Lijun Gu , Paritosh Parmar , Lin Xu , Tianxiao Guo , Xiujin Liu , Weiwei Fu , Yang Zhang , Tianyou Zheng

Easy access to precise 3D tracking of movement could benefit many aspects of rehabilitation. A challenge to achieving this goal is that while there are many datasets and pretrained algorithms for able-bodied adults, algorithms trained on…

Computer Vision and Pattern Recognition · Computer Science 2023-10-31 R. James Cotton , Colleen Peyton

Despite significant advancements in general AI, its effectiveness in the medical domain is limited by the lack of specialized medical knowledge. To address this, we formulate GMAI-VL-5.5M, a multimodal medical dataset created by converting…

Computer Vision and Pattern Recognition · Computer Science 2025-03-28 Tianbin Li , Yanzhou Su , Wei Li , Bin Fu , Zhe Chen , Ziyan Huang , Guoan Wang , Chenglong Ma , Ying Chen , Ming Hu , Yanjun Li , Pengcheng Chen , Xiaowei Hu , Zhongying Deng , Yuanfeng Ji , Jin Ye , Yu Qiao , Junjun He

State-of-the-art approaches for conditional human body rendering via Gaussian splatting typically focus on simple body motions captured from many views. This is often in the context of dancing or walking. However, for more complex use…

Computer Vision and Pattern Recognition · Computer Science 2025-05-06 Maksym Ivashechkin , Oscar Mendez , Richard Bowden

We present GaussianAvatar, an efficient approach to creating realistic human avatars with dynamic 3D appearances from a single video. We start by introducing animatable 3D Gaussians to explicitly represent humans in various poses and…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Liangxiao Hu , Hongwen Zhang , Yuxiang Zhang , Boyao Zhou , Boning Liu , Shengping Zhang , Liqiang Nie

Visual task adaptation has been demonstrated to be effective in adapting pre-trained Vision Transformers (ViTs) to general downstream visual tasks using specialized learnable layers or tokens. However, there is yet a large-scale benchmark…

Computer Vision and Pattern Recognition · Computer Science 2024-04-22 Shentong Mo , Xufang Luo , Yansen Wang , Dongsheng Li

Simultaneous localization and mapping is essential for position tracking and scene understanding. 3D Gaussian-based map representations enable photorealistic reconstruction and real-time rendering of scenes using multiple posed cameras. We…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Lisong C. Sun , Neel P. Bhatt , Jonathan C. Liu , Zhiwen Fan , Zhangyang Wang , Todd E. Humphreys , Ufuk Topcu

Multi-output Gaussian processes (GPs) are a flexible Bayesian nonparametric framework that has proven useful in jointly modeling the physiological states of patients in medical time series data. However, capturing the short-term effects of…

Recently, generalizable human Gaussian splatting from sparse-view inputs has been actively studied for the photorealistic human rendering. Most existing methods rely on explicit geometric constraints or predefined structural representations…

Computer Vision and Pattern Recognition · Computer Science 2026-04-29 Jingi Kim , Wonjun Kim

Compressed video action recognition has recently drawn growing attention, since it remarkably reduces the storage and computational cost via replacing raw videos by sparsely sampled RGB frames and compressed motion cues (e.g., motion…

Computer Vision and Pattern Recognition · Computer Science 2025-10-06 Bing Li , Jiaxin Chen , Dongming Zhang , Xiuguo Bao , Di Huang

Visual search in 3D environments requires embodied agents to actively explore their surroundings and acquire task-relevant evidence. However, existing visual search and embodied AI benchmarks, including EQA, typically rely on static…

Computer Vision and Pattern Recognition · Computer Science 2026-04-24 Koya Sakamoto , Taiki Miyanishi , Daichi Azuma , Shuhei Kurita , Shu Morikuni , Naoya Chiba , Motoaki Kawanabe , Yusuke Iwasawa , Yutaka Matsuo

Understanding dynamic scenes from casual videos is critical for scalable robot learning, yet four-dimensional (4D) reconstruction under strictly monocular settings remains highly ill-posed. To address this challenge, our key insight is that…

Computer Vision and Pattern Recognition · Computer Science 2026-02-17 Can Li , Jie Gu , Jingmin Chen , Fangzhou Qiu , Lei Sun

Self-supervised learning is an efficient pre-training method for medical image analysis. However, current research is mostly confined to specific-modality data pre-training, consuming considerable time and resources without achieving…

Computer Vision and Pattern Recognition · Computer Science 2023-12-01 Yiwen Ye , Yutong Xie , Jianpeng Zhang , Ziyang Chen , Qi Wu , Yong Xia

Real-time multi-agent collaboration for ego-motion estimation and high-fidelity 3D reconstruction is vital for scalable spatial intelligence. However, traditional methods produce sparse, low-detail maps, while recent dense mapping…

Computer Vision and Pattern Recognition · Computer Science 2024-12-16 Xiaohao Xu , Feng Xue , Shibo Zhao , Yike Pan , Sebastian Scherer , Xiaonan Huang

Accurate analysis of cardiac motion is crucial for evaluating cardiac function. While dynamic cardiac magnetic resonance imaging (CMR) can capture detailed tissue motion throughout the cardiac cycle, the fine-grained 4D cardiac motion…

Computer Vision and Pattern Recognition · Computer Science 2025-07-23 Xueming Fu , Pei Wu , Yingtai Li , Xin Luo , Zihang Jiang , Junhao Mei , Jian Lu , Gao-Jun Teng , S. Kevin Zhou

Medicine is inherently multimodal, with rich data modalities spanning text, imaging, genomics, and more. Generalist biomedical artificial intelligence (AI) systems that flexibly encode, integrate, and interpret this data at scale can…

We introduce MIGS (Multi-Identity Gaussian Splatting), a novel method that learns a single neural representation for multiple identities, using only monocular videos. Recent 3D Gaussian Splatting (3DGS) approaches for human avatars require…

Computer Vision and Pattern Recognition · Computer Science 2024-07-18 Aggelina Chatziagapi , Grigorios G. Chrysos , Dimitris Samaras

We introduce MGP-VAE (Multi-disentangled-features Gaussian Processes Variational AutoEncoder), a variational autoencoder which uses Gaussian processes (GP) to model the latent space for the unsupervised learning of disentangled…

Computer Vision and Pattern Recognition · Computer Science 2020-07-21 Sarthak Bhagat , Shagun Uppal , Zhuyun Yin , Nengli Lim

Emotion recognition is relevant for human behaviour understanding, where facial expression and speech recognition have been widely explored by the computer vision community. Literature in the field of behavioural psychology indicates that…

Computer Vision and Pattern Recognition · Computer Science 2024-05-24 Maria Luísa Lima , Willams de Lima Costa , Estefania Talavera Martinez , Veronica Teichrieb