English
Related papers

Related papers: Multivariate Gaussian Representation Learning for …

200 papers

Recently, effective coordination in embodied multi-agent systems has remained a fundamental challenge, particularly in scenarios where agents must balance individual perspectives with global environmental awareness. Existing approaches…

Robotics · Computer Science 2025-11-04 Ziye Wang , Li Kang , Yiran Qin , Jiahua Ma , Zhanglin Peng , Lei Bai , Ruimao Zhang

Multimodal alignment constructs a joint latent vector space where modalities representing the same concept map to neighboring latent vectors. We formulate this as an inverse problem and show that, under certain conditions, paired data from…

Machine Learning · Computer Science 2025-06-10 Abhi Kamboj , Minh N. Do

3D semantic occupancy prediction is a pivotal task in autonomous driving, providing a dense and fine-grained understanding of the surrounding environment, yet single-modality methods face trade-offs between camera semantics and LiDAR…

Computer Vision and Pattern Recognition · Computer Science 2026-02-02 A. Enes Doruk , Hasan F. Ates

Surgical reconstruction of dynamic tissues from endoscopic videos is a crucial technology in robot-assisted surgery. The development of Neural Radiance Fields (NeRFs) has greatly advanced deformable tissue reconstruction, achieving…

Computer Vision and Pattern Recognition · Computer Science 2025-11-03 Wenfeng Huang , Xiangyun Liao , Yinling Qian , Hao Liu , Yongming Yang , Wenjing Jia , Qiong Wang

For many survey-based spatial modelling problems, responses are observed as spatially aggregated over survey regions due to limited resources. Covariates, from weather models and satellite imageries, can be observed at many different…

Applications · Statistics 2022-04-04 Harrison Zhu , Adam Howes , Owen van Eer , Maxime Rischard , Yingzhen Li , Dino Sejdinovic , Seth Flaxman

Surgical scene simulation plays a crucial role in surgical education and simulator-based robot learning. Traditional approaches for creating these environments with surgical scene involve a labor-intensive process where designers hand-craft…

Robotics · Computer Science 2024-08-07 Zhenya Yang , Kai Chen , Yonghao Long , Qi Dou

Temporally localizing actions in a video is a fundamental challenge in video understanding. Most existing approaches have often drawn inspiration from image object detection and extended the advances, e.g., SSD and Faster R-CNN, to produce…

Computer Vision and Pattern Recognition · Computer Science 2019-09-10 Fuchen Long , Ting Yao , Zhaofan Qiu , Xinmei Tian , Jiebo Luo , Tao Mei

Several recent works have directly extended the image masked autoencoder (MAE) with random masking into video domain, achieving promising results. However, unlike images, both spatial and temporal information are important for video…

Computer Vision and Pattern Recognition · Computer Science 2023-08-25 David Fan , Jue Wang , Shuai Liao , Yi Zhu , Vimal Bhat , Hector Santos-Villalobos , Rohith MV , Xinyu Li

Medicine is inherently multimodal and multitask, with diverse data modalities spanning text, imaging. However, most models in medical field are unimodal single tasks and lack good generalizability and explainability. In this study, we…

Computer Vision and Pattern Recognition · Computer Science 2024-10-01 Lijian Xu , Hao Sun , Ziyu Ni , Hongsheng Li , Shaoting Zhang

Recent advancements in foundation models for 2D vision have substantially improved the analysis of dynamic scenes from monocular videos. However, despite their strong generalization capabilities, these models often lack 3D consistency, a…

Computer Vision and Pattern Recognition · Computer Science 2025-12-04 Haoran Zhou , Gim Hee Lee

The emergence of neural rendering has significantly advanced the rendering quality of 3D human avatars, with the recently popular 3DGS technique enabling real-time performance. However, SMPL-driven 3DGS human avatars still struggle to…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Wangze Xu , Yifan Zhan , Zhihang Zhong , Xiao Sun

The predictive functions that permit humans to infer their body state by sensorimotor integration are critical to perform safe interaction in complex environments. These functions are adaptive and robust to non-linear actuators and noisy…

Robotics · Computer Science 2019-10-24 Pablo Lanillos , Gordon Cheng

Medical image segmentation is evolving from task-specific models toward generalizable frameworks. Recent research leverages Multi-modal Large Language Models (MLLMs) as autonomous agents, employing reinforcement learning with verifiable…

Computer Vision and Pattern Recognition · Computer Science 2026-02-04 Shengyuan Liu , Liuxin Bao , Qi Yang , Wanting Geng , Boyun Zheng , Chenxin Li , Wenting Chen , Houwen Peng , Yixuan Yuan

Visuospatial neglect is a disorder characterised by impaired awareness for visual stimuli located in regions of space and frames of reference. It is often associated with stroke. Patients can struggle with all aspects of daily living and…

Human-Computer Interaction · Computer Science 2023-10-24 Ivan De Boi , Elissa Embrechts , Quirine Schatteman , Rudi Penne , Steven Truijen , Wim Saeys

In robot-assisted minimally invasive surgery, accurate 3D reconstruction from endoscopic video is vital for downstream tasks and improved outcomes. However, endoscopic scenarios present unique challenges, including photometric…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Taoyu Wu , Yiyi Miao , Jiaxin Guo , Ziyan Chen , Sihang Zhao , Zhuoxiao Li , Zhe Tang , Baoru Huang , Limin Yu

While recent multimodal models have shown progress in vision-language tasks, small-scale variants still struggle with the fine-grained temporal reasoning required for video understanding. We introduce ReasonAct, a method that enhances video…

Computer Vision and Pattern Recognition · Computer Science 2025-11-27 Jiaxin Liu , Zhaolu Kang

To develop intelligent speech assistants and integrate them seamlessly with intra-operative decision-support frameworks, accurate and efficient surgical phase recognition is a prerequisite. In this study, we propose a multimodal framework…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-24 Kubilay Can Demir , Belen Lojo Rodriguez , Tobias Weise , Andreas Maier , Seung Hee Yang

For medical imaging AI models to be clinically impactful, they must generalize. However, this goal is hindered by (i) diverse types of distribution shifts, such as temporal, demographic, and label shifts, and (ii) limited diversity in…

Image and Video Processing · Electrical Eng. & Systems 2024-07-15 Kumail Alhamoud , Yasir Ghunaim , Motasem Alfarra , Thomas Hartvigsen , Philip Torr , Bernard Ghanem , Adel Bibi , Marzyeh Ghassemi

Continuous human motion understanding remains a core challenge in computer vision due to its high dimensionality and inherent redundancy. Efficient compression and representation are crucial for analyzing complex motion dynamics. In this…

Computer Vision and Pattern Recognition · Computer Science 2025-09-24 Gabriel Maldonado , Narges Rashvand , Armin Danesh Pazho , Ghazal Alinezhad Noghre , Vinit Katariya , Hamed Tabkhi

Clinical machine learning models are increasingly trained using large scale, multimodal foundation paradigms, yet deployment environments often differ systematically from the data generating settings used during training. Such shifts arise…

Machine Learning · Computer Science 2026-03-10 Yuanyun Zhang , Shi Li
‹ Prev 1 4 5 6 7 8 10 Next ›